Book cover recognition method and device, electronic equipment and storage medium

By performing text recognition and feature extraction on book cover images, combined with publisher classification and feature search, the problem of inaccurate recognition in OCR technology when the image quality is poor has been solved, achieving higher recognition accuracy.

CN116612488BActive Publication Date: 2025-11-25深圳市星桐科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310603847.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-25
Publication Date
2025-11-25
Estimated Expiration
2043-05-25

AI Technical Summary

Technical Problem

Existing OCR technology cannot accurately identify book information when recognizing book covers with poor image quality.

Method used

By performing text recognition on book cover images, the text content is obtained and then input into a pre-trained publisher classification model. Feature extraction is performed in conjunction with a feature extraction model, and the book information is determined by searching using the text content and feature vectors.

Benefits of technology

It improves the accuracy of book cover recognition, enabling accurate identification of book information even with poor image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612488B_ABST
    Figure CN116612488B_ABST
Patent Text Reader

Abstract

The present disclosure provides a book cover recognition method and device, electronic equipment and storage medium, the method comprising: performing text recognition on a book cover image to be recognized to obtain text content contained in the book cover image; inputting the book cover image into a pre-trained press classification model to obtain a target press contained in the book cover image; based on the text content, the target press and the book cover image, performing feature extraction using a pre-trained feature extraction model to obtain a feature vector corresponding to the book cover image; performing text search based on the text content to obtain a text search result, and performing vector search based on the feature vector to obtain a vector search result; and determining target book information corresponding to the book cover image according to the text search result and the vector search result. The present scheme can search for more reliable vector search results, thereby improving the accuracy of book cover recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and in particular to a method, apparatus, electronic device and storage medium for recognizing book covers. Background Technology

[0002] Currently, the recognition of book covers is usually achieved using Optical Character Recognition (OCR) technology. OCR technology identifies the text contained in the book cover and then returns information such as the book's name and publisher.

[0003] OCR technology performs well on book covers with good image quality and can return accurate book information. However, when the image quality is poor, OCR technology usually cannot accurately identify the book information. Summary of the Invention

[0004] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, the present disclosure provides a method, apparatus, electronic device and storage medium for identifying book covers.

[0005] According to one aspect of this disclosure, a method for identifying a book cover is provided, comprising:

[0006] Text recognition is performed on the book cover image to be identified to obtain the text content contained in the book cover image;

[0007] The book cover image is input into a pre-trained publisher classification model to obtain the target publisher contained in the book cover image;

[0008] Based on the text content, the target publisher, and the book cover image, a pre-trained feature extraction model is used to extract features to obtain the feature vector corresponding to the book cover image.

[0009] A text search is performed based on the text content to obtain a text search result, and a vector search is performed based on the feature vector to obtain a vector search result. The text search result and the vector search result each contain at least one book information.

[0010] Based on the text search results and the vector search results, the target book information corresponding to the book cover image is determined.

[0011] According to another aspect of this disclosure, a book cover identification device is provided, comprising:

[0012] The first recognition module is used to perform text recognition on the book cover image to be recognized, and obtain the text content contained in the book cover image;

[0013] The second recognition module is used to input the book cover image into a pre-trained publisher classification model to obtain the target publisher contained in the book cover image;

[0014] The feature extraction module is used to extract features based on the text content, the target publisher, and the book cover image using a pre-trained feature extraction model, and obtain the feature vector corresponding to the book cover image.

[0015] The search module is used to perform a text search based on the text content to obtain a text search result, and to perform a vector search based on the feature vector to obtain a vector search result, wherein the text search result and the vector search result each contain at least one book information.

[0016] The determination module is used to determine the target book information corresponding to the book cover image based on the text search results and the vector search results.

[0017] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0018] Processor; and

[0019] Stored program memory,

[0020] The program includes instructions that, when executed by the processor, cause the processor to perform the book cover recognition method according to one of the foregoing aspects.

[0021] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the book cover recognition method according to the foregoing aspect.

[0022] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the book cover recognition method described in the foregoing aspect.

[0023] One or more technical solutions provided in this disclosure involve performing text recognition on a book cover image to obtain the text content contained in the book cover image, inputting the book cover image into a pre-trained publisher classification model to obtain the target publisher corresponding to the book cover, then performing feature extraction based on the text content, target publisher, and book cover image using a pre-trained feature extraction model to obtain the feature vector corresponding to the book cover image, and then performing a text search based on the text content to obtain a text search result, and performing a vector search based on the feature vector to obtain a vector search result, and finally determining the target book information corresponding to the book cover image based on the text search result and the vector search result. By adopting the solution of this disclosure, the corresponding target publisher is obtained by identifying the publisher of the book cover image, and the identified target publisher is used for feature extraction to obtain the feature vector corresponding to the book cover image. This results in a more robust feature vector, which leads to more reliable vector search results when searching based on the feature vector, thus improving the accuracy of book cover recognition. Attached Figure Description

[0024] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0025] Figure 1 A flowchart illustrating a method for recognizing a book cover according to an exemplary embodiment of the present disclosure is shown;

[0026] Figure 2 A flowchart illustrating a method for recognizing a book cover according to another exemplary embodiment of this disclosure is shown;

[0027] Figure 3 A flowchart illustrating a method for recognizing a book cover according to yet another exemplary embodiment of this disclosure is shown;

[0028] Figure 4 A schematic block diagram of a book cover recognition device according to an exemplary embodiment of the present disclosure is shown;

[0029] Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0033] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0034] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0035] The following description, with reference to the accompanying drawings, outlines the methods, devices, electronic equipment, and storage media for identifying book covers provided in this disclosure.

[0036] In educational settings, it's common to encounter situations where a photo of a book cover is taken and the system identifies the book's title, publisher, and other information. For example, in a book knowledge activation application, a student uploads an image of a supplementary textbook cover. The Q&A system recognizes the cover image to determine the book's title, publisher, and other information, then automatically activates the knowledge points within the book. The student can then view the book's pages within the Q&A system.

[0037] Currently, OCR technology is commonly used for recognizing book covers. For high-quality book cover images taken with a mobile phone, OCR technology can produce relatively accurate results and return the correct book title, publisher, and other book information. However, for blurry book cover images or those taken directly from a computer screen and containing moiré patterns, the recognition effect of OCR technology is poor, and it cannot accurately identify the book information.

[0038] To address the aforementioned problems, this disclosure provides a method for recognizing book covers. The method involves performing text recognition on the book cover image to obtain the text content contained within the cover. The book cover image is then input into a pre-trained publisher classification module to identify the target publisher corresponding to the cover. Next, based on the text content, target publisher, and book cover image, a pre-trained feature extraction model is used to extract features, resulting in a feature vector corresponding to the book cover image. Subsequently, a text search is performed based on the text content to obtain text search results, and a vector search is performed based on the feature vector to obtain vector search results. Finally, based on the text search results and vector search results, the target book information corresponding to the book cover image is determined. By employing the scheme of this disclosure, publisher recognition is performed on the book cover image to obtain the corresponding target publisher, and the identified target publisher is used for feature extraction to obtain the feature vector corresponding to the book cover image. This results in a more robust feature vector, leading to more reliable vector search results during feature vector-based searches and improving the accuracy of book cover recognition.

[0039] Figure 1 A flowchart of a book cover recognition method according to an exemplary embodiment of the present disclosure is shown. The method can be executed by a book cover recognition device provided in the embodiments of the present disclosure, wherein the device can be implemented in software and / or hardware, and can generally be integrated into an electronic device, including mobile phones, tablets, wearable devices, etc.

[0040] like Figure 1 As shown, the method for identifying the book cover may include the following steps:

[0041] Step 101: Perform text recognition on the book cover image to be identified to obtain the text content contained in the book cover image.

[0042] The book cover image to be identified can be, for example, the cover image of a textbook, a tutoring book, a picture book, etc., and can be input by taking a picture and uploading it, or by taking a screenshot.

[0043] In this embodiment of the disclosure, for the data cover image to be identified, text recognition can be performed to obtain the text content contained in the book cover image.

[0044] The text content may include, but is not limited to, the book title, the grade level to which the book is applicable, the publisher, and the author.

[0045] For example, OCR technology can be used to perform text recognition on a book cover image to be recognized. This OCR recognition of the book cover image can be divided into two parts: text detection and text line recognition. In the text detection part, the book cover image to be recognized can be input into a pre-trained text detection model for forward inference. The text detection model outputs the position coordinate information of each text line in the book cover image. Considering that the text being detected is contained in the book cover, and that user-uploaded book cover images are diverse, with text lines not always horizontal, but also including vertical and slanted artistic text, all pose challenges to text detection. Experiments have shown that the Mask R-CNN model achieves the best detection efficiency and accuracy in book cover text detection compared to other models. Therefore, in this embodiment, the Mask R-CNN detection model is used to detect the position coordinate information of text lines in the book cover image. It is understood that other text detection models, such as the EAST model and the PSENet model, can also be used in practical applications. Next, in the text line recognition section, the corresponding text image can be extracted from the book cover image based on the position coordinates of the text lines output by the text detection section. The vertical text is then uniformly converted to horizontal text; for example, the vertical text can be rotated 90 degrees counterclockwise. Next, the extracted text images can be corrected by preprocessing the corrected text data, such as scaling to a standard size, adding whitespace, or standardizing the image data. Then, the preprocessed text data is input into a pre-trained text recognition model for forward inference. The text recognition model outputs the text information corresponding to the text images, thus obtaining the text content contained in the book cover image. In this embodiment, a CRNN+Attention model can be used as the text recognition model. In practical applications, other text recognition models, such as the CRNN+CTC model, can also be used.

[0046] In one optional embodiment of this disclosure, before performing text recognition on the book cover image to be recognized, the book cover image to be recognized may be preprocessed. The preprocessing may include, but is not limited to, scaling the image to a fixed size, standardizing the image data, etc.

[0047] Step 102: Input the book cover image into a pre-trained publisher classification model to obtain the target publisher contained in the book cover image.

[0048] In this embodiment of the disclosure, for a book cover image to be identified, in addition to performing text recognition to obtain the text content contained in the book cover image, publisher identification is also performed to obtain the publisher contained in the book cover image, referred to as the target publisher. Specifically, the book cover image to be identified can be input into a pre-trained publisher classification model, which identifies the publisher contained in the book cover image and outputs the target publisher contained in the book cover image.

[0049] The target publisher may be the name of the identified publisher, the category number of the identified publisher, or other forms of publisher representation. This disclosure does not limit the form of representation of the target publisher.

[0050] For example, the publisher classification model can use the ResNet18 model, which is highly efficient and performs well in image classification. Other lightweight classification models, such as MobileNet and SqueezeNet, can also be used in practical applications.

[0051] In one optional embodiment of this disclosure, before performing text recognition on the book cover image to be recognized, the book cover image to be recognized may be preprocessed. The preprocessing may include, but is not limited to, scaling the image to a fixed size, standardizing the image data, etc.

[0052] It should be noted that in this embodiment, the execution order of steps 101 and 102 is not important; they can be executed sequentially or simultaneously. Figure 1 The embodiments shown are merely examples of how step 102 is performed after step 101 to illustrate this disclosure, and should not be construed as limiting the scope of this disclosure.

[0053] Step 103: Based on the text content, the target publisher, and the book cover image, feature extraction is performed using a pre-trained feature extraction model to obtain the feature vector corresponding to the book cover image.

[0054] In this embodiment of the disclosure, after obtaining the text content contained in the book cover image and the target publisher, feature extraction can be performed using a pre-trained feature extraction model based on the text content, the target publisher, and the book cover image to be identified, to obtain the feature vector corresponding to the book cover image.

[0055] For example, several sample cover images can be collected in advance, and text recognition and publisher recognition can be performed on the sample cover images to obtain the corresponding sample text content and sample publisher. The sample cover image, the corresponding sample text content, and the sample publisher are used as a training sample, and the feature vector corresponding to the training sample is obtained as the training target. The initial feature extraction model is trained using the training sample and the corresponding training target to obtain the trained feature extraction model. Then, when obtaining the feature vector corresponding to the book cover image to be identified, the book cover image to be identified, the corresponding text content, and the target publisher can be input together into the trained feature extraction model, and the output of the feature extraction model is obtained as the feature vector corresponding to the book cover image.

[0056] Step 104: Perform a text search based on the text content to obtain text search results, and perform a vector search based on the feature vector to obtain vector search results.

[0057] In this embodiment of the disclosure, after obtaining the text content contained in the book cover image, a text search can be performed based on the text content to obtain text search results. The text search results may include one or more book information entries. This disclosure does not limit the number of book information entries included in the obtained text search results. Book information may include, but is not limited to, book title, publisher, and applicable grade level.

[0058] For example, queries can be performed on a pre-indexed text search library based on text content. During the query, the similarity between the text content and each index in the text search library is calculated, and the search results corresponding to each index are sorted in descending order of similarity. The top N search results are returned as the text search results, where N is a positive integer. The text search library can be an Elasticsearch library. In this embodiment, when indexing the text search library, precise book cover text content can be used, such as publisher, grade, and book title. The book cover text content used for indexing is manually verified to ensure the accuracy of the index. OCR text recognition content is not used for indexing because errors in OCR-recognized text content can affect search accuracy; therefore, manually verified precise text content is used for indexing.

[0059] In this embodiment of the disclosure, after obtaining the feature vector corresponding to the book cover image, a vector search is performed based on the feature vector to obtain a vector search result. The vector search result may include one or more book information entries. This disclosure does not limit the number of book information entries included in the obtained vector search result. Book information may include, but is not limited to, book title, publisher, and applicable grade level.

[0060] For example, a query can be performed in a pre-indexed vector search library based on the acquired feature vector. During the query, the similarity between the feature vector and each index in the vector search library is calculated, and the search results corresponding to each index are sorted in descending order of similarity. The top N search results are returned as the vector search results, where N is a positive integer. Each index in the vector search library is represented in vector form, and the feature vectors corresponding to the book information in the vector search library can be used as the indexes for the search content.

[0061] It is understood that in the embodiments of this disclosure, the number of search results contained in the text search results and the vector search results may be the same or different, and this disclosure does not limit the number of search results contained in the text search results and the vector search results respectively.

[0062] Step 105: Determine the target book information corresponding to the book cover image based on the text search results and the vector search results.

[0063] In this embodiment of the disclosure, after obtaining the text search results and the vector search results, the target book information corresponding to the book cover image can be determined based on the text search results and the vector search results.

[0064] For example, candidate book information containing the target publisher can be filtered from both text search results and vector search results, and then the book information with the highest similarity can be selected as the target book information. It is understood that the similarity here refers to the similarity calculated when searching in the text search library and the vector search library to obtain text search results and vector search results.

[0065] For example, the book information with the highest similarity can be selected from text search results and vector search results as the target book information corresponding to the book cover image.

[0066] The book cover recognition method of this disclosure involves performing text recognition on the book cover image to obtain the text content contained in the book cover image, inputting the book cover image into a pre-trained publisher classification model to obtain the target publisher corresponding to the book cover, then performing feature extraction based on the text content, target publisher, and book cover image using a pre-trained feature extraction model to obtain the feature vector corresponding to the book cover image, and then performing a text search based on the text content to obtain a text search result, and performing a vector search based on the feature vector to obtain a vector search result, and finally determining the target book information corresponding to the book cover image based on the text search result and the vector search result. By adopting the scheme of this disclosure, the corresponding target publisher is obtained by identifying the publisher of the book cover image, and the identified target publisher is used for feature extraction to obtain the feature vector corresponding to the book cover image. This results in a more robust feature vector, which leads to more reliable vector search results when searching based on the feature vector, thus improving the accuracy of book cover recognition.

[0067] In one optional embodiment of this disclosure, the feature extraction model includes a first output branch and a second output branch. The first output branch is used to output a feature vector, and the second output branch is used to output the publisher identification result. When the feature extraction model is trained, iterative training is performed based on the feature vector output by the first output branch and the publisher identification result output by the second output branch.

[0068] In this embodiment, the feature extraction model includes two output branches, denoted as the first output branch and the second output branch. The first output branch outputs a feature vector of fixed dimensions, and the second output branch outputs the publisher identification result. During the training of the feature extraction model, both the first and second output branches are trained simultaneously. The initial feature extraction model is iteratively trained based on the feature vector output by the first output branch and the publisher identification result output by the second output branch until a well-trained feature extraction model is obtained. Therefore, by setting two different output branches to complete multi-task learning, the feature learning capability of the feature extraction model can be improved, enabling the final feature extraction model to extract better features.

[0069] For example, a Swin transformer model can be used, with an output branch added to its network structure to obtain an initial feature extraction model. This initial feature extraction model can then be trained to obtain a new feature extraction model. The added output branch is a fully connected layer used to output the publisher identification result, such as the publisher category number. In practical applications, other network models can also be used.

[0070] Furthermore, in one alternative embodiment of this disclosure, such as Figure 2 As shown, in Figure 1 Based on the illustrated embodiment, step 103 may include the following sub-steps:

[0071] Step 201: Convert the word vectors corresponding to the text content with the target publisher to obtain a text data matrix, wherein the number of columns in the text data matrix is ​​the same as the number of columns in the image data matrix corresponding to the book cover image.

[0072] For example, the word vectors corresponding to the text content can be obtained by converting them using the word2vector tool, or by other methods, such as using a pre-trained word vector model. This disclosure does not limit the method of obtaining the word vectors corresponding to the text content.

[0073] In this embodiment of the disclosure, the word vectors corresponding to the text content and the target publisher can be converted into a text data matrix.

[0074] For example, assuming the target publisher is represented by a publisher classification number (i.e., the target publisher is a number), when converting the word vectors corresponding to the text content with the target publisher, the target publisher can be copied according to the dimension of the word vectors corresponding to the text content, and then concatenated with the word vectors corresponding to the text content to obtain a text data matrix. For instance, assuming the dimension of the word vectors corresponding to the text content is 1*C, and the word vectors corresponding to each word in the text content form a word vector matrix with a dimension of W*C, where W is the number of rows, consistent with the number of words in the text content, and C is the number of columns, consistent with the number of columns in the image data matrix corresponding to the book cover image, then the target publisher can be copied (C-1) times according to the number of columns C. The C target publishers form a vector with a dimension of 1*C, which is then concatenated after the word vector matrix to obtain a text data matrix with a dimension of (W+1)*C.

[0075] For example, assuming the target publisher is represented by its name, when converting the word vectors corresponding to the text content with the target publisher's name, the target publisher can first be converted into its corresponding word vector, and then the word vectors corresponding to the text content and the target publisher can be concatenated to obtain a text data matrix. It should be noted that in this embodiment, the concatenation direction between the word vectors corresponding to the text content and the target publisher is not limited; they can be concatenated in either the row or column direction, as long as the number of columns in the resulting text data matrix matches the number of columns in the image data matrix corresponding to the book cover image.

[0076] Step 202: Concatenate the image data matrix and the text data matrix in the row direction to generate book cover data.

[0077] In this embodiment of the disclosure, the obtained text data matrix can be concatenated with the image data matrix corresponding to the book cover image in the row direction to obtain the book cover data. Since the number of columns in the text data matrix is ​​the same as the number of columns in the image data matrix corresponding to the book cover image, the two can be concatenated in the row direction without changing the number of columns in the concatenated matrix.

[0078] For example, assuming the text data matrix has a dimension of D*C and the image data matrix corresponding to the book cover image has a dimension of H*C, concatenating the image data matrix and the text data matrix along the row direction yields the book cover data with a dimension of (H+D)*C. Here, H is the height of the image data matrix corresponding to the book cover image, which can be represented by the number of pixel rows in the height direction of the book cover image, and C is the width of the image data matrix corresponding to the book cover image, which can be represented by the number of pixel columns in the width direction of the book cover image.

[0079] Step 203: Input the book cover data into a pre-trained feature extraction model, and obtain the feature vector output by the first output branch of the feature extraction model as the feature vector corresponding to the book cover image.

[0080] In this embodiment of the disclosure, after obtaining book cover data based on text content, target publisher, and book cover image, the book cover data can be input into a pre-trained feature extraction model, and the feature vector output by the first output branch of the feature extraction model can be obtained. The obtained feature vector is used as the feature vector corresponding to the book cover image.

[0081] The book cover recognition method of this disclosure converts the word vectors corresponding to the text content with the target publisher to obtain a text data matrix. Then, the image data matrix and the text data matrix are concatenated in the row direction to generate book cover data. The book cover data is then input into a pre-trained feature extraction model, and the feature vector output by the first output branch of the feature extraction model is obtained as the feature vector corresponding to the book cover image. Thus, the model can automatically extract text features, publisher features and image features related to the book cover, so that the feature vectors learned by the model have good generalization ability.

[0082] In one alternative embodiment of this disclosure, such as Figure 3 As shown, based on the foregoing embodiments, step 105 may include the following sub-steps:

[0083] Step 301: Obtain the search result with the highest similarity to the text content from the text search results as the first candidate book information.

[0084] Step 302: If the first candidate book information meets the first preset condition, the first candidate book information is determined as the target book information corresponding to the book cover image.

[0085] The first preset condition includes:

[0086] The publisher information corresponding to the first candidate book information is consistent with the target publisher;

[0087] The text data corresponding to the first candidate book information is consistent with the text content; and

[0088] The confidence level corresponding to the text content is greater than the preset confidence threshold.

[0089] In this embodiment of the disclosure, for the obtained text search results, the book information with the highest similarity to the text content can be obtained from the text search results first, and the book information with the highest similarity to the text content can be used as the first candidate book information. Next, it can be determined whether the first candidate book information meets the first preset condition, and if the first candidate book information meets the first preset condition, the first candidate book information is determined as the target book information corresponding to the book cover image.

[0090] In other words, for the first candidate book information determined from the text search results, it can be determined whether the publisher information corresponding to the first candidate book information, such as the publisher's name and classification number, is consistent with the target publisher obtained by publisher recognition of the closed book image. It can also be determined whether the text data corresponding to the first candidate book information, such as publisher information, grade information, and book title, is consistent with the text content obtained by text recognition of the closed book image. Finally, it can be determined whether the confidence level of the text content is greater than a preset confidence threshold. If all three conditions are met, the first candidate book information is determined to meet the first preset condition, and the first candidate book information is directly returned as the target book information. Specifically, when performing text recognition on the book cover image, not only is the corresponding text content output, but also the confidence level of each text block within the text content. The higher the confidence level, the higher the accuracy of the corresponding text block. When the text content contains multiple text blocks, such as publisher information, grade information, and book title each being a separate text block, the confidence level of each text block can be compared to see if they are all greater than the corresponding confidence threshold. Only when all are greater than the threshold is the confidence level of the text content considered to be greater than the preset confidence threshold.

[0091] Step 303: If the first candidate book information does not meet the first preset condition, obtain the book information with the highest similarity to the feature vector in the vector search results as the second candidate book information.

[0092] Step 304: If the second candidate book information meets the second preset condition, the second candidate book information is determined as the target book information corresponding to the book cover image.

[0093] The second preset condition includes:

[0094] The information of the second candidate book is consistent with the information of the first candidate book; and

[0095] The publisher information corresponding to the second candidate book information is consistent with the target publisher.

[0096] In this embodiment of the disclosure, when the first candidate book information does not meet the first preset condition, the book information with the highest similarity to the feature vector can be obtained from the vector search results, and the book information with the highest similarity to the feature vector can be used as the second candidate book information. Then, it can be determined whether the second candidate book information meets the second preset condition, and if the second candidate book information meets the second preset condition, the second candidate book information is determined as the target book information corresponding to the book cover image.

[0097] In other words, in this embodiment of the disclosure, for the second candidate book information determined from the vector search results, it can be determined whether the second candidate book information is consistent with the first candidate book information, that is, whether the second candidate book information is the same book information as the first candidate book information, and whether the publisher information corresponding to the second candidate book information, such as the publisher's name, classification number, etc., is consistent with the target publisher obtained by publisher identification of the closed image of the book. If both of the above conditions are met, it is determined that the second candidate book information meets the second preset condition, and the second candidate book information is directly returned as the target book information.

[0098] Step 305: If the second candidate book information does not meet the second preset condition, traverse the text search results to determine whether the second candidate book information exists in the text search results.

[0099] Step 306: If the second candidate book information exists in the text search results, increase the similarity of the second candidate book information by a first preset value to obtain a new similarity of the second candidate book information.

[0100] The first preset value can be preset according to actual needs.

[0101] It is understandable that the similarity of the second candidate book information is the similarity between the index of the second candidate book information in the vector search library and the feature vector when the vector search results are obtained based on the feature vector.

[0102] Step 307: If the new similarity of the second candidate book information is greater than the first threshold, and the publisher information corresponding to the second candidate book information is consistent with the target publisher, then the second candidate book information is determined as the target book information corresponding to the book cover image.

[0103] The first threshold can be preset according to actual needs, and this disclosure does not restrict the specific value of the first threshold.

[0104] In this embodiment of the disclosure, when the second candidate book information does not meet the second preset condition, the text search results can be traversed to determine whether the second candidate book information exists in the text search results. If the second candidate book information is not found in the text search results, step 308 is executed. If the second candidate book information is found in the text search results, it is determined that the second candidate book information exists in the text search results, and the similarity of the second candidate book information is further increased by a first preset value to obtain a new similarity of the second candidate book information. Next, it is determined whether the new similarity of the second candidate book information is greater than a first threshold, and whether the publisher information (e.g., publisher's name, publisher's classification number, etc.) corresponding to the second candidate book information is consistent with the target publisher. If the new similarity of the second candidate book information is greater than the first threshold, and the publisher information corresponding to the second candidate book information is consistent with the target publisher, then the second candidate book information is determined as the target book information corresponding to the book cover image. If the new similarity of the second candidate book information is not greater than the first threshold, or the publisher information corresponding to the second candidate book information is inconsistent with the target publisher, then step 308 is executed.

[0105] In this embodiment of the disclosure, when the second candidate book information does not meet the second preset condition, it is determined whether the second candidate book information exists in the text search results. If it exists, the similarity of the second candidate book information is increased by a first preset value. This increases the probability that the book information with the highest similarity in the vector search results is identified as the target book information, thereby improving the accuracy of the returned results.

[0106] Step 308: If the second candidate book information does not exist in the text search results, or the new similarity of the second candidate book information is not greater than the first threshold, or the publisher information corresponding to the second candidate book information is inconsistent with the target publisher, the vector search results are traversed to determine whether the first candidate book information exists in the vector search results.

[0107] Step 309: If the first candidate book information exists in the vector search results, increase the similarity of the first candidate book information by a second preset value to obtain a new similarity of the first candidate book information.

[0108] The second preset value can be preset according to actual needs.

[0109] It is understandable that the similarity of the first candidate book information is the similarity between the index of the first candidate book information in the text search library and the text content when obtaining text search results based on text content.

[0110] Step 310: If the new similarity of the first candidate book information is greater than the second threshold, and the publisher information corresponding to the first candidate book information is consistent with the target publisher, then the first candidate book information is determined as the target book information corresponding to the book cover image.

[0111] The second threshold can be preset according to actual needs, and this disclosure does not restrict the specific value of the second threshold.

[0112] In this embodiment of the disclosure, when there is no second candidate book information in the text search results, or the new similarity of the second candidate book information is not greater than the first threshold, or the publisher information corresponding to the second candidate book information is inconsistent with the target publisher, the vector search results can be traversed to determine whether the first candidate book information exists in the vector search results. If the first candidate book information is found in the vector search results, it is determined that the first candidate book information exists in the vector search results. At this time, the similarity of the first candidate book information can be increased by a second preset value to obtain a new similarity of the first candidate book information. This new similarity is then compared with the second threshold. If the new similarity of the first candidate book information is greater than the second threshold, and the publisher information corresponding to the first candidate book information (e.g., the publisher's name, the publisher's classification number, etc.) is consistent with the target publisher, then the first candidate book information can be determined as the target book information corresponding to the book cover image.

[0113] In this embodiment of the disclosure, by traversing the vector search results to determine whether there is first candidate book information in the vector search results, and if there is, increasing the similarity of the first candidate book information by a second preset value, the probability of the book information with the highest similarity in the text search results being identified as the target book information can be increased, thereby improving the accuracy of the returned results.

[0114] In this embodiment, steps 301-302 above constitute judgment strategy one when determining target book information based on text search results and vector search results. When integrating text search results and vector search results to determine target book information, judgment strategy one is executed first. If judgment strategy one is satisfied, the first candidate book information is directly returned as the target book information. If judgment strategy one is not satisfied, judgment strategy two can be further judged (corresponding to steps 303-304 above). If judgment strategy two is satisfied, the target book information is returned. If judgment strategy two is not satisfied, judgment strategy three can be further judged (corresponding to steps 305-307 above). In the judgment strategy two, if no second candidate book information is found in the text search results, judgment strategy two is also determined to be unsatisfactory. If judgment strategy three is satisfied, the target book information is returned. If judgment strategy three is not satisfied, judgment strategy four can be further judged (corresponding to steps 308-310 above). If judgment strategy four is satisfied, the target book information is returned. If judgment strategy four is not satisfied, and the publisher information corresponding to the first candidate book information is inconsistent with the target publisher, this disclosure proposes judgment strategy five, including:

[0115] If the publisher information corresponding to the first candidate book information is inconsistent with the target publisher, the publisher in the text content is replaced with the target publisher to obtain new text content;

[0116] Based on the new text content, a new text search is performed, and the search results are sorted in descending order of similarity to the new text content to obtain the top M book information, where M is a positive integer;

[0117] Based on the M book information and the text search results, generate new text search results;

[0118] Based on preset similarity adjustment rules, the similarity between the new text search results and the vector search results is adjusted;

[0119] Based on the adjusted similarity, the book information with the highest similarity between the new text search results and the vector search results is determined as the target book information corresponding to the book cover image.

[0120] In one optional embodiment of this disclosure, when adjusting the similarity between new text search results and vector search results based on preset similarity adjustment rules, it can be achieved by at least one of the following methods: (1) increasing the similarity of book information in the new text search results that is consistent with the target publisher by a third preset value, and increasing the similarity of book information in the vector search results that is consistent with the target publisher by a fourth preset value; (2) traversing the new text search results and the vector search results, and when the third candidate book information in the new text search results is consistent with the fourth candidate book information in the vector search results, increasing the similarity of the third candidate book information by a fifth preset value, and increasing the similarity of the fourth candidate book information by a sixth preset value.

[0121] The third, fourth, fifth, and sixth preset values ​​can all be preset according to actual needs.

[0122] It is understood that the methods for adjusting the similarity between new text search results and vector search results are not limited to the two methods provided in the above embodiments of this disclosure, and the similarity of search results can also be adjusted in other ways.

[0123] In this embodiment, when the publisher information corresponding to the first candidate book information is inconsistent with the target publisher, the publisher in the text content can be replaced with the target publisher to obtain new text content. A new text search is then performed in the text search library based on this new text content. The new search results are sorted in descending order of similarity based on the calculated similarity between each index in the text search library and the new text content, and the top M book information items are obtained. Next, a new text search result can be generated based on the newly searched M book information items and the original text search results. For example, assuming the original text search results contain N book information items, these N book information items can also be sorted in descending order of similarity, and the bottom M book information items in the original text search results can be removed. The remaining (NM) book information items are then merged with the newly searched M book information items to obtain N book information items as the new text search results, where N and M are both positive integers, and N is greater than M. Next, based on preset similarity adjustment rules, the similarity between the new text search results and the vector search results can be adjusted. For example, book information with publisher information matching the target publisher can be found from the new text search results, and the similarity of these book information can be increased by a third preset value. Similarly, book information with publisher information matching the target publisher can be found from the vector search results, and the similarity of these book information can be increased by a fourth preset value. And / or, the new text search results and the vector search results can be traversed. When one or more book information items in the new text search results (referred to as third candidate book information for ease of description) match one or more book information items in the vector search results (referred to as fourth candidate book information for ease of description), the similarity of the third candidate book information items is increased by a fifth preset value, and the similarity of the fourth candidate book information items is increased by a sixth preset value. In other words, if the same book information exists in both the new text search results and the vector search results, the similarity of that book information in the new text search results is increased by a fifth preset value, and the similarity of that book information in the vector search results is increased by a sixth preset value. Afterwards, based on the adjusted similarity of each book information item, the book information item with the highest similarity between the new text search results and the vector search results can be determined as the target book information corresponding to the book cover image.

[0124] Due to the large variety and quantity of books, it is impossible to accurately determine which book a cover belongs to by classification when identifying books. Text search is inaccurate when the image quality is poor, making text search results unreliable. Similarly, vector search results are also unreliable. However, there is a high probability that there are accurate results among the obtained text search results and vector search results. Therefore, this disclosure provides a variety of judgment strategies to find more accurate target book information from text search results and vector search results.

[0125] Furthermore, although there are many types of books, the number of publishers is limited. Therefore, the solution provided in this disclosure takes into account the limited number of publishers and uses a publisher classification model to obtain the target publisher contained in the book cover image. Since the publisher classification model has a high accuracy in practical applications and strong anti-interference ability against images, even when the image quality is poor, the classification result still has a very high accuracy rate. Therefore, using the publisher classification model to identify publishers can obtain reliable publisher identification results. This solution introduces reliable information through the publisher classification model, which is not only used for feature vector extraction to guide vector search, but also used to replace the publisher information in the text content of text recognition and re-perform text search, thereby achieving the purpose of guiding text search. This can correct the problem that text search cannot obtain accurate book information due to publisher identification errors, making the identified target book information more reliable.

[0126] In this embodiment, judgment strategies one through four are all precise judgments. Under normal circumstances, even without incorporating the publisher classification model results, they still have a high probability of error. However, this disclosure proposes judgment strategy five, which guides the reordering of text and vector searches by incorporating publisher classification results. This allows book cover images that cannot be identified based on the first four judgment strategies to obtain correct identification results with a high probability based on judgment strategy five provided by this disclosure. Typically, without incorporating publisher classification results, simply combining text search results and vector search results is unlikely to yield accurate results. This is because the book information with the highest similarity in both text search results and vector search results for those book cover images that have entered judgment strategy five is no longer reliable. The correct book information is likely hidden in the remaining book information in the text search results and vector search results (excluding the book information with the highest similarity). This disclosure's solution guides the reordering of the remaining book information in text search results and vector search results using reliable publisher classification results, thereby obtaining correct identification results. In the fifth judgment strategy disclosed herein, by increasing the similarity of the remaining book information with the publisher information that is consistent with the publisher classification results, the book information that is ranked lower will have the opportunity to be ranked higher again, and the book information that is ranked higher but is inconsistent with the publisher classification results will be moved to the lower position. In this way, after being guided by the publisher classification results, the text search results and vector search results are combined to further move the correct results forward, and finally obtain the correct recognition results, thereby improving the accuracy of the returned target book information.

[0127] In one optional embodiment of this disclosure, if the fourth judgment strategy is not satisfied, and the first candidate book information is not found in the vector search results, or the new similarity of the first candidate book information is not greater than the second threshold, then the text search results and vector search results can be merged, and the book information with the highest similarity can be selected as the target book information corresponding to the book cover image. When merging the text search results and vector search results, for book information coexisting in both text search results and vector search results, the higher similarity can be used as the similarity of the merged book information, or the average of the two similarities can be used as the similarity of the merged book information. Alternatively, all book information coexisting in both text search results and vector search results can be found, and then the book information containing the target publisher and with a high similarity can be selected as the target book information corresponding to the book cover image.

[0128] This exemplary embodiment also provides a book cover recognition device.

[0129] Figure 4 A schematic block diagram of a book cover recognition device according to an exemplary embodiment of the present disclosure is shown, such as Figure 4 As shown, the book cover recognition device 40 includes: a first recognition module 410, a second recognition module 420, a feature extraction module 430, a search module 440, and a determination module 450.

[0130] The first recognition module 410 is used to perform text recognition on the book cover image to be recognized, and obtain the text content contained in the book cover image.

[0131] The second recognition module 420 is used to input the book cover image into a pre-trained publisher classification model to obtain the target publisher contained in the book cover image;

[0132] Feature extraction module 430 is used to extract features based on the text content, the target publisher and the book cover image using a pre-trained feature extraction model to obtain the feature vector corresponding to the book cover image;

[0133] Search module 440 is used to perform text search based on the text content to obtain text search results, and to perform vector search based on the feature vector to obtain vector search results, wherein the text search results and the vector search results each contain at least one book information;

[0134] The determining module 450 is used to determine the target book information corresponding to the book cover image based on the text search results and the vector search results.

[0135] Optionally, the feature extraction model includes a first output branch and a second output branch. The first output branch is used to output a feature vector, and the second output branch is used to output the publisher identification result. When training the feature extraction model, iterative training is performed based on the feature vector output by the first output branch and the publisher identification result output by the second output branch.

[0136] Optionally, the feature extraction module 430 is further configured to:

[0137] The word vectors corresponding to the text content are converted with the target publisher to obtain a text data matrix, wherein the number of columns in the text data matrix is ​​the same as the number of columns in the image data matrix corresponding to the book cover image;

[0138] The image data matrix and the text data matrix are concatenated in the row direction to generate book cover data;

[0139] The book cover data is input into a pre-trained feature extraction model, and the feature vector output by the first output branch of the feature extraction model is obtained as the feature vector corresponding to the book cover image.

[0140] Optionally, the determining module 450 is further configured to:

[0141] The book information with the highest similarity to the text content in the text search results is selected as the first candidate book information;

[0142] If the first candidate book information meets the first preset condition, the first candidate book information is determined as the target book information corresponding to the book cover image;

[0143] The first preset condition includes:

[0144] The publisher information corresponding to the first candidate book information is consistent with the target publisher;

[0145] The text data corresponding to the first candidate book information is consistent with the text content; and

[0146] The confidence level corresponding to the text content is greater than the preset confidence threshold.

[0147] Optionally, the determining module 450 is further configured to:

[0148] If the first candidate book information does not meet the first preset condition, the book information with the highest similarity to the feature vector in the vector search results is obtained as the second candidate book information.

[0149] If the second candidate book information meets the second preset condition, the second candidate book information is determined as the target book information corresponding to the book cover image;

[0150] The second preset condition includes:

[0151] The information of the second candidate book is consistent with the information of the first candidate book; and

[0152] The publisher information corresponding to the second candidate book information is consistent with the target publisher.

[0153] Optionally, the determining module 450 is further configured to:

[0154] If the second candidate book information does not meet the second preset condition, the text search results are traversed to determine whether the second candidate book information exists in the text search results;

[0155] If the second candidate book information exists in the text search results, the similarity of the second candidate book information is increased by a first preset value to obtain a new similarity of the second candidate book information;

[0156] If the new similarity of the second candidate book information is greater than the first threshold, and the publisher information corresponding to the second candidate book information is consistent with the target publisher, the second candidate book information is determined as the target book information corresponding to the book cover image.

[0157] Optionally, the determining module 450 is further configured to:

[0158] If the second candidate book information is not found in the text search results, or the new similarity of the second candidate book information is not greater than the first threshold, or the publisher information corresponding to the second candidate book information is inconsistent with the target publisher, the vector search results are traversed to determine whether the first candidate book information exists in the vector search results.

[0159] If the first candidate book information exists in the vector search results, the similarity of the first candidate book information is increased by a second preset value to obtain a new similarity of the first candidate book information;

[0160] If the new similarity of the first candidate book information is greater than the second threshold, and the publisher information corresponding to the first candidate book information is consistent with the target publisher, the first candidate book information is determined as the target book information corresponding to the book cover image.

[0161] Optionally, the determining module 450 is further configured to:

[0162] If the publisher information corresponding to the first candidate book information is inconsistent with the target publisher, the publisher in the text content is replaced with the target publisher to obtain new text content;

[0163] Based on the new text content, a new text search is performed, and the search results are sorted in descending order of similarity to the new text content to obtain the top M book information, where M is a positive integer;

[0164] Based on the M book information and the text search results, generate new text search results;

[0165] Based on preset similarity adjustment rules, the similarity between the new text search results and the vector search results is adjusted;

[0166] Based on the adjusted similarity, the book information with the highest similarity between the new text search results and the vector search results is determined as the target book information corresponding to the book cover image.

[0167] Optionally, the determining module 450 is further configured to:

[0168] The similarity between the publisher information of books in the new text search results and the target publisher is increased by a third preset value, and the similarity between the publisher information of books in the vector search results and the target publisher is increased by a fourth preset value.

[0169] And / or,

[0170] Traverse the new text search results and the vector search results. If the third candidate book information in the new text search results is consistent with the fourth candidate book information in the vector search results, increase the similarity corresponding to the third candidate book information by a fifth preset value and increase the similarity corresponding to the fourth candidate book information by a sixth preset value.

[0171] The book cover recognition device provided in this disclosure can execute any book cover recognition method applicable to electronic devices provided in this disclosure, and has the corresponding functional modules and beneficial effects of executing the method. Content not described in detail in the device embodiments of this disclosure can be referred to the descriptions in any method embodiments of this disclosure.

[0172] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program, when executed by the at least one processor, causing the electronic device to perform a book cover recognition method according to embodiments of this disclosure.

[0173] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a book cover recognition method according to embodiments of this disclosure.

[0174] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a book cover recognition method according to embodiments of this disclosure.

[0175] refer to Figure 5 The present invention describes a structural block diagram of an electronic device 1100 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0176] like Figure 5 As shown, the electronic device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of the device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0177] Multiple components in electronic device 1100 are connected to I / O interface 1105, including: input unit 1106, output unit 1107, storage unit 1108, and communication unit 1109. Input unit 1106 can be any type of device capable of inputting information to electronic device 1100. Input unit 1106 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 1107 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1108 may include, but is not limited to, disk and optical disk. Communication unit 1109 allows electronic device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0178] The computing unit 1101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above. For example, in some embodiments, the book cover recognition method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1100 via ROM 1102 and / or communication unit 1109. In some embodiments, the computing unit 1101 can be configured to perform the book cover recognition method by any other suitable means (e.g., by means of firmware).

[0179] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0180] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0181] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0182] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0183] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0184] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

Claims

1. A method for identifying book covers, wherein, The method includes: Text recognition is performed on the book cover image to be identified to obtain the text content contained in the book cover image; The book cover image is input into a pre-trained publisher classification model to obtain the target publisher contained in the book cover image; The word vectors corresponding to the text content are converted with the target publisher to obtain a text data matrix, wherein the number of columns in the text data matrix is ​​the same as the number of columns in the image data matrix corresponding to the book cover image; The image data matrix and the text data matrix are concatenated in the row direction to generate book cover data; The book cover data is input into a pre-trained feature extraction model, and the feature vector output by the first output branch of the feature extraction model is obtained as the feature vector corresponding to the book cover image. A text search is performed based on the text content to obtain text search results, and a vector search is performed based on the feature vector to obtain vector search results. The text search results and the vector search results each contain at least one book information. Based on the text search results and the vector search results, the target book information corresponding to the book cover image is determined.

2. The book cover recognition method as described in claim 1, wherein, The feature extraction model includes a first output branch and a second output branch. The first output branch is used to output a feature vector, and the second output branch is used to output the publisher identification result. When training the feature extraction model, iterative training is performed based on the feature vector output by the first output branch and the publisher identification result output by the second output branch.

3. The book cover identification method as described in any one of claims 1-2, wherein, The step of determining the target book information corresponding to the book cover image based on the text search results and the vector search results includes: The book information with the highest similarity to the text content in the text search results is selected as the first candidate book information; If the first candidate book information meets the first preset condition, the first candidate book information is determined as the target book information corresponding to the book cover image; The first preset condition includes: The publisher information corresponding to the first candidate book information is consistent with the target publisher; The text data corresponding to the first candidate book information is consistent with the text content; and The confidence level corresponding to the text content is greater than the preset confidence threshold.

4. The book cover recognition method as described in claim 3, wherein, The step of determining the target book information corresponding to the book cover image based on the text search results and the vector search results further includes: If the first candidate book information does not meet the first preset condition, the book information with the highest similarity to the feature vector in the vector search results is obtained as the second candidate book information. If the second candidate book information meets the second preset condition, the second candidate book information is determined as the target book information corresponding to the book cover image; The second preset condition includes: The information of the second candidate book is consistent with the information of the first candidate book; and The publisher information corresponding to the second candidate book information is consistent with the target publisher.

5. The book cover recognition method as described in claim 4, wherein, The step of determining the target book information corresponding to the book cover image based on the text search results and the vector search results further includes: If the second candidate book information does not meet the second preset condition, the text search results are traversed to determine whether the second candidate book information exists in the text search results; If the second candidate book information exists in the text search results, the similarity of the second candidate book information is increased by a first preset value to obtain a new similarity of the second candidate book information; If the new similarity of the second candidate book information is greater than the first threshold, and the publisher information corresponding to the second candidate book information is consistent with the target publisher, the second candidate book information is determined as the target book information corresponding to the book cover image.

6. The book cover recognition method as described in claim 5, wherein, The step of determining the target book information corresponding to the book cover image based on the text search results and the vector search results further includes: If the second candidate book information is not found in the text search results, or the new similarity of the second candidate book information is not greater than the first threshold, or the publisher information corresponding to the second candidate book information is inconsistent with the target publisher, the vector search results are traversed to determine whether the first candidate book information exists in the vector search results. If the first candidate book information exists in the vector search results, the similarity of the first candidate book information is increased by a second preset value to obtain a new similarity of the first candidate book information; If the new similarity of the first candidate book information is greater than the second threshold, and the publisher information corresponding to the first candidate book information is consistent with the target publisher, the first candidate book information is determined as the target book information corresponding to the book cover image.

7. The book cover recognition method as described in claim 6, wherein, The step of determining the target book information corresponding to the book cover image based on the text search results and the vector search results further includes: If the publisher information corresponding to the first candidate book information is inconsistent with the target publisher, the publisher in the text content is replaced with the target publisher to obtain new text content; Based on the new text content, a new text search is performed, and the search results are sorted in descending order of similarity to the new text content to obtain the top M book information, where M is a positive integer; Based on the M book information and the text search results, generate new text search results; Based on preset similarity adjustment rules, the similarity between the new text search results and the vector search results is adjusted; Based on the adjusted similarity, the book information with the highest similarity between the new text search results and the vector search results is determined as the target book information corresponding to the book cover image.

8. The book cover recognition method as described in claim 7, wherein, The adjustment of the similarity between the new text search result and the vector search result based on the preset similarity adjustment rules includes: The similarity between the publisher information of books in the new text search results and the target publisher is increased by a third preset value, and the similarity between the publisher information of books in the vector search results and the target publisher is increased by a fourth preset value. And / or, Traverse the new text search results and the vector search results. If the third candidate book information in the new text search results is consistent with the fourth candidate book information in the vector search results, increase the similarity corresponding to the third candidate book information by a fifth preset value and increase the similarity corresponding to the fourth candidate book information by a sixth preset value.

9. A book cover recognition device, wherein, The device includes: The first recognition module is used to perform text recognition on the book cover image to be recognized, and obtain the text content contained in the book cover image; The second recognition module is used to input the book cover image into a pre-trained publisher classification model to obtain the target publisher contained in the book cover image; The feature extraction module is used to convert the word vectors corresponding to the text content with the target publisher to obtain a text data matrix, wherein the number of columns of the text data matrix is ​​the same as the number of columns of the image data matrix corresponding to the book cover image; the image data matrix and the text data matrix are concatenated in the row direction to generate book cover data; the book cover data is input into a pre-trained feature extraction model, and the feature vector output by the first output branch of the feature extraction model is obtained as the feature vector corresponding to the book cover image; The search module is used to perform a text search based on the text content to obtain a text search result, and to perform a vector search based on the feature vector to obtain a vector search result, wherein the text search result and the vector search result each contain at least one book information. The determination module is used to determine the target book information corresponding to the book cover image based on the text search results and the vector search results.

10. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the book cover recognition method according to any one of claims 1-8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the book cover recognition method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Intelligent reading recommendation method and device and electronic equipment

    CN107679070A

  • Book searching method, device and system

    CN111460185A