Text and image feature-based dual-stage book matching method and system
By combining the two-stage matching method of text and image features, OCR, improved TextCNN and ColBERT and self-built BookNet networks, the problems of low recognition accuracy and efficiency in book management are solved, and efficient and accurate book matching is achieved.
Patent Information
- Application Number
- CN202510623682.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-22
AI Technical Summary
The prior art has problems of low recognition accuracy and low efficiency in library management, especially in complex scenarios, which are difficult to meet the needs of efficient real-time and automated identification.
The two-stage matching method based on text and image features is adopted, and the book cover text information is extracted through OCR technology, and the improved TextCNN is used for preliminary classification, combined with ColBERT to calculate the text similarity, and image matching is performed through BookNet convolutional network model, and the book matching result is finally determined by the weighted average method.
It significantly improves the accuracy and efficiency of book retrieval and can achieve efficient and accurate book matching in complex scenarios.
Smart Images

Figure CN120526440A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a two-stage book matching method and system based on text and image features. Background Art
[0002] With the rapid development of the information society, the demand for book management in libraries, bookstores, personal collections, and other places is increasing. Traditional book management methods, such as manual search, barcode scanning, or optical character recognition (OCR), while convenient to a certain extent, suffer from significant shortcomings in efficiency, accuracy, and adaptability. This is especially true when faced with large numbers of books or complex environments, and they still struggle to meet the requirements for real-time, automated, and efficient recognition. Therefore, developing an efficient and accurate book matching technology and system is of great application value.
[0003] In existing literature, Chen DM et al. proposed a smartphone-based mobile book recognition system in their paper "Building Book Inventories using Smartphones." This system uses a smartphone to capture images and detect long straight lines using the Hough transform. Combined with the Segmentation-Based Speeded-Up Robust Features (SURF) algorithm, the segmented book spines are matched against a spine feature database for recognition. This method overcomes the traditional strict requirements for book placement order and outperforms traditional optical character recognition (OCR) algorithms to a certain extent. However, the recognition accuracy of this method is low, at only 74.5%.
[0004] In another related study, Cui Chen, in his paper "Research and Implementation of Key Technologies for Image-Based Book Spine Detection and Recognition," employed SIFT features, the approximate nearest neighbor algorithm (ANN), and the random sampling consensus (RANSAC) algorithm to design a multi-target template matching method for matching book spines in bookcase scene images. This method can achieve relatively accurate matching of book spines in dense vertical arrangements, tilted displays, and under varying lighting conditions. However, when faced with a large number of books, this method is inefficient and cannot meet the requirements for efficient recognition.
[0005] In summary, the existing technology still has certain deficiencies in terms of accuracy and efficiency of book recognition. Therefore, it is particularly important to propose a two-stage book matching method and system based on text and image features, which has great application prospects. Summary of the Invention
[0006] In response to the existing problems of low book matching accuracy and efficiency, this paper provides a two-stage book matching method and system based on text and image features. By combining text information with image features, this method employs a two-stage matching strategy to address issues of recognition accuracy and matching efficiency in complex scenarios.
[0007] In view of this, the present invention proposes a two-stage book matching method based on text and image features, comprising:
[0008] S1: Using an OCR method to extract text information from the segmented book cover image, wherein the text information is the title field information of the book.
[0009] S2: Based on the text information, the improved TextCNN is used to perform preliminary classification of the books;
[0010] S3: Using a text similarity analysis algorithm, calculate the text similarity between the book to be matched and all books in the preliminary classification;
[0011] S4: Perform image matching on candidate books with high text similarity, use the BookNet convolutional network model to extract book cover image features, and complete image matching using the cosine similarity algorithm;
[0012] S5: Comprehensive text similarity and image matching to determine the final matching books.
[0013] Specifically, in S1, the OCR method uses the PaddleOCR model to extract text information from the segmented book cover image, including:
[0014] S11: Send the segmented single book image to the PaddleOCR model to extract the text of each information
[0015] information;
[0016] S12: Divide the text into a title field and an author field based on the length information and the area occupied by the extracted text information;
[0017] S13: Output the information of the book title field.
[0018] Specifically, in S2, the improved TextCNN includes:
[0019] Add parallel LSTM branches to the TextCNN model to capture long-term sequence relationships of sentences;
[0020] Add a feature adaptive gating mechanism to the TextCNN model to dynamically adjust the weights of CNN and LSTM features;
[0021] The improved TextCNN model adopts adversarial training to improve the robustness of the model by adding perturbations.
[0022] Specifically, in S2, the improved TextCNN is used to perform preliminary classification of books based on the text information, including:
[0023] S21: Using the Chinese 24-category book classification system as the classification standard, each book is classified into one of the 24 categories, so that the improved TextCNN can classify the books;
[0024] S22: The text information extracted by S1 is used as the input of the improved TextCNN. The text information is processed by the CNN branch and LSTM branch in the improved TextCNN model to extract local features and long sequence features.
[0025] S23: Input the extracted local features and long sequence features into the adaptive gating mechanism to obtain feature information after feature fusion;
[0026] S24: Classify the fused feature information according to the trained weight file and divide each book into one of the 24 categories;
[0027] S25: The adversarial training network is used for training, and the perturbation value is 0.5.
[0028] Specifically, the text similarity analysis algorithm uses the ColBERT interactive model to calculate the text similarity between the book to be matched and all books in the preliminary classification, including:
[0029] S31: Encode each book title and obtain the vector representation of each word;
[0030] S32: Calculate the dot product similarity of all possible word vector combinations;
[0031] S33: Finally, the similarity score is output, and similar books are sorted from high to low according to the similarity score.
[0032] Specifically, the BookNet convolutional network model consists of a wavelet transform module, a feature extraction module, and an inverse wavelet transform module;
[0033] The candidate books with high text similarity are matched with images, the book cover image features are extracted using the BookNet convolutional network model, and the image matching is completed using the cosine similarity algorithm, including:
[0034] S41: The segmented single book image is input into the BookNet network and decomposed into four images of different frequency bands by the wavelet transform module;
[0035] S42: The images of the four different frequency bands are subjected to feature extraction by the four-branch structure of the feature extraction module to obtain semantic information of the different frequency bands and obtain feature maps of the four branches;
[0036] S43: The feature maps of the four branches are passed through the wavelet inverse transform module to output the fused feature map;
[0037] S44: Using the cosine similarity algorithm, the extracted feature information is subjected to image similarity analysis with the similar books selected in S3;
[0038] S45: Output books with high similarity for subsequent decision making.
[0039] Specifically, the matching decision adopted by S5 is: using the weighted average method, the formula is: S final =w text ×S text +w image ×S image , where w text is 0.25, w image is 0.75;
[0040] S final The scores are sorted from large to small, and books with high similarity are output according to the scores.
[0041] In a second aspect, the present invention provides a two-stage book matching system based on text and image features, comprising:
[0042] OCR module: used to extract text information from the segmented book cover image;
[0043] Classification module: Based on the extracted text information, the improved TextCNN is used to perform preliminary classification of books;
[0044] Text similarity calculation module: used to calculate the text similarity between the book to be matched and all books in the category;
[0045] Image matching module: performs image matching on candidate books with high text similarity, extracts cover image features and compares them;
[0046] Decision module: used to combine the text similarity analysis results with the image matching results to determine the final matching books.
[0047] This paper proposes a two-stage book matching method and system based on text and image features. By integrating text and image features and employing a two-stage matching strategy, this method significantly improves the accuracy and efficiency of book retrieval. Application of this method effectively addresses the real-life issues of low library book retrieval efficiency and insufficient matching accuracy. This method boasts high matching accuracy and efficiency and is suitable for use in libraries, book management, and other fields.
[0048] The two-stage matching framework proposed in this paper integrates text features and image features. From the theoretical and design perspectives, this method has obvious advantages: 1. The two-stage cascade design can effectively reduce the system's computational burden and improve retrieval efficiency; 2. The complementary fusion of text and image features can meet the recognition needs of various complex scenarios; 3. The modular design of the framework makes it scalable and adaptable.
[0049] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0052] Figure 1 This is a two-stage book matching method based on text and image features and the overall system framework of the present invention;
[0053] Figure 2 Schematic diagram of the improved TextCNN network model of the present invention;
[0054] Figure 3 Schematic diagram of the BookNet network model of the present invention;
[0055] Figure 4 This is a visualization diagram of the results of running the method of the present invention. DETAILED DESCRIPTION
[0056] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of systems consistent with certain aspects of the present invention, as detailed in the appended claims.
[0057] First, this implementation provides a two-stage book matching method based on text and image features. The overall framework flow chart is as follows: Figure 1 Specifically, the method includes the following steps:
[0058] Step 1: Use OCR technology to extract text information from the segmented book cover image:
[0059] 1) The OCR technology uses the PaddleOCR model. First, the segmented single book image is sent to the PaddleOCR model to extract the text information of each book;
[0060] 2) Based on the extracted text information, the text is divided into the title field and the author field according to the length information and the area occupied;
[0061] 3) Output the information of the title field.
[0062] Step 2: Based on the text information extracted by S1, the improved TextCNN is used to perform preliminary classification of books. The improved TextCNN network is as follows: Figure 2 As shown: The improved TextCNN model specifically includes:
[0063] Add parallel LSTM branches to capture the long sequence relationships of sentences, making up for the shortcomings of TextCNN in processing long sequence relationships;
[0064] Added a feature adaptive gating mechanism to dynamically adjust the weights of CNN and LSTM features;
[0065] Adversarial training is used to improve the robustness of the model by adding perturbations.
[0066] Use the improved TextCNN to perform preliminary classification of books, including:
[0067] 1) The classification standard adopts the Chinese 24-category book classification method, classifying each book into one of the 24 categories for TextCNN to classify;
[0068] 2) In the improved TextCNN, the book title field information output in step 1 is used as input;
[0069] 3) The title field information is extracted through the CNN branch and the LSTM branch in parallel to realize local feature and long sequence feature extraction;
[0070] 4) Input the extracted local features and long sequence features into the adaptive gating mechanism to obtain feature information after feature fusion;
[0071] 5) Classify the fused feature information according to the trained weight file and divide each book into one of the 24 categories;
[0072] 6) The network is trained in adversarial training mode, and a perturbation value of 0.5 is added to improve the robustness of the model.
[0073] Step 3: Use the text similarity analysis algorithm to calculate the text similarity between the book to be matched and all books in the category. The text similarity analysis algorithm uses the ColBERT interactive model:
[0074] 1) Encode each book title and obtain the vector representation of each word
[0075] 2) Calculate the dot product similarity of all possible word vector combinations
[0076] 3) Finally, the similarity score is output and similar books are sorted from high to low according to the similarity score.
[0077] Step 4: For candidate books with high text similarity, image matching is performed, and the self-built BookNet convolutional network model is used to extract book cover image features and complete the matching through the cosine similarity algorithm. The self-built BookNet convolutional network model is as follows: Figure 3 As shown:
[0078] 1) The self-built BookNet convolutional network model consists of a wavelet transform module, a feature extraction module, and an inverse wavelet transform module;
[0079] 2) The segmented single book image is input into the BookNet network and decomposed into four images of different frequency bands by the wavelet transform module;
[0080] 3) The images of four different frequency bands are subjected to feature extraction by the four-branch structure of the feature extraction module to obtain the semantic information of their different frequency bands and obtain the feature maps of their four branches;
[0081] 4) The feature maps of the four branches are passed through the inverse wavelet transform module to output the fused feature map;
[0082] 5) Use the cosine similarity algorithm to analyze the image similarity of the extracted feature information with the similar books selected in step 3
[0083] 6) Output books with high similarity for subsequent decision making.
[0084] Step 5: Combine the text similarity analysis and image matching results to determine the final matching books. The results are visualized as shown in the figure below. Figure 4 As shown:
[0085] 1) The matching decision adopts the weighted average method, the formula is: S final =w text ×S text +w image ×S image , where w text is 0.25, w image is 0.75;
[0086] 2) S final The scores are sorted from large to small, and books with high similarity are output according to the scores.
[0087] This paper proposes a two-stage book matching method and system based on text and image features that effectively addresses the problems of low library book retrieval efficiency and insufficient matching accuracy. By integrating text and image features and employing a two-stage matching strategy, this method significantly improves the accuracy and efficiency of book retrieval.
[0088] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These changes and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A two-stage book matching method based on text and image features, characterized by: include: S1: Using an OCR method to extract text information from the segmented book cover image, wherein the text information is the title field information of the book. S2: Based on the text information, the improved TextCNN is used to perform preliminary classification of the books; S3: Using a text similarity analysis algorithm, calculate the text similarity between the book to be matched and all books in the preliminary classification; S4: Perform image matching on candidate books with high text similarity, use the BookNet convolutional network model to extract book cover image features, and complete image matching using the cosine similarity algorithm; S5: Comprehensive text similarity and image matching to determine the final matching books.
2. The two-stage book matching method based on text and image features according to claim 1 is characterized in that: In S1, the OCR method uses the PaddleOCR model to extract text information from the segmented book cover image, including: S11: Send the segmented single book image to the PaddleOCR model to extract the text information of each book; S12: Divide the text into a title field and an author field based on the length information and the area occupied by the extracted text information; S13: Output the information of the book title field.
3. The two-stage book matching method based on text and image features according to claim 1 is characterized in that: In S2, the improved TextCNN includes: Add parallel LSTM branches to the TextCNN model to capture long-term sequence relationships of sentences; Add a feature adaptive gating mechanism to the TextCNN model to dynamically adjust the weights of CNN and LSTM features; The improved TextCNN model adopts adversarial training to improve the robustness of the model by adding perturbations.
4. The two-stage book matching method based on text and image features according to claim 1 is characterized in that: In S2, the improved TextCNN is used to perform preliminary classification of books based on the text information, including: S21: Using the Chinese 24-category book classification system as the classification standard, each book is classified into one of the 24 categories, so that the improved TextCNN can classify the books; S22: The text information extracted by S1 is used as the input of the improved TextCNN. The text information is processed by the CNN branch and LSTM branch in the improved TextCNN model to extract local features and long sequence features. S23: Input the extracted local features and long sequence features into the adaptive gating mechanism to obtain feature information after feature fusion; S24: Classify the fused feature information according to the trained weight file and divide each book into one of the 24 categories; S25: The network training method of adversarial training is adopted, and the perturbation value added is 0.
5.
5. The two-stage book matching method based on text and image features according to claim 1 is characterized in that: The text similarity analysis algorithm uses the ColBERT interactive model to calculate the text similarity between the book to be matched and all books in the preliminary classification, including: S31: Encode each book title and obtain the vector representation of each word; S32: Calculate the dot product similarity of all possible word vector combinations; S33: Finally, the similarity score is output, and similar books are sorted from high to low according to the similarity score.
6. The two-stage book matching method based on text and image features according to claim 1 is characterized in that: The BookNet convolutional network model consists of a wavelet transform module, a feature extraction module, and an inverse wavelet transform module; The candidate books with high text similarity are matched with images, the book cover image features are extracted using the BookNet convolutional network model, and the image matching is completed using the cosine similarity algorithm, including: S41: The segmented single book image is input into the BookNet network and decomposed into four images of different frequency bands by the wavelet transform module; S42: The images of the four different frequency bands are subjected to feature extraction by the four-branch structure of the feature extraction module to obtain semantic information of the different frequency bands and obtain feature maps of the four branches; S43: The feature maps of the four branches are passed through the wavelet inverse transform module to output the fused feature map; S44: Using the cosine similarity algorithm, the extracted feature information is subjected to image similarity analysis with the similar books selected in S3; S45: Output books with high similarity for subsequent decision making.
7. The two-stage book matching method based on text and image features according to claim 1 is characterized in that: S5 uses the weighted average method to make the matching decision. The formula is: final =w test ×S text +w image ×S image , where w text is 0.25, w image is 0.75; S final The scores are sorted from large to small, and books with high similarity are output according to the scores.
8. A two-stage book matching system based on text and image features, characterized by: include OCR module: used to extract text information from the segmented book cover image; Classification module: Based on the extracted text information, the improved TextCNN is used to perform preliminary classification of books; Text similarity calculation module: used to calculate the text similarity between the book to be matched and all books in the category; Image matching module: performs image matching on candidate books with high text similarity, extracts cover image features and compares them; Decision module: used to combine the text similarity analysis results with the image matching results to determine the final matching books.