Automatic book cataloguing method

By improving the YOLO model and clustering algorithm, the problem of inaccurate classification of book covers and short text titles is solved, and high accuracy management of automatic book cataloging is achieved.

CN120356197AActive Publication Date: 2025-07-22INNER MONGOLIA UNIV OF SCI & TECH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510222828.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-22
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the prior art, the complex pattern of the book cover background and decorative text lead to inaccurate OCR recognition, and the short length of the title text leads to inaccurate classification, especially the short text title cannot be accurately classified.

Method used

The improved YOLO text detection model is adopted, combined with the ASPP module and the improved clustering algorithm and neural network to achieve accurate positioning and classification of cataloged text areas.

Benefits of technology

It improves the accuracy of book classification cataloging, avoids interference from background patterns and decorative texts, ensures the accurate classification of titles and introduction texts, and achieves efficient management of books.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356197A_ABST
    Figure CN120356197A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of book cataloguing, in particular to an automatic book cataloguing method, which comprises the following steps of: generating a multi-row cataloguing text block diagram from a book cover image through improved YOLO; generating a name text and a description text from the catalogue text block diagram through an OCR (Optical Character Recognition) model; classifying the description text to generate a name text and a brief introduction text; if the text length of the name text is greater than the text length lower limit, performing clustering calculation on the name text and classified names of the book management library through a name classification model based on an improved clustering algorithm to generate book classification tags; if the text length of the name text is smaller than or equal to the text length lower limit, generating a book classification label from the brief introduction text through a text classification model based on an improved neural network; and inputting the text and the book classification label into a library automation system. According to the method, the interference of background patterns or decorative characters on OCR (Optical Character Recognition) is avoided, so that the classified cataloguing result of the book is more accurate, and accurate cataloguing management of the book is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of book cataloging, and particularly relates to a method for automatic book cataloging. Background Art

[0002] Book cataloging is the process of entering book information into a system, which is a crucial part of the library's automated bibliographic management system. The quality of the entered information directly determines the accuracy and integrity of the library's collection information. Therefore, most libraries currently use optical character recognition (OCR) and natural language processing (NLP) technologies to intelligently extract and classify the entered information to achieve the efficiency and accuracy of automatic book cataloging.

[0003] However, there are the following problems in the implementation of automatic book cataloging: The background pattern of the book cover is relatively complex and rich, and there are decorative texts, resulting in inaccurate OCR recognition of the title text, or the recognition of decorative texts as the title, causing errors in the entered information; The length of the title text of the book is too short, resulting in inaccurate classification. For example, for a book titled "Memory", NLP performs semantic classification by vectorizing the title text. Since the feature vector generated by less than six characters has too low a dimension, the NLP designed for large-dimensional vector data cannot accurately classify it.

[0004] For example, the patent document with the application number 202311551618.2 discloses a method, device and medium for automatic book cataloging for a cataloging robot. However, during its automatic cataloging process, it does not perform intelligent recognition and extraction of the cataloging text, nor can it classify the title of short texts.

[0005] Therefore, how to achieve intelligent recognition and extraction of cataloging text and classification of the title of short texts to achieve high accuracy in automatic book cataloging is a technical problem to be solved at present. Summary of the Invention

[0006] To this end, the present invention provides a method for automatic book cataloging. By improving the text detection model of YOLO, accurate positioning and recognition of the cataloging text area are achieved, avoiding the interference of background patterns or decorative texts on OCR, and also enabling the position information of the text area to be used for the classification and cataloging judgment of the text. By performing clustering calculations on longer title texts through a clustering algorithm, or by improving the neural network to perform semantic recognition on the introduction text, the classification and cataloging results of the book are made more accurate, achieving accurate cataloging management of the book.

[0007] To achieve the above object, the present invention proposes a method for automatic book cataloging, including:

[0008] Generate a multi-line catalog text block diagram from the book cover image through a text detection model based on improved YOLO, where the text detection model embeds an improved ASPP module;

[0009] Convert the catalog text block diagram through an OCR model to generate a name text and a description text;

[0010] Classify the description text into a title text and a brief introduction text according to the positional relationship between the name text and the description text, as well as the name text area of the name text and the title text area of the description text;

[0011] If the text length of the title text is greater than the lower limit of the text length, cluster the title text and the classified titles in the book management library through a title classification model based on an improved clustering algorithm to generate a book classification label;

[0012] If the text length of the title text is less than or equal to the lower limit of the text length, generate the book classification label from the brief introduction text through a text classification model based on an improved neural network;

[0013] Input the name text, the brief introduction text, the title text, and the book classification label into the library automation system.

[0014] Furthermore, the branch of the improved ASPP module embeds a coordinate attention mechanism, where ASPP is used to localize the blurred catalog text block diagram of the image features obtained at different sampling rates through the coordinate attention mechanism.

[0015] Furthermore, the improved ASPP module has a first branch to a fifth branch, a feature splicing layer, and a feature convolution output layer. The first branch is provided with a first-size convolutional layer and the coordinate attention mechanism. The second to fourth branches are respectively provided with a second-size convolutional layer and the coordinate attention mechanism. The fifth branch is provided with a pooling layer and a deconvolution layer. The convolutional kernel of the first-size convolutional layer is smaller than the convolutional kernel of the second-size convolutional layer. The outputs of the first to fifth branches are connected to the feature splicing layer, and the output of the feature splicing layer is connected to the feature convolution output layer;

[0016] Localize the blurred catalog text block diagram of the image features obtained at different sampling rates through the first to fourth branches;

[0017] Generate a localization splicing feature by splicing the features of the localization through the feature splicing layer;

[0018] Extract and output the features of the localization splicing feature through the feature convolution output layer.

[0019] Further, the text detection model includes a smooth label sub-model, an improved loss function, and a decoupled head for performing inclined bounding on the catalog text block diagram. The process of the decoupled head for outputting the catalog text block diagram includes:

[0020] The decoupled head classifies the rotation angle of the catalog information box through the smooth label sub-model to generate an angle label value;

[0021] The decoupled head learns the inclination angle of the true box through the improved loss function, where the improved loss function is constructed based on the angle label value and the cross-entropy function.

[0022] In the above solution, through the ASPP module embedded with the coordinate attention mechanism, the YOLO model pays more attention to the accurate position information of different sampling rates of the image. Through the ASPP module and the improved loss function, the YOLO model can detect the catalog text with an inclined design, generate an inclined block diagram, and thus avoid OCR from recognizing the inclined text as text in different lines.

[0023] Further, the clustering algorithm is the Kmeans clustering algorithm. The process of generating a book classification label by clustering the title text and the classified titles in the book management library through a title classification model based on the improved clustering algorithm includes:

[0024] Generating a clustering cluster including all text vectors from the title text and the classified titles through a word embedding model;

[0025] Dividing the clustering cluster into two sub-clustering clusters according to the average value of all text vectors in the clustering cluster, retaining the sub-clustering cluster that reduces the total error, and removing the sub-clustering cluster that cannot reduce the total error;

[0026] Performing binary partitioning of the retained sub-clustering cluster by repeatedly calculating the average value until the number of clustering clusters reaches a set value, and outputting the book classification label.

[0027] Further, if the text vector corresponding to the classified title is less than the product of the average value and the weight, the classified title is removed.

[0028] In the above solution, the Kmeans clustering algorithm with binary partitioning is used to perform clustering calculation on the long title text, overcoming the problem that the Kmeans clustering algorithm is prone to converge to the global minimum due to the lack of rich semantic information in the title text itself, and realizing accurate classification of books based on the semantic information of the title text.

[0029] Further, the process of generating a book classification label from the abstract text through a text classification model based on an improved neural network includes:

[0030] Extract the keywords of the abstract text;

[0031] Generate keyword vectors from the keywords through a word embedding model;

[0032] Generate the book classification labels from the keyword vectors based on a neural network module and a bidirectional long short-term memory network module.

[0033] Further, the process of generating book classification labels from keyword vectors based on a neural network module and a bidirectional long short-term memory network module includes:

[0034] Generate a first vector and a second vector from the keyword vectors through the neural network module and the bidirectional long short-term memory network module respectively;

[0035] Perform weighted summation on the first vector and the second vector to generate a comprehensive feature vector;

[0036] Generate the book classification labels from the comprehensive feature vector through a fully connected layer and an activation function in sequence.

[0037] In the above solution, the relevance between the words before and after the abstract text is extracted and learned through a neural network module and a bidirectional long short-term memory network module, achieving a better text classification effect for the abstract text.

[0038] Further, the process of classifying the description text into a title text and an abstract text according to the positional relationship between the name text and the description text, the name text area of the name text, and the title text area of the description text includes:

[0039] Divide the description text into multiple lines of line text, and calculate the height ratio between the multiple lines of line text;

[0040] If the height ratio is greater than the ratio threshold, determine the line text with the highest height as the title text, and the remaining line text as the abstract text;

[0041] If the height ratio is less than or equal to the ratio threshold, determine the title text and the abstract text according to the weighted calculation result of the height ratio, the center point distance between the name text and the description text.

[0042] Further, the catalog text block diagram includes a name text block diagram, a description text block diagram, a price block diagram, a publisher block diagram, a publication date block diagram, a collection location block diagram, and a call number block diagram;

[0043] The catalog text block diagram is also converted through an OCR model to generate other information text;

[0044] Extract keywords from the other information text to generate descriptive text, price, publisher, publication date, collection location, and call number;

[0045] Extract keywords from the name text to generate author name, translator name, and editor name.

[0046] In the above solution, accurate classification of the book cover text is achieved based on position information and size information.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows

[0048] 1. By improving the text detection model of YOLO, accurate positioning and recognition of the cataloging text area are achieved, avoiding interference of background patterns or decorative text on OCR, and enabling the position information of the text area to be used for classification and cataloging judgment of the text. Through the clustering algorithm, clustering calculation is performed on the long title text, or through improving the neural network, semantic recognition is performed on the brief introduction text, making the classification and cataloging results of the books more accurate and achieving accurate cataloging management of the books.

[0049] 2. Through the ASPP module embedded with the coordinate attention mechanism, the YOLO model pays more attention to the accurate position information of different sampling rates of the image. Through the ASPP module and the improved loss function, the YOLO model can detect the cataloging text with an inclined design, generate an inclined frame diagram, and thus avoid OCR from recognizing the inclined text as text in different lines.

[0050] 3. Through the Kmeans clustering algorithm of binary partitioning, clustering calculation is performed on the long title text, overcoming the problem that the Kmeans clustering algorithm is prone to converge to the global minimum due to the lack of rich semantic information in the title text itself, and achieving accurate classification of books based on the semantic information of the title text.

[0051] 4. It is achieved to extract and learn the correlation between the front and back words of the brief introduction text through the neural network module and the bidirectional long short-term memory network module, achieving a better text classification effect for the brief introduction text. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic flowchart of the book automatic cataloging method according to an embodiment of the present invention;

[0053] Figure 2 It is a schematic structural diagram of the improved YOLO text detection model of the book automatic cataloging method according to an embodiment of the present invention;

[0054] Figure 3 It is a schematic structural diagram of the improved ASPP module of the improved YOLO text detection model of the book automatic cataloging method according to an embodiment of the present invention;

[0055] Figure 4 This is a schematic structural diagram of a text classification model of an improved neural network for the book automatic cataloging method according to an embodiment of the present invention. Detailed implementation manners

[0056] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.

[0058] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0059] In addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0060] As Figures 1 to 4 shown, the present invention provides a book automatic cataloging method. By improving the text detection model of YOLO, the accurate positioning and recognition of the cataloging text area are realized, the interference of background patterns or decorative texts to OCR is avoided, and the position information of the text area can also be used for the classification and cataloging judgment of the text. By using a clustering algorithm to perform clustering calculations on longer title texts, or by improving the neural network to perform semantic recognition on the introduction texts, the classification and cataloging results of the books are made more accurate, and accurate cataloging management of the books is realized.

[0061] As Figures 1 to 4 shown, this embodiment proposes a book automatic cataloging method, including:

[0062] Generate a multi-line catalog text block diagram from the book cover image through a text detection model based on improved YOLO (You Only Look Once), where the text detection model embeds an improved ASPP (Atrous Spatial Pyramid Pooling) module for blurry catalog text block diagram localization;

[0063] Convert the catalog text block diagram through an OCR (Optical Character Recognition) model to generate a name text and a description text;

[0064] Classify the description text into a title text and an introduction text according to the positional relationship between the name text and the description text, as well as the name text area of the name text and the title text area of the description text;

[0065] If the text length of the title text is greater than the lower limit of the text length, then cluster the title text and the classified titles in the library management database through a title classification model based on an improved clustering algorithm to generate a book classification label;

[0066] If the text length of the title text is less than or equal to the lower limit of the text length, then generate the book classification label through a text classification model based on an improved neural network for the introduction text;

[0067] Input the name text, the introduction text, the title text, and the book classification label into the library automation system.

[0068] It can be understood that YOLO with the improved ASPP module realizes the processing of blurry and damaged book information, ensures the integrity of the catalog information, and can learn and output the text distribution characteristics of the book cover, so as to generate a title text and an introduction text based on the text distribution characteristics. The description text is the text longer than other information on the book cover, including the title text and the introduction text. The title text includes the book name and the subtitle. The introduction text is the content introduction of the book. Compared with other information on the book cover, such as price, author, translator, editor, publisher, publication date, collection location, call number, etc., it is difficult to classify and extract through feature words such as "written by" and "translated by". Therefore, this embodiment uses the block diagram generated by the YOLO model for auxiliary classification to improve the accuracy of its distinction.

[0069] Furthermore, as Figure 3 shown, the branch of the improved ASPP module embeds a coordinate attention mechanism, where ASPP is used to make the image features obtained at different sampling rates perform blurry catalog text block diagram localization through the coordinate attention mechanism.

[0070] It is understandable that YOLO is preferably the YOLOv5 model as shown in Figure 2 Although its performance for object detection is relatively high, the bounding boxes output by it are rectangular boxes that are parallel to each other without rotation angles, resulting in insufficient accuracy when outputting bounding boxes with inclined angles. Therefore, the ASPP module can accurately locate text block diagrams that are relatively blurred and have certain damage, and can also output bounding boxes with high-accuracy inclined angles.

[0071] It is understandable that the Coordinate Attention module processes the input feature map in terms of width, height, and rotation angle using global average pooling, and encodes each channel. This method can make the model pay more attention to the accurate position information of the image and can obtain the attention on the width, height, and rotation angle of the image. Therefore, the Coordinate Attention module can improve the accuracy of the model.

[0072] Furthermore, as shown in Figure 3 the improved ASPP module has a first branch to a fifth branch, a feature splicing layer, and a feature convolution output layer. The first branch is provided with a first-size convolutional layer and the coordinate attention mechanism. The second to fourth branches are respectively provided with a second-size convolutional layer and the coordinate attention mechanism. The fifth branch is provided with a pooling layer and a deconvolution layer. The convolutional kernel of the first-size convolutional layer is smaller than that of the second-size convolutional layer. The outputs of the first to fifth branches are connected to the feature splicing layer, and the output of the feature splicing layer is connected to the feature convolution output layer;

[0073] The first to fourth branches are used to locate the blurred cataloged text block diagrams for the image features obtained at different sampling rates; the feature splicing layer is used to splice the features of the positioning to generate positioning spliced features; the feature convolution output layer is used to perform feature extraction and output on the positioning spliced features.

[0074] Specifically, as shown in Figure 3 the convolutional kernel of the first-size convolutional layer is 1x1, the convolutional kernel of the second-size convolutional layer is 3x3, the dilation rates of the convolutional layers of the first to fourth branches are 1, 6, 12, and 18 in sequence, and the convolutional kernels of the deconvolution layer of the fifth branch and the convolutional kernel of the feature convolution output layer are 1x1.

[0075] Furthermore, as shown in Figure 2 the text detection model includes a smooth label sub-model, an improved loss function, and a decoupled head, which are used to perform inclined box selection on the cataloged text block diagrams. The process of the decoupled head for outputting the cataloged text block diagrams includes:

[0076] The decoupling head classifies the rotation angle of the catalog information box through the smooth label sub-model to generate an angle label value;

[0077] The decoupling head learns the tilt angle of the true box through an improved loss function, where the improved loss function is constructed based on the angle label value and the cross-entropy function.

[0078] Specifically, the smooth label sub-model is preferably a Circular Smooth Label (CSL), and then the tilt angle regression problem is converted into a classification problem for processing.

[0079] Specifically, the improved loss function is:

[0080]

[0081] In the formula, loss(z,y) represents the improved loss function, y i,j represents the true label value of the i-th sample at the j-th angle in the circular smooth label, δ represents the sigmoid activation function, and z i,j represents the predicted label value of the i-th sample at the j-th angle in the circular smooth label. The formula sums the logarithmic calculations of the true label values and predicted label values at each angle to generate the loss function.

[0082] Preferably, the YOLOv5 model has multiple decoupling heads, and the improved loss function is only used for one of the decoupling heads, and the other decoupling heads can output normal block diagrams.

[0083] In the above solution, through the ASPP module embedded with the coordinate attention mechanism, the YOLO model pays more attention to the accurate position information of different sampling rates of the image. Through the ASPP module and the improved loss function, the YOLO model can detect the catalog text with a tilt design, generate a tilted block diagram, and thus avoid OCR from recognizing the tilted text as text in different lines.

[0084] Further, the clustering algorithm is the Kmeans clustering algorithm. The process of generating a book classification label by clustering the title text and the classified titles in the book management library through a title classification model based on the improved clustering algorithm includes:

[0085] Generating a clustering cluster including all text vectors by the title text and the classified titles through a word embedding model;

[0086] Dividing the clustering cluster into two sub-clustering clusters according to the average value of all text vectors in the clustering cluster, retaining the sub-clustering cluster that reduces the total error, and removing the sub-clustering cluster that cannot reduce the total error;

[0087] The binary partition that cyclically calculates the average value of the reserved sub-cluster clusters is performed until the number of clustering clusters reaches the set value, and the book classification label is output.

[0088] Specifically, the word embedding model is preferably a Word2Vec model, which is used to map vocabulary to a high-dimensional vector space. The Kmeans clustering algorithm preferably sets the number of clustering clusters K to 8, namely eight categories: literature, science, history and geography, philosophy, economy and politics, medicine, education, and industrial technology. Then, 8 initial points are randomly selected as the clustering centers, and the title text and the classified titles are all assigned to the cluster to which the centroid closest to them belongs. Then, the centroid of each cluster is updated to the average value of all points in the cluster for binary selection, thereby overcoming the problem that the kmeans algorithm is prone to converge to the global minimum.

[0089] By comparing the errors of the two sub-cluster clusters after binary division with the total error before binary division, it is judged whether the sub-cluster clusters reduce the total error.

[0090] Further, if the text vector corresponding to the classified title is less than the product of the average value and the weight, the classified title is removed, and the specific calculation is through the following formula:

[0091] d(p i )<ηMed(P)

[0092] In the formula, d(p i ) is the text vector corresponding to the classified title i, η is the weight, preferably 1.7, and Med(P) is the average value. Among them, the text vector is the distance from the classified title i to the cluster center, and Med(P) is the average value of the distances from all classified titles i in the cluster to the cluster center. Therefore, the classified titles that are too far from the cluster center and are not conducive to classification can be removed, thereby making the Kmeans clustering algorithm have a better clustering and classification effect for the title text of short texts.

[0093] In the above solution, the Kmeans clustering algorithm with binary division is used to perform clustering calculations on the longer title text, overcoming the problem that the semantic information of the title text itself is not rich enough, resulting in the Kmeans clustering algorithm being prone to converge to the global minimum, and realizing accurate classification of books based on the semantic information of the title text.

[0094] Further, as Figure 4 shown, the process of generating the book classification label for the abstract text through the text classification model based on the improved neural network includes:

[0095] Extract the keywords of the abstract text;

[0096] Generate keyword vectors for the keywords through a word embedding model;

[0097] Generate the book classification labels based on the keyword vectors through a neural network module and a bidirectional long short-term memory network (BiLSTM) module.

[0098] Preferably, extract keywords with commonalities from the introduction text through TextRank, and the word embedding model is a Word2Vec model.

[0099] Furthermore, as Figure 4 shown, the process of generating book classification labels based on the keyword vectors through a neural network module and a bidirectional long short-term memory network module includes:

[0100] Pass the keyword vectors through the neural network module and the bidirectional long short-term memory network module respectively to generate a first vector and a second vector;

[0101] Perform weighted summation on the first vector and the second vector to generate a comprehensive feature vector;

[0102] Pass the comprehensive feature vector through a fully connected layer and an activation function in sequence to generate the book classification labels.

[0103] Preferably, the activation function is Softmax.

[0104] Specifically, since the introduction text may involve information of other labels, the neural network module may extract features of other labels, thereby affecting the classification result. According to the different degrees of dependence of different texts on local semantic features, a gating mechanism is introduced, aiming to assign weights to the text significant features captured by the neural network module and the text context features captured by the bidirectional long short-term memory network module, and then fuse these two features according to the weights to improve the accuracy of text classification. The calculation formula of the gating mechanism is:

[0105] a1 = δ(W·T cnn + b)

[0106] T text = a1T cnn + (1 - a1)T Bi

[0107] In the formula, a1 is the weight coefficient, δ represents the sigmoid activation function, W represents the weight matrix, b is the bias term, and T text 、T cnn 、T Bi are the comprehensive feature vector, the first vector, and the second vector respectively.

[0108] In the above solution, the neural network module and the bidirectional long short-term memory network module are used to extract and learn the correlation between the words before and after the introduction text, achieving a better text classification effect for the introduction text.

[0109] Further, the process of classifying the description text into a title text and an introduction text according to the positional relationship between the name text and the description text, the area of the name text of the name text, and the area of the title text of the description text includes:

[0110] Dividing the description text into multiple lines of line text and calculating the height ratio between the multiple lines of line text;

[0111] If the height ratio is greater than the ratio threshold, the line text with the highest height is determined as the title text, and the remaining line text is determined as the introduction text.

[0112] Specifically, the ratio threshold is preferably 0.5, which conforms to the conventional font size relationship between the title and the introduction.

[0113] If the height ratio is less than or equal to the ratio threshold, the title text and the introduction text are determined according to the weighted calculation result of the height ratio, the name text, and the central point distance of the description text. Specifically:

[0114]

[0115] In the formula, f represents the score result of the weighted calculation, 0.3 and 0.7 are the preferred weighting coefficients, p is the height ratio, and d1 and d2 are the central point distances between the name text and the first part and the second part of the description text respectively. When the score result is greater than 1.3, the first part is the title text and the second part is the introduction text, otherwise it is the opposite.

[0116] It can be understood that usually the author's name is closer to the title text and farther from the introduction text. Therefore, the title text and the introduction text can be distinguished through the calculation result of the above formula.

[0117] Further, the catalog text block diagram includes a name text block diagram, a description text block diagram, a price block diagram, a publisher block diagram, a publication date block diagram, a collection location block diagram, and a call number block diagram;

[0118] The catalog text block diagram is also converted through an OCR model to generate other information text;

[0119] Extracting keywords from the other information text to generate a description text, price, publisher, publication date, collection location, and call number;

[0120] Extracting keywords from the name text to generate the author's name, translator's name, and editor's name.

[0121] Specifically, classification can be achieved by matching the feature words of other information texts and name texts through TextRank, and the input is managed by the library automation system.

[0122] In the above solution, accurate classification of the book cover text according to the position information and size information is achieved.

[0123] In this embodiment, by improving the text detection model of YOLO, accurate positioning and recognition of the cataloging text area are achieved, avoiding the interference of background patterns or decorative texts on OCR, and also enabling the position information of the text area to be used for the classification judgment of the text. By using the clustering algorithm to perform clustering calculations on the longer title text, or by improving the neural network to perform semantic recognition on the introduction text, the classification result of the book is more accurate, and accurate classification management of the book is achieved. By embedding the coordinate attention mechanism into the ASPP module, the YOLO model pays more attention to the accurate position information of different sampling rates of the image. Through the ASPP module and the improved loss function, the YOLO model can detect the cataloging text with an inclined design and generate an inclined block diagram, thereby preventing OCR from recognizing the inclined text as text in different lines. By using the binary partition Kmeans clustering algorithm to perform clustering calculations on the longer title text, the problem that the Kmeans clustering algorithm is prone to converge to the global minimum due to the lack of rich semantic information in the title text itself is overcome, and accurate classification of the book based on the semantic information of the title text is achieved. By using the neural network module and the bidirectional long short-term memory network module to extract and learn the correlation between the front and back words of the introduction text, a better text classification effect for the introduction text is achieved.

[0124] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

[0125] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for automatic cataloging of books, characterized in that, Including: Generating a multi-line catalog text block diagram from the book cover image through a text detection model based on improved YOLO, where the text detection model embeds an improved ASPP module; Converting the catalog text block diagram through an OCR model to generate a name text and a description text; Classifying the description text into a title text and an introduction text according to the positional relationship between the name text and the description text, as well as the name text area of the name text and the title text area of the description text; If the text length of the title text is greater than the lower limit of the text length, clustering the title text and the classified titles in the book management library through a title classification model based on an improved clustering algorithm to generate a book classification label; If the text length of the title text is less than or equal to the lower limit of the text length, generating the book classification label from the introduction text through a text classification model based on an improved neural network; Inputting the name text, the introduction text, the title text, and the book classification label into the library automation system.

2. The automatic book cataloging method according to claim 1, characterized in that, The branch of the improved ASPP module embeds a coordinate attention mechanism, where ASPP is used to perform fuzzy catalog text block diagram positioning on the image features obtained at different sampling rates through the coordinate attention mechanism.

3. The automatic book cataloging method according to claim 2, wherein The improved ASPP module has a first branch to a fifth branch, a feature splicing layer, and a feature convolution output layer. The first branch is provided with a first-size convolutional layer and the coordinate attention mechanism. The second to fourth branches are respectively provided with a second-size convolutional layer and the coordinate attention mechanism. The fifth branch is provided with a pooling layer and a deconvolution layer. The convolutional kernel of the first-size convolutional layer is smaller than the convolutional kernel of the second-size convolutional layer. The outputs of the first to fifth branches are connected to the feature splicing layer, and the output of the feature splicing layer is connected to the feature convolution output layer; Performing fuzzy catalog text block diagram positioning on the image features obtained at different sampling rates through the first to fourth branches; Performing feature splicing on the positioning through the feature splicing layer to generate a positioning splicing feature; Performing feature extraction and output on the positioning splicing feature through the feature convolution output layer.

4. The book automatic cataloging method according to claim 2, characterized in that, The text detection model includes a smooth label sub-model, an improved loss function, and a decoupled head, which are used to perform inclined box selection on the catalog text block diagram. The process of the decoupled head for outputting the catalog text block diagram includes: The decoupled head classifies the rotation angle of the catalog information box through the smooth label sub-model to generate an angle label value; The decoupled head learns the inclination angle of the true box through the improved loss function, where the improved loss function is constructed based on the angle label value and the cross-entropy function.

5. The automatic book cataloging method according to claim 1, characterized in that The clustering algorithm is the Kmeans clustering algorithm. The process of clustering the title text and the classified titles in the book management library through a title classification model based on an improved clustering algorithm to generate a book classification label includes: Generating a clustering cluster including all text vectors from the title text and the classified titles through a word embedding model; Divide the cluster into two sub - clusters according to the average value of all text vectors in the cluster, retain the sub - cluster that reduces the total error, and remove the sub - cluster that cannot reduce the total error; Perform a binary partition by cyclically calculating the average value of the retained sub - clusters until the number of clustering clusters reaches a set value, and output the book classification label.

6. The automatic book cataloging method according to claim 5, wherein If the text vector corresponding to the classified title is less than the product of the average value and the weight, remove the classified title.

7. The automatic book cataloging method according to claim 1, wherein The process of generating a book classification label for the introduction text through a text classification model based on an improved neural network includes: Extract the keywords of the introduction text; Generate keyword vectors for the keywords through a word embedding model; Generate the book classification label based on the keyword vectors using a neural network module and a bidirectional long - short - term memory network module.

8. The automatic book cataloging method according to claim 7, wherein The process of generating a book classification label based on the keyword vectors using a neural network module and a bidirectional long - short - term memory network module includes: Generate a first vector and a second vector for the keyword vectors through the neural network module and the bidirectional long - short - term memory network module respectively; Perform a weighted sum of the first vector and the second vector to generate a comprehensive feature vector; Generate the book classification label by passing the comprehensive feature vector through a fully - connected layer and an activation function in sequence.

9. The automatic book cataloging method according to any one of claims 1 to 8, characterized in that The process of classifying the description text into a title text and an introduction text according to the positional relationship between the name text and the description text, the area of the name text of the name text, and the area of the title text of the description text includes: Divide the description text into multi - line text lines and calculate the height ratio between the multi - line text lines; If the height ratio is greater than the ratio threshold, determine the text line with the highest height as the title text, and the remaining text lines as the introduction text; If the height ratio is less than or equal to the ratio threshold, determine the title text and the introduction text according to the weighted calculation result of the height ratio, the distance between the center points of the name text and the description text.

10. The automatic book cataloging method according to any one of claims 1 to 8, characterized in that, The catalog text block diagram includes a name text block diagram, a description text block diagram, a price block diagram, a publisher block diagram, a publication date block diagram, a collection location block diagram, and a call number block diagram; The catalog text block diagram is also converted through an OCR model to generate other information text; Extract keywords from the other information text to generate a description text, price, publisher, publication date, collection location, and call number; Extract keywords from the name text to generate author name, translator name, and editor name.

Citation Information

Patent Citations

  • Book name positioning and part-of-speech tagging method and system

    CN110197175A

  • Bill element extraction method and device, electronic equipment and readable storage medium

    CN111914835A

  • Identification method and device, computer equipment and storage medium

    CN114299500A

  • Library book classification method based on content keywords and neural network

    CN115062105A

  • Automatic cataloguing method, system and equipment for cases and storage medium

    CN115880704A