Book CIP detection and identification optimization method and system based on multi-modal algorithm
Through the preprocessing of multimodal algorithms, CIP area detection and text recognition optimization, the instability problem of OCR technology in the positioning and recognition of book CIP information is solved, and high-precision CIP information extraction is realized, which is suitable for automated book management systems.
Patent Information
- Application Number
- CN202510281044.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-22
AI Technical Summary
When processing book CIP information, existing OCR technology is affected by layout, font size, book inclination and lighting conditions, resulting in unstable positioning of text boxes, making it difficult to accurately locate CIP areas, especially in multi-page scanning or batch detection, and it is difficult to identify special characters such as Roman characters.
Multimodal algorithm is used to obtain book images for pre-processing, use the YOLO algorithm to accurately locate the CIP region, and combine the OCR model of Transformer or CRNN for text recognition, and then conduct context correlation correction and specific syntax structure processing to improve the recognition accuracy.
It realizes accurate positioning and identification of book CIP information in complex environments, improves overall recognition accuracy, reduces manual intervention, and is suitable for automated book management systems.
Smart Images

Figure CN120356226A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Cataloging in Publication (CIP) recognition methods for books, and particularly to an optimized method and system for detecting and recognizing book CIP based on multimodal algorithms. Background Art
[0002] CIP (Cataloging in Publication) information is crucial metadata in the book publishing process, covering basic information such as book title, author, publication year, ISBN (International Standard Book Number), etc. Accurately collecting and managing this information is essential for book publishing, sales, and library management. In practical applications, traditional OCR (Optical Character Recognition) technology faces some challenges when processing CIP information. Existing OCR technology is often affected by layout, font size, book tilt, or lighting conditions when detecting text boxes, resulting in unstable text box positions or shape changes. Especially in scenarios of multi-page scanning or batch detection, the detection results often lack consistency, bringing greater interference to automated information extraction. A large number of Roman characters are included in the CIP information of books, such as ISBN numbers, author names, and publisher names. Due to the significant structural differences between Roman characters and other characters such as Chinese, traditional OCR technology often has difficulty accurately recognizing these characters, resulting in chaotic or incorrect recognition results. Most existing OCR models are trained based on general datasets and cannot be optimized for the special layout and content characteristics of CIP information. For example, ISBN numbers have a specific format, and the layout of table information is fixed, while traditional OCR training data cannot cover these characteristics, resulting in poor model performance in CIP information recognition. In the scanned or photographed images of books, CIP information is usually located in a specific area. However, due to cluttered backgrounds, edge occlusion, or content misalignment, a single OCR technology often cannot accurately locate the area where CIP information is located.
[0003] Therefore, it is necessary to design a new method to accurately locate the CIP area in the book, thereby improving the overall recognition accuracy. Summary of the Invention
[0004] The purpose of the present invention is to overcome the defects of the prior art and provide an optimized method and system for detecting and recognizing book CIP based on multimodal algorithms.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: An optimized method for detecting and recognizing book CIP based on multimodal algorithms, including:
[0006] Obtain the book image to be recognized;
[0007] Preprocess the book image to be recognized to obtain a preprocessing result;
[0008] Detect the CIP area of the book from the preprocessing result to obtain a detection result;
[0009] Input the detection result into a text recognition model to recognize the basic information of the book to obtain a recognition result;
[0010] Post-process and optimize the recognition result to obtain an optimized result;
[0011] Output the optimized result.
[0012] A further technical solution thereof is: the preprocessing of the book image to be recognized to obtain a preprocessing result includes:
[0013] Convert the book image to be recognized into grayscale to obtain a grayscale image;
[0014] Normalize the grayscale image to obtain a preprocessing result.
[0015] A further technical solution thereof is: the detection of the CIP area of the book from the preprocessing result to obtain a detection result includes:
[0016] Input the preprocessing result into a target detection model to detect the CIP area of the book to obtain a detection result;
[0017] Wherein, the target detection model is a model obtained by training the YOLO algorithm with book images containing CIP areas with different layouts, lighting conditions, and tilt angles as a sample set.
[0018] A further technical solution thereof is: the sample set used in the training of the target detection model covers CIP areas in multiple languages, multiple fonts, and multiple layouts.
[0019] A further technical solution thereof is: the text recognition model is obtained by training an OCR model based on Transformer or CRNN with a sample set of a number of CIP areas with different fonts and structures.
[0020] A further technical solution thereof is: the detection of the CIP area of the book from the preprocessing result to obtain a detection result includes:
[0021] Perform multiple detections of the CIP area of the book on the preprocessing result, and use the method of weighted average to unify the results of multiple detections to obtain a detection result.
[0022] Its further technical solution is: post-processing and optimizing the recognition result to obtain an optimized result, including:
[0023] Performing context correlation correction on the recognition result and representing the author name in the recognition result using a specific grammatical structure to obtain an optimized result.
[0024] The present invention also provides a book CIP detection and recognition optimization system based on a multimodal algorithm, including:
[0025] An image acquisition unit for acquiring a book image to be recognized;
[0026] A preprocessing unit for preprocessing the book image to be recognized to obtain a preprocessing result;
[0027] A region detection unit for detecting the book CIP region of the preprocessing result to obtain a detection result;
[0028] A recognition unit for inputting the detection result into a text recognition model to recognize the basic information of the book to obtain a recognition result;
[0029] An optimization unit for post-processing and optimizing the recognition result to obtain an optimized result;
[0030] An output unit for outputting the optimized result.
[0031] Its further technical solution is: the preprocessing unit includes:
[0032] A grayscale processing subunit for performing grayscale processing on the book image to be recognized to obtain a grayscale image;
[0033] A normalization processing subunit for performing normalization processing on the grayscale image to obtain a preprocessing result.
[0034] Its further technical solution is: the region detection unit is used to input the preprocessing result into a target detection model for book CIP region detection to obtain a detection result;
[0035] Wherein, the target detection model is a model obtained by training the YOLO algorithm with book images containing CIP regions with different layouts, lighting conditions, and tilt angles as a sample set.
[0036] The beneficial effects of the present invention compared with the prior art are as follows: By acquiring the book image to be recognized and performing preprocessing, a good foundation is laid for subsequent detection and recognition. Then, using the book CIP area detection algorithm, the CIP area in the book is accurately located from the preprocessing results. Next, the detected CIP area is input into the text recognition model to recognize the basic information of the book; subsequently, through post-processing optimization techniques, the accuracy of the recognition result is further improved; finally, the optimized result is output to ensure more accurate recognition of the overall book information.
[0037] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic diagram of the application scenario of the book CIP detection and recognition optimization method based on multi-modal algorithms provided by the embodiments of the present invention;
[0040] Figure 2 It is a schematic flow diagram of the book CIP detection and recognition optimization method based on multi-modal algorithms provided by the embodiments of the present invention;
[0041] Figure 3 It is a schematic sub-flow diagram of the book CIP detection and recognition optimization method based on multi-modal algorithms provided by the embodiments of the present invention;
[0042] Figure 4 It is a schematic block diagram of the book CIP detection and recognition optimization system based on multi-modal algorithms provided by the embodiments of the present invention;
[0043] Figure 5 It is a schematic block diagram of the preprocessing unit of the book CIP detection and recognition optimization system based on multi-modal algorithms provided by the embodiments of the present invention;
[0044] Figure 6 It is a schematic block diagram of the computer device provided by the embodiments of the present invention. Detailed Embodiments
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0047] It should also be understood that the terms used in this specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0048] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0049] Please refer to Figure 1 and Figure 2 , Figure 1 which is a schematic diagram of the application scenario of the optimized method for detecting and identifying the CIP of a book based on a multimodal algorithm provided by an embodiment of the present invention. Figure 2 which is a schematic flowchart of the optimized method for detecting and identifying the CIP of a book based on a multimodal algorithm. The optimized method for detecting and identifying the CIP of a book based on a multimodal algorithm is applied to a server. The server interacts with a camera or a scanning device, and processes the book image to be recognized through preprocessing steps, including grayscale conversion and normalization, to make it suitable for subsequent analysis; uses an object detection model (such as the YOLO algorithm) to detect the CIP area of the preprocessed image to ensure that the CIP area can be accurately found; when training the object detection model, uses book images containing diverse samples (such as different languages, fonts, and layouts) to improve the generalization ability of the model; adopts multiple CIP area detections, and combines multiple detection results through a weighted average method to obtain a more stable and accurate area positioning; uses an OCR model based on Transformer or CRNN for text recognition to ensure that various CIP text information with different fonts and structures can be recognized; through post-processing optimization, including context association correction and specific grammar processing of the author's name, further improves the accuracy of the recognition result.
[0050] Figure 2 It is a schematic flowchart of an optimized method for book CIP detection and recognition based on a multimodal algorithm provided by an embodiment of the present invention. As Figure 2 shown, this method includes the following steps S110 to S160.
[0051] S110. Obtain the book image to be recognized.
[0052] In this embodiment, book images containing CIP information are obtained from various input sources (such as scanners or cameras). These images may have problems such as noise, uneven illumination, and tilt. These book images can be scanned copies or pictures taken by a camera.
[0053] S120. Preprocess the book image to be recognized to obtain a preprocessing result.
[0054] In this embodiment, the preprocessing result refers to the image after the above grayscale conversion and normalization processing; it has the following characteristics:
[0055] All images are adjusted to the same size, ensuring the consistency and repeatability of subsequent processing.
[0056] Through denoising, grayscale conversion, and brightness equalization processing, the text and other key information in the image become clearer, reducing the risk of misrecognition caused by the quality problems of the original image.
[0057] The preprocessed image is more suitable for subsequent region detection (such as book CIP region localization based on YOLO), text recognition (OCR), and post-processing optimization steps.
[0058] In summary, the preprocessing step aims to improve the image quality so that subsequent analysis and processing are more accurate and effective. The preprocessing result refers to the image after a series of processing operations, and this image has been optimized to better adapt to subsequent algorithm processing (such as YOLO detection and OCR recognition).
[0059] In one embodiment, please refer to Figure 3 , the above step S120 may include steps S121 to S122.
[0060] S121. Perform grayscale conversion on the book image to be recognized to obtain a grayscale image.
[0061] In this embodiment, the grayscale image refers to the book image after grayscale processing.
[0062] Specifically, a grayscale algorithm is adopted to convert the color value of each pixel into a brightness value, forming a grayscale image. This step helps to simplify the image data while retaining sufficient details for subsequent processing, reducing the interference caused by color information and enabling subsequent processing to focus more on text content extraction.
[0063] S122. Normalize the grayscale image to obtain a preprocessing result.
[0064] In this embodiment, the grayscale image is adjusted to a fixed size (e.g., 640×640 pixels) to adapt to the subsequent model input requirements. This is usually achieved through scaling or cropping.
[0065] For overly dark or bright images, apply brightness equalization techniques to improve contrast and clarity. For example, use methods such as histogram equalization or adaptive gamma correction to make the details in the image more obvious.
[0066] After normalization processing, the interference caused by color information is reduced, enabling subsequent processing to focus more on text content extraction.
[0067] S130. Detect the CIP area of the book in the preprocessing result to obtain a detection result.
[0068] In this embodiment, the detection result specifically refers to the area containing CIP information accurately located in the preprocessed image through a target detection model. These areas usually include the positions of important publication information such as ISBN numbers, book titles, and author names.
[0069] Specifically, input the preprocessing result into the target detection model for book CIP area detection to obtain a detection result;
[0070] Among them, the target detection model is a model obtained by training the YOLO algorithm with a sample set of book images with CIP areas of different formats, lighting conditions, and tilt angles. The sample set used in training the target detection model covers CIP areas in multiple languages, multiple fonts, and multiple formats.
[0071] Take the preprocessing result obtained in step S120 (i.e., the grayscale image with standardized size and brightness equalization) as the input and feed it into a pre-trained target detection model. Here, a target detection model optimized and trained based on the YOLO (You Only Look Once) algorithm is used. YOLO is a real-time object detection system widely used in various scenarios for its fast speed and high accuracy.
[0072] The target detection model used is optimized and trained specifically for the CIP information area of books. The training sample set includes book images of CIP areas under different layouts, lighting conditions, and tilt angles, ensuring that the model has good robustness and adaptability.
[0073] The YOLO model first extracts features from the input image to generate feature maps at multiple scales. Based on these feature maps, the model predicts the positions of bounding boxes that may contain CIP information and their class probabilities. Each prediction box not only contains the rectangular coordinates (x, y, w, h) but also comes with a probability score indicating the presence of CIP information within the box. To remove redundant overlapping prediction boxes, the non-maximum suppression algorithm is applied to retain only the most likely few candidate boxes as the final detection results.
[0074] The detection result is not just a set of coordinates of one or more rectangular boxes. It also includes the confidence score corresponding to each box, which reflects the degree of certainty of the model that the box actually contains CIP information. For each box marked as the CIP area, its specific content can be further used for:
[0075] The OCR module in the next step will focus on performing text recognition tasks within these precisely delimited areas, thereby improving the recognition accuracy of the overall system. Necessary corrections are made according to the consistency and accuracy of the detection boxes, such as adjusting boundary offsets or resolving issues of inconsistent multiple detection results.
[0076] In summary, step S130 accurately identifies and locates the CIP information area in the book by applying a highly customized YOLO target detection model to the preprocessed image, laying a solid foundation for subsequent efficient text extraction and recognition. In this process, considering the diversity and complexity in the actual application scenario, through the combination of a carefully designed training sample set and advanced detection techniques, the applicability and reliability of the system in different environments are greatly improved.
[0077] In one embodiment, the preprocessing result is subjected to multiple detections of the book CPI area, and the results of multiple detections are unified using a weighted average method to obtain the detection result.
[0078] Specifically, when implementing a book CIP information detection and recognition system based on a multi-modal algorithm, to improve the accuracy and robustness of the detection result, the CIP area of the same input image is usually detected multiple times, and a weighted average method is used to synthesize the results of multiple detections.
[0079] Through multiple independent detection processes, the accidental errors that may exist due to single detection, such as local noise and lighting changes, can be reduced, which have an impact on the detection result.
[0080] Enhance the robustness of the system: Different detection runs may produce slightly different results due to random initialization or other factors. Detecting multiple times and integrating these results can effectively cope with the complex environmental conditions and variables in the input image.
[0081] 2. Implementation steps
[0082] S201. Perform multiple CIP region detections on the preprocessed image
[0083] First, ensure that the input image has undergone appropriate preprocessing steps, including but not limited to denoising, grayscale conversion, edge enhancement, and normalization, to provide a high-quality basis for subsequent accurate detection. Use the trained YOLO object detection model to repeatedly detect the same image. Before each detection, some slight changes or perturbations can be introduced, such as adjusting certain parameter settings of the detection model or changing the small rotation angle of the input image, to simulate different detection scenarios. Each detection will output one or more bounding boxes containing CIP information. Record the positions (x, y, w, h) of all detected bounding boxes and their corresponding confidence scores. Determine the weight of each detection box in the final result according to its confidence score. Generally speaking, the higher the confidence, the greater the possibility that there is indeed CIP information in the box, so a higher weight should be given.
[0084] For each potential CIP region, collect the bounding box coordinates (x, y, w, h) generated for this region during all detection processes.
[0085] Using the confidence scores of each box as weights, perform weighted average calculations on the corresponding dimensions (i.e., x, y, w, h) of these coordinates. For example, the weighted average of the width w can be calculated by the following formula: where, w i is the width value obtained from the i-th detection, score i is the corresponding confidence score, and n represents the total number of detections.
[0086] After the above weighted average processing, a set of optimized bounding box coordinates is obtained, which represents the information that is most likely to accurately reflect the actual position of the CIP region in the multiple detection results.
[0087] Apply the finally determined CIP region bounding box to the original image to check whether it correctly covers the expected CIP information part. If a large deviation is found, it may be necessary to adjust the detection strategy or retrain the model. Use the optimized bounding box to guide the OCR text recognition module to work, and only focus on extracting CIP-related information within these precisely located regions, thereby improving the recognition accuracy and efficiency of the overall system.
[0088] Through the above method, the information from multiple detections can be effectively integrated, problems that may occur in single detection can be overcome, and thus the performance of the entire book CIP detection and recognition system can be improved. This method is particularly suitable for dealing with challenging application scenarios such as complex backgrounds, uneven lighting, and tilted books.
[0089] S140. Input the detection result into the text recognition model to recognize the basic information of the book, so as to obtain the recognition result.
[0090] In this embodiment, the recognition result refers to the text content of the CIP area recognized by the text recognition model.
[0091] Specifically, the text recognition model is obtained by training an OCR model based on Transformer or CRNN using several CIP areas with different fonts and structures as the sample set.
[0092] Taking the accurate CIP area detection frame generated by the YOLO algorithm as the input and feeding it into the optimized text recognition model, so as to achieve the accurate recognition of the basic information of the book (such as ISBN number, book title, author name, etc.).
[0093] In this embodiment, the text recognition model used is an OCR model constructed based on the Transformer or CRNN (Convolutional Recurrent Neural Network) architecture. Both of these architectures have the ability to process sequence data and are very suitable for recognizing text information with specific structures and formats.
[0094] In order to ensure that the model can accurately recognize CIP information in different fonts, sizes, and layouts, a CIP area sample set covering multiple languages (especially Roman characters), fonts, and layouts is constructed. These samples include not only common printed characters but also consider handwriting styles and text deformations caused by scanning quality, etc.
[0095] Specifically, first, extract the corresponding image patches from the accurate CIP area detection frames obtained in the previous step (i.e., S130, book area detection). Each image patch represents an area that may contain useful information, such as ISBN number, book title, author name, etc. Perform necessary preprocessing operations on the extracted image patches, such as size adjustment, contrast enhancement, etc., so as to better meet the input requirements of the OCR model.
[0096] The preprocessed CIP region image blocks are input one by one into the trained OCR text recognition model. In this process, the OCR model automatically recognizes the characters in the image according to its internal parameter settings and converts them into an editable text form. Considering that CIP information is usually presented in paragraphs or tables, the OCR model processes the image in a line-by-line recognition manner. This means that for each CIP region, the OCR model first determines the boundaries of each line therein, and then sequentially performs character recognition on each line. This strategy helps to solve the problem of line skipping under complex layout and improve the recognition accuracy.
[0097] After the text recognition model completes the recognition of all input image blocks, a series of preliminary recognition results will be generated. These results include all visible characters in the CIP region and their position information.
[0098] S150. Post-process and optimize the recognition results to obtain optimized results.
[0099] In this embodiment, the optimized result refers to the result obtained after correcting and optimizing the recognition result.
[0100] Specifically, perform context-related correction on the recognition result, and represent the author's name in the recognition result using a specific grammatical structure to obtain the optimized result.
[0101] In this embodiment, a series of post-processing techniques are used to correct and optimize the preliminary recognition results output by the OCR model, so as to ensure that the finally output CIP information has higher accuracy and consistency.
[0102] Specifically, a method combining geometric constraints and semantic rules can also be used to further optimize the position of the detection frame and the recognized text content; improve the consistency and accuracy of the recognition result, and correct the text content according to specific rules (such as ISBN format, author name structure, etc.).
[0103] When performing optimization, use context information to verify and correct errors or inconsistencies in the recognition result. For example, the ISBN number must meet certain format requirements (10 or 13 digits), and the book title and author's name also have their specific grammatical structures.
[0104] For the ISBN number, check whether it conforms to the format of the International Standard Book Number. If not, try to repair it based on the known format rules or mark it as an anomaly.
[0105] For the author's name, specific grammatical structures are applied for verification. For example, Western author names usually follow the format of "Given Name Surname" or "Surname, Given Name", while Chinese names are mostly in the form of "Surname Given Name". The system will automatically adjust the recognition results according to these rules or prompt the user for manual confirmation.
[0106] Based on predefined grammatical rules and domain knowledge, logical checks and corrections are performed on the recognized text content.
[0107] Ensure that the recognized ISBN number conforms to the specified length and format. For a 13-digit ISBN, it is also necessary to verify whether its check digit is correct.
[0108] According to the habits of different languages, adjust the order of the author's name. For example, convert "Lastname,Firstname" to "Firstname Lastname" and vice versa.
[0109] Similarly, for the book title, publisher name, etc., corresponding formatting and normalization can also be performed according to the known knowledge base or rule set.
[0110] After completing all the above correction steps, integrate the optimized recognition results to form the final CIP information set.
[0111] These information not only include the basic book metadata (such as ISBN number, book title, author's name), but may also contain additional notes or tags to indicate certain special processing situations (for example, a certain field has been manually corrected).
[0112] S160. Output the optimized result.
[0113] In this embodiment, the finally output information should be sorted in a predetermined format for subsequent data management and application.
[0114] Through the above detailed post-processing optimization process, the method of this embodiment can significantly improve the accuracy and reliability of CIP information recognition. Especially in the face of complex layouts, blurred images or non-standard fonts, it can still maintain a high recognition quality. In addition, this method also provides a solid technical foundation for the implementation of an automated library management system, reducing the need for manual intervention and improving the overall work efficiency.
[0115] In addition, the method of this embodiment uses an OCR model based on Transformer or CRNN (Convolutional Recurrent Neural Network) specifically for recognizing the CIP region detected by YOLO. Through these deep learning architectures, complex texts can be effectively recognized and the accuracy of the model can be improved.
[0116] The text recognition model is fine-tuned with a dedicated CIP dataset to improve the recognition accuracy of specific information (such as Roman characters like ISBN). This targeted training enables the OCR model to better handle specific formats and characters, providing high-precision text recognition results.
[0117] For the text within the detection box, the model recognizes it as line-by-line text, and each line is processed individually to ensure more refined text extraction.
[0118] By designing an automatic line break detection algorithm, the system can handle the line skipping problem in complex layouts, ensuring that each line of content can be correctly recognized in text with a more complex layout structure and avoiding information loss.
[0119] Specifically, for a paragraph of text, the line spacing is a key indicator for judging whether to break a line. A line break usually occurs when the line spacing is greater than a certain threshold. Therefore, first, the vertical distance of each line needs to be calculated, and the line break position is determined based on this distance. If the vertical spacing between adjacent lines is less than the threshold, it is regarded as the same paragraph of text; otherwise, it is considered a line break.
[0120] Text lines may sometimes have a slight left or right offset (e.g., due to font and kerning issues), so it is necessary to detect the horizontal alignment of the lines. A common method is to check the left and right boundaries of each line of text. If the horizontal offset of two consecutive lines of text exceeds the preset threshold, it indicates that a line break has occurred.
[0121] For each line after segmentation, it is possible to determine whether to break a line by analyzing the bounding box of each character or word. Information such as the height and width of the character box is one of the bases for line breaks.
[0122] The text region in the image is extracted through the text recognition model or other text detection algorithms (such as YOLO or EAST). What is obtained in this step is the bounding box of the text block, rather than individual lines.
[0123] Sort the text boxes according to the positions of these bounding boxes. For multi-line text, first sort these boxes by the vertical coordinates to ensure that the text boxes are arranged line by line. If a significant inter-line gap is detected, it indicates that a line break has occurred.
[0124] Then, based on the line spacing and the vertical position of the characters, determine the exact line break position. For example, a variable threshold can be set to judge whether a line break is needed between adjacent lines according to the actual position of the text boxes. For complex layouts, this threshold can be dynamically adjusted according to the actual situation to adapt to different types of layouts.
[0125] For some line jumps caused by text alignment issues or detection errors, rearrangement can be performed through a position-based correction algorithm. For example, if the starting position of a line is far from the end of the previous line, this line can be considered a new start to prevent text lines from being misjudged as the same line.
[0126] In some complex layouts, the text may be arranged in the form of a table or columns. At this time, the row and column boundaries of the table can be identified through a table detection algorithm (such as a line segment detection method based on the Hough transform), and then the text in each cell can be processed separately to ensure the accuracy of recognition.
[0127] In some complex layouts, the text may not be completely horizontally aligned. It can be adjusted by analyzing the tilt angle of each line of text to determine whether the text needs to be rotated or its baseline adjusted.
[0128] By analyzing the context content, the algorithm can determine whether a text paragraph should continue on the same line. For example, if the text content contains some obvious conjunctions, punctuation marks, etc., these clues can be used to enhance the accuracy of line break detection. The timing of line breaks depends not only on the position but also on the structure of the text content.
[0129] Through deep learning models, especially those based on convolutional neural networks (CNNs) or recurrent neural networks (RNNs), more spatial layout features can be learned. By training the model, the system can better capture the line break patterns in complex layouts, thereby optimizing the accuracy of line break detection.
[0130] In complex layouts, a fixed line spacing threshold may not be suitable for different situations. An adaptive algorithm can be designed to dynamically adjust the line spacing threshold according to the distribution of the text, or statistical methods (such as dynamic adjustment based on the mean and variance) can be used to optimize the effect of line break detection.
[0131] For example, if it is detected that the text area is relatively dense, the determination conditions for the line spacing may need to be appropriately relaxed, and vice versa tightened. The boundaries of each line can be optimized by combining global information during processing.
[0132] Finally, the algorithm is evaluated using manually annotated data or preprocessed layout data. The accuracy of line break detection is measured through metrics such as precision and recall. If the detection results are not satisfactory, the algorithm can be further optimized by adjusting the parameters in the algorithm, adding additional context information, or adopting a more refined deep learning model.
[0133] For the above text recognition model, TensorRT technology is used to comprehensively optimize the OCR model, mainly including operator fusion, model quantization (such as INT8 quantization), and memory management optimization. These optimization strategies help reduce computational overhead, improve inference speed, while reducing memory occupancy and enhancing overall performance.
[0134] Through TensorRT optimization, the inference process has been significantly accelerated, with the inference speed increased by about 30% to 50%. This acceleration enables the model to achieve real-time response, suitable for large-scale application scenarios, ensuring fast and accurate recognition results under high load.
[0135] Through multiple technical optimizations of TensorRT, including model quantization (such as using INT8 quantization technology), removing redundant operators, and optimizing the computation graph, the inference efficiency and speed of the OCR model have been greatly improved. This series of optimizations has brought about a 30%-50% increase in inference speed, enabling the OCR system to respond quickly and significantly improve the processing efficiency. Deploying the TensorRT-optimized model in the NVIDIA GPU environment makes full use of the parallel computing power of the GPU to accelerate the inference process. This can support efficient inference calculations, especially suitable for large-scale application scenarios that require rapid processing of large amounts of data.
[0136] The above-mentioned optimized method for book CIP detection and recognition based on multimodal algorithms lays a foundation for subsequent detection and recognition by obtaining the book image to be recognized and performing preprocessing. Then, the book CIP region detection algorithm is used to accurately locate the CIP region in the book from the preprocessing results. Next, the detected CIP region is input into the text recognition model to recognize the basic information of the book. Subsequently, through post-processing optimization techniques, the accuracy of the recognition results is further improved. Finally, the optimized results are output to ensure more accurate overall book information recognition.
[0137] Figure 4 It is a schematic block diagram of an optimized system 300 for book CIP detection and recognition based on multimodal algorithms provided by an embodiment of the present invention. As Figure 4 shown, corresponding to the above-mentioned optimized method for book CIP detection and recognition based on multimodal algorithms, the present invention also provides an optimized system 300 for book CIP detection and recognition based on multimodal algorithms. The optimized system 300 for book CIP detection and recognition based on multimodal algorithms includes units for executing the above-mentioned optimized method for book CIP detection and recognition based on multimodal algorithms, and this system can be configured in a server. Specifically, please refer to Figure 4, the book CIP detection and recognition optimization system 300 based on multimodal algorithms includes an image acquisition unit 301, a preprocessing unit 302, a region detection unit 303, an identification unit 304, an optimization unit 305, and an output unit 306.
[0138] The image acquisition unit 301 is used to acquire the book image to be recognized; the preprocessing unit 302 is used to preprocess the book image to be recognized to obtain a preprocessing result; the region detection unit 303 is used to detect the book CIP region of the preprocessing result to obtain a detection result; the identification unit 304 is used to input the detection result into a text recognition model to recognize the basic information of the book to obtain an identification result; the optimization unit 305 is used to post-process and optimize the identification result to obtain an optimization result; the output unit 306 is used to output the optimization result.
[0139] In one embodiment, as Figure 5 shown, the preprocessing unit 302 includes:
[0140] The grayscale processing sub-unit 3021 is used to perform grayscale processing on the book image to be recognized to obtain a grayscale image;
[0141] The normalization processing sub-unit 3022 is used to perform normalization processing on the grayscale image to obtain a preprocessing result.
[0142] In one embodiment, the region detection unit 303 is used to input the preprocessing result into a target detection model to detect the book CIP region to obtain a detection result;
[0143] Among them, the target detection model is a model obtained by training the YOLO algorithm with book images containing CIP regions with different layouts, lighting conditions, and tilt angles as a sample set.
[0144] In one embodiment, the region detection unit 303 is further used for:
[0145] Performing multiple book CPI region detections on the preprocessing result, and using the method of weighted average to unify the results of multiple detections to obtain a detection result.
[0146] In one embodiment, the optimization unit 305 is used to perform context-related correction on the recognition result and represent the author's name in the recognition result using a specific grammar structure to obtain an optimization result.
[0147] It should be noted that those skilled in the art can clearly understand the specific implementation processes of the above-mentioned book CIP detection and recognition optimization system 300 based on the multimodal algorithm and each unit. They can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and conciseness of description, they will not be elaborated here.
[0148] The above-mentioned book CIP detection and recognition optimization system 300 based on the multimodal algorithm can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 6 the following.
[0149] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a server. Among them, the server can be an independent server or a server cluster composed of multiple servers.
[0150] Referring to Figure 6 , the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory can include a non-volatile storage medium 503 and an internal memory 504.
[0151] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions. When the program instructions are executed, the processor 502 can be made to execute a book CIP detection and recognition optimization method based on the multimodal algorithm.
[0152] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0153] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be made to execute a book CIP detection and recognition optimization method based on the multimodal algorithm.
[0154] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0155] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the following steps:
[0156] Obtain a book image to be recognized; preprocess the book image to be recognized to obtain a preprocessing result; perform book CIP area detection on the preprocessing result to obtain a detection result; input the detection result into a text recognition model to recognize basic book information to obtain a recognition result; perform post-processing optimization on the recognition result to obtain an optimized result; output the optimized result.
[0157] Among them, the text recognition model is obtained by training an OCR model based on Transformer or CRNN using several CIP areas with different fonts and structures as a sample set.
[0158] In one embodiment, when the processor 502 implements the step of preprocessing the book image to be recognized to obtain a preprocessing result, the specific implementation is as follows:
[0159] Perform grayscale processing on the book image to be recognized to obtain a grayscale image; perform normalization processing on the grayscale image to obtain a preprocessing result.
[0160] In one embodiment, when the processor 502 implements the step of performing book CIP area detection on the preprocessing result to obtain a detection result, the specific implementation is as follows:
[0161] Input the preprocessing result into a target detection model to perform book CIP area detection to obtain a detection result;
[0162] Among them, the target detection model is a model obtained by training the YOLO algorithm using book images with CIP areas including different layouts, lighting conditions, and tilt angles as a sample set.
[0163] The sample set used in the training of the target detection model covers CIP areas in multiple languages, multiple fonts, and multiple layouts.
[0164] In one embodiment, when the processor 502 implements the step of performing book CIP area detection on the preprocessing result to obtain a detection result, the specific implementation is as follows:
[0165] Perform multiple book CPI area detections on the preprocessing result, and use the method of weighted average to unify the results of multiple detections to obtain a detection result.
[0166] In one embodiment, when the processor 502 implements the step of performing post-processing optimization on the recognition result to obtain an optimized result, the specific implementation is as follows:
[0167] Perform context - related correction on the recognition result, and represent the author's name in the recognition result using a specific grammatical structure to obtain an optimized result.
[0168] It should be understood that in the embodiments of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general - purpose processors, digital signal processors (DSPs), application - specific integrated circuits (ASICs), field - programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general - purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0169] Those of ordinary skill in the art can understand that all or part of the processes in the methods of implementing the above - mentioned embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer - readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above - mentioned methods.
[0170] Therefore, the present invention also provides a storage medium. The storage medium may be a computer - readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the processor performs the following steps:
[0171] Obtain a book image to be recognized; pre - process the book image to be recognized to obtain a pre - processing result; detect the CIP area of the book for the pre - processing result to obtain a detection result; input the detection result into a text recognition model to recognize the basic information of the book to obtain a recognition result; perform post - processing optimization on the recognition result to obtain an optimized result; output the optimized result.
[0172] Among them, the text recognition model is obtained by training an OCR model based on Transformer or CRNN using several CIP areas with different fonts and structures as a sample set.
[0173] In an embodiment, when the processor executes the computer program to implement the step of pre - processing the book image to be recognized to obtain a pre - processing result, the following steps are specifically implemented:
[0174] Perform grayscale processing on the book image to be recognized to obtain a grayscale image; perform normalization processing on the grayscale image to obtain a preprocessing result.
[0175] In one embodiment, when the processor executes the computer program to implement the step of detecting the CIP area of the book for the preprocessing result to obtain a detection result, the specific implementation is as follows:
[0176] Input the preprocessing result into the target detection model for CIP area detection of the book to obtain a detection result;
[0177] Among them, the target detection model is a model obtained by training the YOLO algorithm with book images containing CIP areas with different formats, lighting conditions, and tilt angles as the sample set.
[0178] The sample set used in the training of the target detection model covers CIP areas in multiple languages, multiple fonts, and multiple formats.
[0179] In one embodiment, when the processor executes the computer program to implement the step of detecting the CIP area of the book for the preprocessing result to obtain a detection result, the specific implementation is as follows:
[0180] Perform multiple CPI area detections on the preprocessing result, and use the method of weighted average to unify the results of multiple detections to obtain a detection result.
[0181] In one embodiment, when the processor executes the computer program to implement the step of post-processing and optimizing the recognition result to obtain an optimized result, the specific implementation is as follows:
[0182] Perform context-related correction on the recognition result, and represent the author's name in the recognition result using a specific grammar structure to obtain an optimized result.
[0183] The storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.
[0184] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0185] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0186] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the system embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0187] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.
[0188] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An optimized method for book CIP detection and recognition based on multimodal algorithms, characterized in that Including: Obtain the book image to be recognized; Preprocess the book image to be recognized to obtain a preprocessing result; Perform book CIP area detection on the preprocessing result to obtain a detection result; Input the detection result into a text recognition model to recognize the basic information of the book to obtain a recognition result; Perform post-processing optimization on the recognition result to obtain an optimized result; Output the optimized result.
2. The optimized method for detecting and identifying the CIP of a book based on a multimodal algorithm according to claim 1, wherein, The preprocessing of the book image to be recognized to obtain a preprocessing result includes: Perform grayscale processing on the book image to be recognized to obtain a grayscale image; Perform normalization processing on the grayscale image to obtain a preprocessing result.
3. The optimized method for detecting and identifying the CIP of a book based on a multimodal algorithm according to claim 1, characterized in that, The book CIP area detection on the preprocessing result to obtain a detection result includes: Input the preprocessing result into a target detection model to perform book CIP area detection to obtain a detection result; Wherein, the target detection model is a model obtained by training the YOLO algorithm with a book image containing CIP areas with different layouts, lighting conditions, and tilt angles as a sample set.
4. The optimized method for detecting and identifying book CIP based on a multimodal algorithm according to claim 3, characterized in that, The sample set used in the training of the target detection model covers CIP areas in multiple languages, multiple fonts, and multiple layouts.
5. The optimized method for book CIP detection and recognition based on multimodal algorithms according to claim 1, characterized in that, The text recognition model is obtained by training an OCR model based on Transformer or CRNN with a number of CIP areas with different fonts and structures as a sample set.
6. The optimized method for detecting and identifying the CIP of a book based on a multimodal algorithm according to claim 1, wherein The book CIP area detection on the preprocessing result to obtain a detection result includes: Perform multiple book CPI area detections on the preprocessing result, and use the method of weighted average to unify the results of multiple detections to obtain a detection result.
7. The optimized method for detecting and identifying the CIP of a book based on a multimodal algorithm according to claim 1, characterized in that, The post-processing optimization of the recognition result to obtain an optimized result includes: Perform context-related correction on the recognition result, and represent the author's name in the recognition result using a specific grammar structure to obtain an optimized result.
8. The optimized system for detecting and identifying the CIP of books based on multimodal algorithms, characterized in that, Including: An image acquisition unit for obtaining the book image to be recognized; A preprocessing unit for preprocessing the book image to be recognized to obtain a preprocessing result; An area detection unit for performing book CIP area detection on the preprocessing result to obtain a detection result; A recognition unit for inputting the detection result into a text recognition model to recognize the basic information of the book to obtain a recognition result; An optimization unit for performing post-processing optimization on the recognition result to obtain an optimized result; An output unit for outputting the optimized result.
9. The optimized book CIP detection and recognition system based on the multimodal algorithm according to claim 8, characterized in that, The preprocessing unit includes: A grayscale processing subunit for performing grayscale processing on the book image to be recognized to obtain a grayscale image; A normalization processing subunit for performing normalization processing on the grayscale image to obtain a preprocessing result.
10. The optimized system for detecting and identifying the CIP of a book based on a multimodal algorithm according to claim 8, characterized in that, The area detection unit for inputting the preprocessing result into a target detection model to perform book CIP area detection to obtain a detection result; Wherein, the target detection model is a model obtained by training the YOLO algorithm with a book image containing CIP areas with different layouts, lighting conditions, and tilt angles as a sample set.