Book note digital recycling and sharing method and system based on OCR-NLP collaboration

Through OCR-NLP collaborative technology, the automatic identification and evaluation of second-hand books and notes content is solved, and the problem of inefficiency of second-hand book recycling platforms is achieved, and intelligent pricing and efficient utilization of knowledge resources are achieved.

CN120340052APending Publication Date: 2025-07-18SHANGHAI NORMAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510408375.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing second-hand book recycling platform is inefficient, relies on manual evaluation and strong subjectivity to pricing, and it is difficult to accurately identify and organize handwritten notes, resulting in the inability to efficiently utilize knowledge resources.

Method used

The OCR hybrid identification model and the NLP note quality assessment model are used to automatically identify books and notes content, and pricing results are generated through OCR identification coefficients, note value coefficients and original book price, and combined with the damage degree judgment model to achieve automatic quality evaluation and intelligent pricing.

Benefits of technology

It has improved the efficiency of second-hand book recycling, ensured the fairness and accuracy of pricing, and built a knowledge resource platform that integrates second-hand book and note recycling, realizing the accurate identification and value evaluation of handwritten notes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340052A_ABST
    Figure CN120340052A_ABST
Patent Text Reader

Abstract

The invention provides a book note digital recycling and sharing method and system based on OCR-NLP collaboration, and relates to the technical field of knowledge management and cyclic utilization. An OCR hybrid recognition model and an NLP note quality evaluation model are constructed; acquiring a book picture uploaded by a seller user; inputting the book picture into an OCR mixed recognition model, and obtaining an OCR recognition coefficient, a book text, a note text and book information; matching corresponding book classifications according to the book information, and carrying out tagging processing and incorporating the book classifications into a knowledge base; inputting the book text and the note text into an NLP note quality evaluation model to obtain a note value coefficient; generating a pricing result according to the OCR identification coefficient, the note value coefficient and the book original price; after the books are attached with corresponding prices, the books are put on the shelf and displayed as commodities, automatic quality evaluation and intelligent pricing of the second-hand books are achieved, fairness and accuracy of pricing are guaranteed, recycling and pricing efficiency is improved, and a knowledge resource platform integrating second-hand book and note recycling and sharing is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge management and recycling, and in particular to a method and system for digitalizing, recycling and sharing books and notes based on OCR-NLP collaboration. Background Art

[0002] With the explosive growth of knowledge, books, as an important carrier of knowledge dissemination, are being updated at an increasingly rapid pace, resulting in a large number of second-hand books being idle or discarded, causing a waste of resources. At the same time, handwritten notes, as an important way to record personal learning and thinking, often contain rich knowledge points and personalized insights, but these notes often exist in fragmented form and are difficult to be effectively organized and used. Traditional second-hand book recycling and knowledge management methods have many shortcomings, such as low recycling efficiency, inaccurate recognition of note content, and difficulty in knowledge integration and reuse.

[0003] Specifically, most existing second-hand book recycling platforms rely on manual evaluation and pricing of books, which is inefficient and highly subjective; and the processing of handwritten notes mostly relies on manual reading and analysis, which is not only time-consuming and labor-intensive, but also difficult to accurately extract key knowledge points in the notes and organize them in a structured manner. In addition, the re-creation and circulation of knowledge resources also lack an efficient and convenient platform, which makes it impossible to effectively utilize a large number of valuable knowledge resources.

[0004] Based on this, the present invention is proposed. Summary of the invention

[0005] The purpose of the present invention is to provide a method and system for digital recycling and sharing of books and notes based on OCR-NLP collaboration, to realize automatic quality evaluation and intelligent pricing of second-hand books, to ensure the fairness and accuracy of pricing, to improve the efficiency of recycling pricing, and to build a knowledge resource platform for recycling and sharing of second-hand books and notes.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] The present invention provides, in a first aspect, a method for digital recycling and sharing of book notes based on OCR-NLP collaboration, mainly including the following steps: constructing an OCR hybrid recognition model and an NLP note quality evaluation model; obtaining book pictures uploaded by seller users that contain book content and note content; inputting the book pictures into the OCR hybrid recognition model to obtain an OCR recognition coefficient, book text, note text, and book information including the book title and original price; matching the corresponding book classification according to the book information, and incorporating it into the knowledge base after being tagged; inputting the book text and note text into the NLP note quality evaluation model to establish the value association between the note text and the book text, and obtaining a note value coefficient; generating a pricing result according to the OCR recognition coefficient, note value coefficient, and original price of the book; attaching the corresponding price to the book and putting it on the shelf and displaying it as a commodity.

[0008] As a preferred solution in the first aspect of the present invention, the method for digital recycling and sharing of book notes based on OCR-NLP collaboration further includes: constructing a damage degree judgment model based on deep learning; identifying the integrity of the book pictures, which is output when the OCR hybrid recognition model identifies that the book pictures are incomplete; inputting the incomplete book pictures into the damage degree judgment model based on deep learning to obtain a damage coefficient; and the pricing result is also generated in combination with the damage coefficient.

[0009] As a preferred solution in the first aspect of the present invention, the method for digital recycling and sharing of book notes based on OCR-NLP collaboration further includes: recommending corresponding commodities according to the purchase intention of buyer users; and entering the transaction process after obtaining the confirmation of the purchase intention of buyer users.

[0010] As a preferred solution in the first aspect of the present invention, when inputting the book pictures into the OCR hybrid recognition model to obtain the book text, note text, and book information including the book title and original price, the specific acquisition of the book text and note text includes the following: feature extraction and decoupling, extracting general text features through a shared underlying convolutional network, and fusing multi-scale features through a feature pyramid network; separating the printed and handwritten features through a gating mechanism, allocating them to different recognition branches, and dynamically allocating recognition weights; dual-branch parallel recognition, setting a printed text recognition branch and a handwritten text recognition branch to respectively recognize the printed text and the handwritten text; cross-modal association and structured output, establishing a cross-modal association by mapping the recognition results of the printed text and the handwritten text through a spatial coordinate system, constructing a tree-shaped semantic structure with the printed paragraph as the root node and the handwritten annotation as the child node and outputting it to obtain the book text and note text.

[0011] Further, the printed text recognition branch generates a printed text sequence through a deformable convolutional network and CTC decoding optimization with increased character spacing constraints; the handwritten text recognition branch generates a handwritten text sequence through adversarial generative data augmentation and an attention gating unit. The printed text recognition branch captures geometric invariant features in regular layouts through a deformable convolutional network, combines prior knowledge of character spacing to optimize text alignment accuracy, and generates a printed text sequence. The handwritten text recognition branch introduces adversarial generative data augmentation to simulate diverse writing styles and ink penetration effects, and strengthens the context association of cursive characters through an attention gating unit to generate a handwritten text sequence.

[0012] As a preferred solution in the first aspect of the present invention, inputting the book text and the note text into the NLP note quality assessment model, establishing the value association between the note text and the book text, and obtaining the note value coefficient specifically includes the following: text input, inputting the book text and the note text into the NLP note quality assessment model; semantic anchor alignment, using an improved dynamic time warping algorithm, combining the semantic constraints of the knowledge graph and character coordinate information, locating the association anchor points between the note text and the book text, and outputting the semantic matching degree; logical gain analysis, calling a subject-specific model to verify the logical rigor of the content of the aligned note text, where the application conditions of verification theorems in mathematics are verified, and the plot evidence in literature is verified, and the logical gain value is output; knowledge topology evaluation, mapping the note text after logical gain analysis to the corresponding subject knowledge graph, quantifying the contribution of the content of the note text to the corresponding subject knowledge graph through a graph attention network, and outputting the knowledge contribution value; comprehensive scoring and value coefficient output, constructing a dynamic weight allocation network through a gated recurrent unit, and the dynamic weight allocation network dynamically adjusts the weights of the semantic matching degree, the logical gain value, and the knowledge contribution degree according to subject characteristics, and calculates the weighted note comprehensive score and normalizes it to output the note value coefficient.

[0013] In a second aspect, the present invention provides a digital recycling and sharing system for book notes based on OCR-NLP collaboration for implementing the above method, including: a model construction module for constructing an OCR hybrid recognition model and an NLP note quality evaluation model; a picture acquisition module for acquiring book pictures uploaded by seller users that contain book content and note content; a text recognition module for inputting the book pictures into the OCR hybrid recognition model to obtain an OCR recognition coefficient, book text, note text, and book information including the book title and the original price of the book; a classification and induction module for matching the corresponding book classification according to the book information and incorporating it into the knowledge base after being tagged; a note quality evaluation module for inputting the book text and the note text into the NLP note quality evaluation model to establish a value association between the note text and the book text and obtain a note value coefficient; a pricing module for generating a pricing result based on the OCR recognition coefficient, the note value coefficient, and the original price of the book; and a listing module for listing the book with the corresponding price and displaying it as a commodity.

[0014] Compared with the prior art, the present invention has the following beneficial technical effects:

[0015] The present invention uses an OCR hybrid recognition model to accurately recognize handwritten notes, and deeply analyzes the recognition results through an NLP note quality evaluation model, establishes a value association between the note text and the book text, obtains a note value coefficient, and generates a pricing result based on the OCR recognition coefficient, the note value coefficient, and the original price of the book, so as to realize the automatic quality assessment and intelligent pricing of second-hand books, improve the recycling efficiency and reduce human intervention, and ensure the fairness and accuracy of pricing. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the method for digital recycling and sharing of book notes based on OCR-NLP collaboration provided in Embodiment 1 of the present invention;

[0018] Figure 2 It is a module diagram of the digital recycling and sharing system for book notes based on OCR-NLP collaboration provided in Embodiment 1 of the present invention;

[0019] Figure 3 It is a flowchart of the method for digital recycling and sharing of book notes based on OCR-NLP collaboration provided in Embodiment 2 of the present invention;

[0020] Figure 4 This is the module diagram of the book note digital recycling and sharing system based on OCR-NLP collaboration provided in Embodiment 2 of the present invention;

[0021] Figure 5 This is the flowchart of the book note digital recycling and sharing method based on OCR-NLP collaboration provided in Embodiment 3 of the present invention;

[0022] Figure 6 This is the data flow of the book note digital recycling and sharing method based on OCR-NLP collaboration provided in Embodiment 3 of the present invention;

[0023] Figure 7 This is the system flowchart of the book note digital recycling and sharing system based on OCR-NLP collaboration provided in Embodiment 3 of the present invention.

[0024] The reference numerals in the drawings are as follows: model construction module 100, picture acquisition module 200, text recognition module 300, classification and induction module 400, note quality evaluation module 500, pricing module 600, listing module 700, recommendation module 800, transaction module 900. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] Embodiment 1

[0027] Please refer to Figure 1 , and a preferred implementation manner is provided. A book note digital recycling and sharing method based on OCR-NLP collaboration provided by this implementation manner is mainly implemented through the following steps:

[0028] S100. Construct an OCR hybrid recognition model and an NLP note quality evaluation model;

[0029] S200. Obtain the book pictures uploaded by the seller user, which contain book content and note content;

[0030] S300. Input the book pictures into the OCR hybrid recognition model to obtain the OCR recognition coefficient, book text, note text, and book information including the book name and the original price of the book;

[0031] S400. Match the corresponding book classification according to the book information, and incorporate it into the knowledge base after tagging;

[0032] S500. Input the book text and the note text into the NLP note quality assessment model, establish the value association between the note text and the book text, and obtain the note value coefficient;

[0033] S600. Generate a pricing result based on the OCR recognition coefficient, the note value coefficient, and the original price of the book;

[0034] S700. Attach the corresponding price to the book, put it on the shelf, and display it as a commodity.

[0035] Please refer to Figure 2 Correspondingly, to implement the method for digital recycling and sharing of book notes based on OCR-NLP collaboration in this embodiment, a system for digital recycling and sharing of book notes based on OCR-NLP collaboration is provided, which mainly consists of the following modules:

[0036] The model construction module 100 is used to construct an OCR hybrid recognition model and an NLP note quality assessment model;

[0037] The image acquisition module 200 is used to acquire the book image uploaded by the seller user, which contains the book content and the note content;

[0038] The text recognition module 300 is used to input the book image into the OCR hybrid recognition model to obtain the OCR recognition coefficient, the book text, the note text, and the book information including the book name and the original price of the book;

[0039] The classification and induction module 400 is used to match the corresponding book classification according to the book information, and after being tagged, it is incorporated into the knowledge base;

[0040] The note quality assessment module 500 is used to input the book text and the note text into the NLP note quality assessment model, establish the value association between the note text and the book text, and obtain the note value coefficient;

[0041] The pricing module 600 is used to generate a pricing result based on the OCR recognition coefficient, the note value coefficient, and the original price of the book;

[0042] The shelving module 700 is used to attach the corresponding price to the book, put it on the shelf, and display it as a commodity.

[0043] This embodiment uses the OCR hybrid recognition model to accurately recognize the handwritten notes, and through the NLP note quality assessment model, deeply analyzes the recognition results, establishes the value association between the note text and the book text, obtains the OCR recognition coefficient and the note value coefficient, and generates a pricing result based on the OCR recognition coefficient, the note value coefficient, and the original price of the book, so as to realize the automatic quality assessment and intelligent pricing of second-hand books, improve the recycling efficiency, reduce human intervention, and ensure the fairness and accuracy of pricing.

[0044] Example 2

[0045] Please refer to Figure 3 , based on Example 1, the method for digital recycling and sharing of book notes based on OCR-NLP collaboration provided in this example also incorporates the degree of damage to the book into the pricing factor. It is mainly achieved through the following steps:

[0046] S110. Build an OCR hybrid recognition model, an NLP note quality assessment model, and a damage degree judgment model based on deep learning.

[0047] S200. Obtain the book pictures uploaded by the seller user, which contain book content and note content;

[0048] S310. Input the book picture into the OCR hybrid recognition model to first perform integrity recognition of the book picture;

[0049] S320. Output the book picture when the OCR hybrid recognition model recognizes that the book picture is incomplete;

[0050] S330. When the book picture is complete, obtain the OCR recognition coefficient, book text, note text, and book information including the book name and original price through the OCR hybrid recognition model;

[0051] S400. Match the corresponding book category according to the book information, and incorporate it into the knowledge base after tag processing;

[0052] S500. Input the book text and note text into the NLP note quality assessment model, establish the value association between the note text and the book text, and obtain the note value coefficient;

[0053] S510. Input the incomplete book picture into the damage degree judgment model based on deep learning to obtain the damage coefficient;

[0054] S601. Generate a pricing result according to the OCR recognition coefficient, note value coefficient, damage coefficient, and book original price;

[0055] S700. Attach the corresponding price to the book and put it on the shelf and display it as a commodity.

[0056] S800. Recommend corresponding commodities according to the purchase intention of the buyer user;

[0057] S900. Enter the transaction process after obtaining the confirmation of the purchase intention of the buyer user.

[0058] Please refer to Figure 4, correspondingly, to implement the method for digital recycling and sharing of book notes based on OCR-NLP collaboration in this embodiment, a system for digital recycling and sharing of book notes based on OCR-NLP collaboration is provided, which mainly consists of the following modules:

[0059] The model construction module 100 is used to construct an OCR hybrid recognition model and an NLP note quality evaluation model;

[0060] The picture acquisition module 200 is used to acquire book pictures uploaded by seller users, which contain book content and note content;

[0061] The text recognition module 300 is used to input the book picture into the OCR hybrid recognition model to obtain the OCR recognition coefficient, book text, note text, and book information including the book title and original price;

[0062] The classification and induction module 400 is used to match the corresponding book classification according to the book information, and after being processed by tagging, it is incorporated into the knowledge base;

[0063] The note quality evaluation module 500 is used to input the book text and note text into the NLP note quality evaluation model, establish the value association between the note text and the book text, and obtain the note value coefficient;

[0064] The pricing module 600 is used to generate a pricing result according to the OCR recognition coefficient, note value coefficient, damage coefficient, and book original price;

[0065] The listing module 700 is used to list the book with the corresponding price and display it as a commodity;

[0066] The recommendation module 800 is used to recommend corresponding commodities according to the purchase intention of buyer users;

[0067] The transaction module 900 is used to enter the transaction process after obtaining the confirmation of the purchase intention of the buyer user.

[0068] Generally speaking, the overall function design of the system for digital recycling and sharing of book notes based on OCR-NLP collaboration in this embodiment generally includes three major parts: quality inspection and pricing part, recommendation part, and transaction part. Among them, the quality inspection part consists of the model construction module 100, picture acquisition module 200, text recognition module 300, classification and induction module 400, note quality evaluation module 500, pricing module 600, and listing module 700. Seller users and buyer users conduct buying and selling transactions through the platform.

[0069] In this embodiment, a recommendation module and a transaction module are set up, integrating functions such as second-hand book recycling, in-depth development of handwritten notes, and knowledge resource sharing, to build a circulation platform that integrates knowledge acquisition, collation, and sharing. At the same time, in this embodiment, the degree of damage to the book is incorporated into the pricing factor, which can further improve the fairness and reasonableness of book pricing.

[0070] Embodiment 3

[0071] Please refer to Figure 5 and Figure 6 , based on Embodiment 2, the method for digital recycling and sharing of book notes based on OCR-NLP collaboration provided in this embodiment also considers book information rectification, supplementary recording mechanism, and pricing fine-tuning mechanism to more flexibly meet the expectations of seller users. It is mainly implemented through the following steps:

[0072] S110. Build an OCR hybrid recognition model, an NLP note quality evaluation model, and a damage degree judgment model based on deep learning.

[0073] S200. Obtain the book picture containing book content and note content uploaded by the seller user.

[0074] S210. Provide a book information supplementary input box to obtain the book information supplemented by the seller user; further, in a more preferred implementation manner, set that the seller user can select the physical state of the book, that is, the degree of damage (no damage / slight damage / serious damage) through a drop-down menu: slight damage is defined as the corner of the first page being folded and the edge being worn but not affecting reading; serious damage includes situations such as missing pages, large-area stains, or loose binding. The system will cross-verify the user's self-evaluation result with the integrity of the book picture recognized by the 0CR hybrid model. At the same time, the user needs to specify the subject attribute from 22 major categories of the Chinese Library Classification, such as "R Medicine and Health", and this information will restrict the subsequent knowledge graph call range;

[0075] S310. Input the book picture into the OCR hybrid recognition model to perform integrity recognition of the book picture;

[0076] S320. Output the book picture when the OCR hybrid recognition model recognizes that the book picture is incomplete;

[0077] When the book picture is complete, obtain the book text, note text, and book information including the book title and original price of the book through the OCR recognition coefficient of the OCR hybrid recognition model; the OCR recognition coefficient is an index of the recognizable degree of handwriting quantified by the OCR hybrid recognition model; evaluate the recognizable degree of the handwriting of the handwritten note through the OCR recognition coefficient output by the OCR hybrid model.

[0078] S410. Match the corresponding book classification according to the book information;

[0079] S420. Send the corresponding classified books obtained by matching to the seller user for confirmation or correction;

[0080] S430. When receiving the confirmation feedback from the seller user, attach the corresponding classification label; when receiving the correction feedback from the seller user, re-match the corresponding book classification according to the corrected or supplemented book information by the seller user, attach the corresponding classification label, and then incorporate it into the knowledge base;

[0081] S500. Input the book text and note text into the NLP note quality assessment model, establish the value association between the note text and the book text, and obtain the note value coefficient;

[0082] S510. Input the incomplete book picture into the damage degree judgment model based on deep learning to obtain the damage coefficient;

[0083] S601. Generate the pricing result according to the OCR recognition coefficient, note value coefficient, damage coefficient and the original price of the book; here, a key evaluation pricing algorithm model can be introduced to generate the pricing result as the book price and the recommended pricing range;

[0084] S610. Send the pricing result to the seller user for confirmation;

[0085] S620. When receiving the feedback of the seller user confirming the pricing, list the book with the corresponding price and display it as a commodity;

[0086] S630. When receiving the feedback from the seller user who has objections to the pricing, provide a pricing adjustment range not exceeding the floating ratio threshold for the seller user to adjust the pricing, or, re-price according to the seller user's choice, and let the seller user re-upload the book picture containing the note content to re-obtain the book picture for re-pricing;

[0087] S700. List the book with the corresponding price and display it as a commodity.

[0088] S800. Recommend corresponding commodities according to the purchase intention of the buyer user;

[0089] S900. Enter the transaction process after obtaining the confirmation of the purchase intention from the buyer user.

[0090] Please refer to Figure 7 As shown, in a more detailed implementation manner, each step executed by the seller, platform, and buyer corresponding to the quality inspection part, recommendation part, and transaction part of the system is given.

[0091] In the actual implementation of the above steps S200 to S430, the book pictures mainly include the book cover and inner pages. The system can first require the seller user to upload a photo of the book cover. The system uses an OCR hybrid recognition model to identify the book title, author, and ISBN code, and automatically matches the corresponding book classification. If it is inaccurate, the seller can make manual adjustments. The system extracts the ISBN code from the book picture and connects to the publisher's database or the metadata system of the National Library to obtain the authoritative electronic text and knowledge framework of the book, and constructs a cross-modal semantic association channel between the handwritten notes and the standard original text. This mechanism intelligently matches the handwritten note content recognized by OCR with the corresponding chapters of the book original text, verifies the logical relevance, innovation, and other quality dimensions of the notes and knowledge points through natural language processing, and finally converts the high-quality notes that conform to the disciplinary knowledge system into pricing gain factors to achieve the precise quantitative impact of content value on the selling price of second-hand books. Secondly, the system can require the seller user to upload at least 3 photos of inner pages (including handwritten notes), which need to cover the front, middle, and back parts of the book to prevent the seller user from only showing high-quality pages. If the photos are blurred or the notes are unclear, the system will prompt to retake. Finally, the seller user can be required to fill in relevant supplementary information.

[0092] In a more preferred implementation manner, in the above steps S410 to S430, during this process, a multi-source data complementary verification mechanism can be started: compare the note text recognized by OCR with the standard electronic text retrieved by ISBN to detect the integrity of the knowledge coverage of the user's notes, such as whether the annotations cover the core chapters; during the process of requiring the user to supplement information, the user can be required to fill in the degree of damage, and the degree of damage filled in by the user is cross-verified with the missing area of the page recognized by OCR. If there is a contradiction, manual review will be triggered; subject classification, such as "0 Mathematics, Physics, and Chemistry", restricts the scope of knowledge graph invocation to avoid cross-disciplinary misjudgment. After the verification passes, the classification label and verification log are synchronously stored in the knowledge base to provide a reliable data source for the subsequent pricing model.

[0093] The implementation of the above steps S601 to S630 enables the system to have the functions of pricing and confirming the price: the seller user can view the book quality inspection result and the final pricing result. The system supports the seller user to make a manual fine-tuning of up and down 10%. If the difference is too large, the seller user can restart the pricing process. If the pricing deviation is large for multiple times, an artificial review module can be designed for processing.

[0094] The implementation of the above step S800 enables the system to have the function of intelligent commodity recommendation. The buyer user fills in information such as the purchase purpose, required subject, price range, and book quality range. The system automatically generates a recommended commodity list according to the filled results. At the same time, an AI assistant can also be designed to further assist in the recommendation. Users with unclear needs can ask questions, such as asking for recommended books in a certain field, book sales rankings, etc., and the system will make further precise recommendations.

[0095] In a preferred embodiment, in the above step S330, the book text and the note text are specifically obtained through an OCR hybrid recognition model. First, regarding the architecture of the OCR hybrid recognition model, the deep learning model architecture adopted by the system in this embodiment is based on an improved CRNN-Transformer hybrid framework, and the hybrid recognition performance of handwritten notes (note text) and printed text (book text) is improved through a multi-level optimization strategy. Specifically, it is implemented through the following key modules (multi-modal feature decoupling encoder, dual-branch recognition network, and cross-modal association module) and steps:

[0096] S331. Feature extraction and decoupling: Extract general text features (underlying feature extraction) through a shared underlying convolutional network, and the feature pyramid network fuses multi-scale features; separate the printed and handwritten features through a gating mechanism, allocate them to different recognition branches, and dynamically allocate recognition weights. Here, the weights mainly refer to the feature contribution weights of different recognition branches, that is, printed and handwritten. The gating unit will generate two sets of weight coefficients to respectively control the weight ratio of printed and handwritten features in the final feature fusion. Its core role is to dynamically adjust the importance of the two types of features in the fusion process through the gating mechanism, solve the recognition problem when printed and handwritten texts appear in combination, and improve the adaptability of the model to the mixed text scenario. Specifically:

[0097] 1. Multi-modal feature decoupling encoder:

[0098] (1) Underlying feature extraction:

[0099] Adopt an improved ResNet-34 as the backbone network through a shared underlying convolutional network, and introduce a channel attention gating mechanism on the basis of the original residual block. Define the output of the "l"th layer residual block as:

[0100]

[0101] where is the residual function, σ is the Sigmoid activation function, W g ∈R (C×1) is a learnable parameter, and GAP represents global average pooling. This structure enables the model to share general text features at the underlying layer while dynamically suppressing noise channels.

[0102] (2) Feature pyramid:

[0103] Fuse multi-scale features through the feature pyramid network (FPN). Multi-scale refers to the size of the feature map, that is, the spatial resolution. Specifically, it means that feature maps at different levels have different spatial scaling ratios, 1 / 4, 1 / 8, 1 / 16. Define the generation process of the kth level feature map:

[0104] P k =Conv1×1 (C k ) + UpSample(P k+1 )

[0105] where C k is the output of the k-th layer of the backbone network, and bilinear interpolation is used for upsampling. Finally, pyramid features containing three scales of [1 / 4, 1 / 8, 1 / 16] are obtained. The purpose of fusing multi-scale features is to adapt to the diverse target size requirements, enabling the model to accurately detect targets of different sizes simultaneously.

[0106] S332. Dual-branch parallel recognition. Set up a printed text recognition branch and a handwritten text recognition branch to recognize printed text and handwritten text respectively. Further, the printed text recognition branch optimizes and generates a printed text sequence through a deformable convolutional network and CTC decoding with increased character spacing constraints; the handwritten text recognition branch generates a handwritten text sequence through adversarial generative data augmentation and an attention gating unit. The printed text recognition branch captures geometric invariant features in regular layouts through a deformable convolutional network, combines prior knowledge of character spacing to optimize text alignment accuracy, and generates a printed text sequence. The handwritten text recognition branch introduces adversarial generative data augmentation to simulate diverse writing styles and ink penetration effects, and strengthens the context association of cursive characters through an attention gating unit to generate a handwritten text sequence. Specifically as follows:

[0107] 2. Dual-branch recognition network

[0108] (1) Printed text recognition branch:

[0109] The deformable convolutional network extracts printed text character features and corrects curved printed text:

[0110] Offset learning is introduced in the 3x3 convolutional kernel. The offset of the i-th printed text character position is calculated as:

[0111] Δp i = W offset ·x(p i )

[0112] where W offset is the offset prediction network, and x(p i ) is the coordinate offset of the printed text character position p i on the input feature map. The local area of the feature map covered by the input 3×3 convolutional kernel, that is, the 3×3 pixel block scanned by the standard convolutional kernel, outputs the offset of the printed text character position (i.e., 2D coordinates). The character sequence of the printed text after correction of the character position offset is recognized through the deformable convolutional network. Experiments show that this method improves the recognition accuracy of curved text (physical curvature of printed text in the image) by 12.7%.

[0113] CTC Decoding Optimization

[0114] Add a character spacing constraint term to the CTC loss:

[0115]

[0116] Let \(L_{CTC}\) denote the Connectionist Temporal Classification loss, which is used to optimize the model training for text recognition tasks. Here, \(p(y|x)\) is the probability of the model predicting the character sequence \(y\) under the input character sequence \(x\) (the character sequence after correction by the character position offset), \(\lambda = 0.1\), is the sum of the mean squared errors of the center distances between adjacent characters, \(d\) t is the center distance between adjacent characters, \(\mu\) d is the average spacing of the training set, \(T\) is the length of the input sequence, \(t\) is the current time step index, and the traversal range is from 1 to \(T - 1\).

[0117] (2) Handwritten Recognition Branch:

[0118] The adversarial data augmentation network extracts handwritten text features:

[0119] By adopting conditional GAN and fusing explicit conditions such as character labels and implicit conditions such as stroke styles, precise control over the generation process of handwritten samples is achieved. The generative adversarial network that introduces external conditional information such as labels and text features to guide the generation process can generate targeted samples to solve the problem of insufficient OCR training data, and by introducing physical feature conditions such as ink concentration, the recognition ability of the model for low-quality scanned documents is improved. The loss function of the generator \(G\) is:

[0120]

[0121] where \(D\) is the discriminator used to judge whether the input sample is the real sample \(x\) real or the generated sample \(G(z|t)\), \(\lambda\) rec = 10 controls the weight of the reconstruction loss, denotes averaging over the distribution of the noise \(z\) to ensure the robustness of the generator to different noise inputs, \(\log(1 - D(G(z|y)))\) is the adversarial loss objective of the generator. The generator hopes that the discriminator considers the samples it generates as real samples, so it needs to minimize \(\log(1 - D(G(z|y)))\). \(z\sim p\) z The random noise vector \(z\) follows the prior distribution \(p\) z . \(\|G(z|y)-x\) real \|_1 measures the pixel-level difference between the generated sample \(G(z|y)\) and the real sample \(x\) real .

[0122] Attention Gating Unit:

[0123] Introduce a gating mechanism in the BiLSTM. Its input is the concatenated hidden state and the current feature, and the output is the weighted hidden state. This design significantly improves the robustness of cursive handwriting recognition. Update gate calculation:

[0124] g t = σ(W g [h t-1 , x t ),

[0125] g t is the update gate, which controls the fusion ratio of the input information x t at the current moment and the historical state h t-1 , and its value range is (0, 1). σ(W g [h t-1 , x t ) is the Sigmoid activation function, W g is the learnable weight matrix, [h t-1 , x t is the concatenation operation, which concatenates the previous hidden state h t-1 and the current input x t along the feature dimension to form the input of the gating signal. ⊙ is the element-wise multiplication, is the candidate hidden state, and (1 - g t ) is the complementary gating signal.

[0126] S333. Cross-modal association and structured output. Map the recognition results of printed text and handwritten text via a spatial coordinate system to establish cross-modal association, construct a tree-like semantic structure with the printed paragraph as the root node and the handwritten annotation as the child node and output it to obtain the book text and note text. Support the joint semantic parsing of printed and handwritten texts through cross-modal association.

[0127] 3. Cross-modal association module

[0128] (1) Spatial-semantic alignment

[0129] Input the character-level coordinates of the printed text and handwritten text obtained by the recognition of the printed text branch and handwritten text branch, that is, the extracted character positions and geometric information, generate the bounding box coordinates, and construct the printed-handwritten association matrix Generate the initial association strength by combining the cosine similarity.

[0130]

[0131] where p i and q j are the normalized coordinates of the printed character i and the handwritten character j, is the square of the Euclidean distance between two-character coordinates, σ is the Gaussian kernel parameter controlling distance attenuation, and θ i,j is the angle between the direction of the printed character i and the handwritten character j.

[0132] (2) Dynamic routing algorithm

[0133] An iterative routing mechanism is used to optimize the association matrix, and the t-th iteration update is as follows:

[0134]

[0135] where represents the association strength between capsule i and capsule j (i and j are the row and column index elements in the printed-handwritten association matrix ) in the t-th iteration. Each iteration gradually strengthens the effective cross-modal associations by accumulating the previous association strength . is an exponential function used to enhance the weight of high association strength. After 3 iterations, the algorithm converges to a stable state. The tabular image in the input scanned document and the corresponding text description are input, and the dynamic association matrix between table cells and text paragraphs is output.

[0136] 4. Multi-task joint optimization

[0137] The total loss function of the model combines four optimization objectives: (1) CTC loss to solve the problem of misalignment between the lengths of the input sequence and the target sequence, improve text readability, and reduce the phenomena of "crowded characters" or "broken characters". (2) Attention decoding loss to align the input features and the target output through the attention mechanism and improve the model's ability to capture key information. (3) GAN loss to improve the model's adaptability to complex scenarios through the generative adversarial network. (4) Orthogonal loss to enforce the orthogonality of the printed and handwritten feature spaces and prevent the confusion of the two types of features.

[0138]

[0139] where is the CTC plus loss weight, is the GAN plus loss weight, is the attention decoding loss, is the orthogonal loss weight, enforcing the orthogonality of the two types of text feature spaces. The weights are set as λ1 = 1, λ2 = 0.5, λ3 = 0.1, λ4 = 0.05.

[0140] In a preferred embodiment, in the above step S500, the note value coefficient is specifically obtained through the NLP note quality assessment model. First, regarding the architecture of the NLP note quality assessment model, at the natural language processing technology level, the system constructs a multi-level semantic understanding framework, and through in-depth knowledge reasoning and context awareness mechanisms, realizes the value association quantification between handwritten notes and the original text of the book. Based on the structured data output by the OCR hybrid recognition model - including the separated printed text paragraphs, the corresponding handwritten note (annotation) text content and its spatial coordinate mapping relationship, the system first uses the Sentence-BERT model to embed the two types of texts into the 768-dimensional semantic vector space respectively. In this process, through the domain adaptation fine-tuning strategy, specific subject corpora are used to optimize the semantic discrimination of the embedded vectors. The value evaluation of handwritten notes is mainly realized through three-stage cognitive calculations of semantic anchor alignment, logical gain analysis, and knowledge topology evaluation. The specific steps are as follows:

[0141] S501. Text input, the book text and the note text are input into the NLP note quality assessment model. The system uses the NLP note quality assessment model to perform multi-level comparisons between the handwritten note text content extracted by OCR and the original text (printed text), establishes a text-note mapping relationship based on the paragraph and title levels, detects the completeness of knowledge point supplementation through paragraph-level semantic similarity calculation, positively detects the coverage of the note content on the book knowledge points, and reversely verifies the key knowledge points in the book that are not extracted by the note.

[0142] S502. Semantic anchor alignment, using the improved dynamic time warping algorithm, combined with the semantic constraints of the knowledge graph and character coordinate information, locates the association anchors between the note text (handwritten text) and the book text (printed text), and outputs the semantic matching degree. In this embodiment, the system uses the improved dynamic time warping (DTW) algorithm to slide and match the handwritten annotation meaning units along the paragraph sequence of the printed original text. The handwritten annotation meaning unit is the text or symbol with specific semantic information marked by the user in the paper book. By introducing the semantic constraint term based on the knowledge graph, the distance calculation method of the traditional DTW is optimized: for example, when the handwritten content involves "Fourier transform properties", the algorithm preferentially searches for matching anchors in the paragraphs of the "integral transform" chapter in the original text, rather than mechanically traversing the whole book. At the same time, a spatio-temporal constraint matrix is constructed in combination with the character coordinate information to suppress unreasonable associations across pages, such as misassociating the annotation on page 5 to the content on page 50.

[0143] S503. Logical gain analysis: Invoke the subject-specific model to verify the logical rigor of the aligned note text content. For mathematics, verify the application conditions of theorems; for literature, verify the plot evidence. Output the logical gain value. In this embodiment, the system invokes the subject-specific logical verification model. Taking mathematics notes as an example, a symbolic reasoning engine based on the Coq theorem prover converts the handwritten derivation steps into formal language and automatically detects the rigor of links such as equation transformation and condition application. When it detects, for example, obtaining from "L'Hopital's rule", the model will verify whether the theorem application conditions (0 / 0 type indeterminate form, function differentiability) are clearly satisfied in the context and mark the reasoning jumps with unstated premises. For literary annotations, through knowledge graph embedding based on TransE, analyze whether assertions such as "contradictions in character development" in the handwritten comments have sufficient plot evidence support.

[0144] S504. Knowledge topology assessment: Map the note text after logical gain analysis to the corresponding subject knowledge graph, and quantify the contribution of the note text content to the corresponding subject knowledge graph through a graph attention network, and output the knowledge contribution value. In this embodiment, the system maps to the corresponding subject knowledge graph according to the Chinese Library Classification (such as chemical substance reaction network, historical event causal chain), and quantifies the contribution of the handwritten content to the knowledge system through a graph attention network (GAT). Define the knowledge node activation function:

[0145]

[0146] where α ij is the knowledge point activation function, exp(·) is the exponential function, LeakyReLU is the activation function, a T is the weight vector, h i is the knowledge point embedding vector of the original text (printed text), h j is the text content embedding vector of the handwritten note (annotation), h k is the original feature vector of the knowledge node k, Wh i 、Wh j and Wh k are the corresponding knowledge node weight matrices. When the handwritten note (annotation) text forms a new edge connection, such as supplementing the causal relationship between "penicillin discovery" and "treatment of World War II wounded", the system triggers the knowledge density gain factor, causing a non-linear increase in the note value coefficient. At the same time, detect the proportion of redundant content, and filter out low-value annotations through TF-IDF weighted similarity and edit distance threshold.

[0147] S505. Comprehensive score and value coefficient output. A dynamic weight allocation network is constructed through a gated recurrent unit. The dynamic weight allocation network dynamically adjusts the weights of semantic matching degree, logical gain value, and knowledge contribution degree according to the subject characteristics, and calculates the comprehensive score of the note through weighted calculation and outputs the note value coefficient after normalization. In this embodiment, the final value evaluation model integrates three indicators: semantic matching degree (weight 40%), logical gain value (35%), and knowledge contribution degree (25%). A dynamic weight allocation network is constructed through a gated recurrent unit (GRU), and the importance of the indicators is automatically adjusted according to the subject characteristics - for example, the weight of logical rigor is increased in legal notes, while the evaluation of innovation is strengthened in art reviews. The output result is normalized to a value coefficient in the range of 0-1 by the S-shaped function and mapped to the quality multiplier term in the pricing formula. This solution enables the resale price of high-quality notes (value coefficient > 0.8) to reach more than twice the basic price, while low-quality notes (< 0.4) can only obtain 50% of the benchmark price, accurately reflecting the market premium of knowledge added value.

[0148] In a preferred embodiment, in the above step S510, the damage coefficient of the book is obtained through a damage degree judgment model based on deep learning. First, regarding the architecture of the damage degree judgment model based on deep learning. In the optimization of the model structure for damage degree recognition, the core goal is to achieve high-precision quantitative evaluation of complex damage patterns by improving the balance between feature expression ability and computational efficiency. The damage degree judgment model based on deep learning in this embodiment adopts a lightweight dynamic perception network, a cross-layer feature pyramid, a dual-path attention mechanism, and a multi-granularity Transformer interactive encoder, and the following is a detailed description:

[0149] 1. Lightweight dynamic perception network

[0150] (1) Adaptive receptive field module

[0151] In the scenario of book damage recognition, the adaptive receptive field module can significantly enhance the model's ability to capture complex damage patterns. When the damaged areas in the book images uploaded by seller users may appear as subtle creases, local tears, or large-area stains, this module can dynamically adjust the perception granularity like a "smart magnifying glass": in the face of tiny cracks at the paper edges, it automatically shrinks the receptive field to focus on the microscopic texture of fiber fractures; when encountering the wavy deformation formed by page folds, it expands the perception range to capture the overall deformation trend; and when dealing with the blurred stains caused by ink penetration, it can, through the multi-scale feature competition mechanism, synchronously analyze the diffusion degree of the stain edges and the color concentration of the central area. This dynamic scale adaptation ability enables the model to not only identify millimeter-level subtle damage features but also evaluate the structural integrity of the entire page of paper, and under the limited computing resources of mobile devices, achieve a comprehensive judgment of the book damage degree from local to global, providing accurate quantitative basis for the evaluation of the priority of ancient book restoration or the grading of second-hand books. Let the input feature map be The dynamic receptive field adjustment is achieved through deformable convolution. The offset generation network outputs the offset field (K is the convolution kernel size), and the dynamic convolution kernel position correction formula is:

[0152]

[0153] where p k is the sampling point of the original convolution kernel, Δp k (p0) is the offset, and p0 is the current position. Feature resampling is achieved through bilinear interpolation:

[0154]

[0155] where Y(p0) is the value after feature resampling, w k is the convolution kernel weight, is the input feature map value, and k represents the sampling point index of the convolution kernel, which is used to dynamically adjust the sampling position of the deformable convolution. This module makes the convolution kernel adaptively fit the damaged edge (such as deforming along the crack direction) by dynamically adjusting the sampling grid, and compared with the fixed convolution kernel, it improves the feature capture ability under the same number of parameters.

[0156] (2) Dynamic channel pruning mechanism

[0157] The dynamic channel pruning mechanism acts as an "intelligent feature filter" in the damaged assessment of second-hand books. It can automatically adjust the computational intensity of the neural network according to the complexity of the damaged features in each book image. When the book cover uploaded by the user has only slight creases, the mechanism will intelligently close the channels unrelated to deep texture analysis, just like an experienced ancient book restorer quickly scanning the key areas, and only retain the core path for identifying superficial damages. When faced with complex damages such as multiple ink penetrations and broken binding threads in the inner pages, it will dynamically activate more feature channels, just like adjusting the magnification of a microscope to deeply analyze the fiber fracture levels and ink diffusion gradients. This adaptive computing mode based on image content can not only quickly lock the macroscopic features of the spine glue opening in low-resolution pictures taken by mobile phones, but also accurately quantify the reticular brittle fracture details caused by paper acidification in high-definition scanned pictures, enabling the model to always maintain the balance between the professionalism of the assessment results and the response speed when dealing with complex lighting and multi-angle shooting scenarios common in flea markets, providing the second-hand book trading platform with an automated product quality appraisal ability that takes into account both efficiency and accuracy.

[0158] It introduces a lightweight gating network Adopts Gumbel-Softmax approximation for discrete selection:

[0159]

[0160] Among them, is the output probability value, α c is the learnable channel importance parameter, the temperature coefficient T controls the degree of discretization, ∈~Gumbel(0,1) is the random noise sampled from the standard Gumbel distribution (mean 0, variance 1), used to introduce randomness and prevent the model from falling into local optima. After pruning, the feature channel is calculated as:

[0161]

[0162] Among them, is the pruned feature channel, which is weighted and retained for the original channel X c through the probability g c , C max is the maximum channel number limit, ensuring that the number of channels retained after pruning does not exceed the hardware or model capacity constraints. This mechanism can reduce a part of the channel computational amount while maintaining the accuracy.

[0163] 2. Cross-layer feature pyramid and dual-path attention mechanism

[0164] This module realizes the multi-scale depth fusion of damaged features by constructing a cross-layer dense connection structure. Specifically, the high-resolution features extracted by the shallow network (such as the fine scratch texture on the surface of wood products) are cascaded across layers with the semantic features captured by the deep network (such as the overall morphology of the plastic fracture surface). Through lateral connections and upsampling operations, a hybrid feature pyramid containing microscopic details and macroscopic structures is formed, improving the pixel-level localization accuracy of the model for damaged areas.

[0165] On this basis, a spatial-channel dual-path attention mechanism is designed to strengthen key regions

[0166] The spatial attention branch uses a learnable spatial filter to generate a heat map in the image plane. For example, for the worn area of a leather sofa, this branch will enhance the response value of the damaged central area while weakening the interference of the leather texture.

[0167] The channel attention branch, through global feature analysis, identifies the key color channels related to damage. When identifying the crack of a red plastic bottle, this branch will focus on strengthening the contrast difference between the red component in the RGB channel and the crack area. The outputs of the two branches are fused through adaptive weights, and finally enhanced features with both spatial sensitivity and channel discriminability are generated, improving the damage recognition accuracy of the model in complex backgrounds.

[0168] 3. Multi-granularity Transformer Interaction Encoder

[0169] The multi-granularity Transformer interaction encoder constructs a hierarchical visual understanding system in the detection of damaged second-hand books. When the model processes the page images uploaded by users, this encoder will simultaneously establish a microscopic perspective from the level of paper fibers to a macroscopic perspective of the overall book structure: in the area where the binding thread is broken, the fine-grained branch captures the worn details of the thread bifurcation like a high-power magnifying glass; the medium-grained branch outlines the three-dimensional contour of the book spine deformation like the light of a repair table lamp; and the coarse-grained branch views the overall warping degree of the cover like a spotlight in an exhibition hall. The feature streams at these different observation scales do not work in isolation, but through a cross-granularity attention dialogue mechanism, enabling the oxidation traces of yellowed paper to be associated with both the microscopic evidence of fiber embrittlement and the structural changes of loose binding at the same time. When the model evaluates the spread range of water stains, it can not only count the macroscopic proportion of polluted pixels but also distinguish the fiber swelling characteristics caused by ink penetration. This multi-perspective collaborative analysis method enables the algorithm to not only identify local damages such as label tearing commonly seen in second-hand book transactions but also judge the level of deformation of the entire volume caused by long-term compression, achieving hierarchical analysis capabilities comparable to professional paper cultural relic detectors with limited computing resources on mobile devices.

[0170] The encoder divides the feature map into N w ×N wWindow, and the local self-attention calculation within the window is as follows:

[0171]

[0172] Among them, Q, K, and V are the Query, Key, and Value matrices, respectively representing the interaction relationships between each position in the input sequence and other positions. Q is further explained as: the query vector encoding the local texture features of the book cover, T represents the window size, that is, the pixel region dimension covered by each window. is the scaling factor, is the relative position encoding matrix. The cross-window global attention is achieved through shifted window partitioning to establish the association between windows:

[0173]

[0174] Among them, Head1,..., Head h represents the output of the h-th attention head. Each head independently learns the damaged features from different perspectives, such as horizontal cracks and vertical stains. W O is the output weight matrix that maps the concatenated multi-head features to the final detection result and outputs the damage coefficient.

[0175] In the above step S601, a pricing result is generated based on the note value coefficient, the damage coefficient, and the original price of the book. A preferred implementation of giving the price is as follows:

[0176] (1) Pricing criteria for damage degree and note quality

[0177] According to the damaged book photos uploaded by the user, after evaluating the damage degree and note quality, they are divided into four damage levels, namely A, B, C, and D, according to the following criteria. The specific criteria are shown in Table 1. Among them, the 0CR recognition coefficient is an index quantifying the legibility of the handwriting through a deep learning model, ranging from 0 to 1, reflecting the physical readability of the handwritten notes in the book uploaded by the user. The 0CR recognition coefficient is closely related to the damage degree. For example, slight damage or stains will reduce the clarity of the handwriting, and this coefficient can be directly evaluated and generated by the 0CR hybrid recognition model.

[0178] Table 1: Book Pricing Table

[0179]

[0180] (2) Explanation of pricing details

[0181] The OCR recognition coefficient quantifies the legibility of the handwriting (in the range of 0 - 1) through a deep learning model, reflecting the physical readability of the notes: ≥0.9 means that more than 90% of the characters can be accurately converted, corresponding to Class A clarity; ≤0.5 indicates that more than half of the content cannot be recognized and is classified as Class D.

[0182] The damage coefficient of books is based on a deep learning-based damage degree judgment model. According to the physical damage degree and content integrity, it is divided into four grades. Class A corresponds to a damage coefficient ≥ 0.90, which is applicable to brand-new books with no creases on the cover and no stains on the inner pages; Class B has a damage coefficient of 0.70 - 0.89, which is applicable to books with slightly worn corners but still readable as a whole; Class C has a damage coefficient of 0.50 - 0.69, which is applicable to books with obvious creases on the cover and yellowed inner pages, etc.; Class D has a damage coefficient < 0.50, which is applicable to books with chaotic content logic, a large amount of plagiarism, or physical damage that cannot be repaired.

[0183] The note value coefficient evaluates the knowledge increment (in the range of 0 - 1) of the content through an NLP model, including logicality, innovation, and knowledge matching degree. Above 0.9, it needs to meet high-value features such as original illustrations, rigorous derivations, or interdisciplinary associations. Below 0.5, it is determined as an invalid note.

[0184] When the OCR recognition coefficient, note value coefficient, and damage coefficient all belong to the same grade, for example, OCR = 0.85 / Class B, value = 0.80 / Class B, damage = 0.75 / Class B, it is priced directly according to that grade. For example, Class B corresponds to 30% - 50%. If there is a cross-grade situation such as OCR = 0.92 / Class A, value = 0.68 / Class C, damage = 0.70 / Class B, the system automatically adopts the lowest grade, which is Class C, corresponding to 10% - 30%. In special scenarios, an additional correction factor can fine-tune the price: If a hand-drawn mind map is detected (it needs to simultaneously meet OCR ≥ 0.8, note value ≥ 0.75, and damage coefficient ≥ 0.7), the price can be increased by 5%. If stains cover key formulas resulting in an OCR local coefficient < 0.3, an additional 10% reduction is required. For notes on scarce disciplines (such as rare out-of-print textbooks), when the note value coefficient reaches Class A and the knowledge scarcity score is manually reviewed, the price can be increased to 150% of the original price, but it is necessary to simultaneously verify whether the OCR and damage coefficient meet the basic thresholds (OCR ≥ 0.6, damage ≤ 0.4). The final price needs to comprehensively consider the three coefficient grades, correction factors, and manual review results to ensure the balance between the value of scarce resources and the risk of physical damage.

[0185] Based on the above embodiments, the method and system for digital recycling and sharing of book notes based on OCR-NLP collaboration of the present invention can be positioned as a trading platform for enhancing the value of second-hand books based on artificial intelligence multimodal recognition technology. The main technological innovation lies in using OCR technology to accurately identify the damage degree and content value of old book notes in combination with multi-dimensional information input by users, using NLP technology to supplement and determine the content value, and finally generating a multi-dimensional comprehensive pricing for second-hand books. The system combines the existing integrated optical character recognition (OCR) to quickly identify and analyze the clarity and damage degree of handwritten notes, and then through intelligent recognition systems such as natural language processing (NLP) technology, efficiently recognize the handwritten content, accurately extract the key knowledge points in the notes, and at the same time judge the internal logic and fluency of the sentences and the text matching rate to judge the quality of the notes. At the same time, classification and structured arrangement are carried out. Users can submit idle books through a convenient online channel, and the platform will automatically complete the evaluation of the book condition, digital processing of the content, and knowledge tagging and archiving to form a searchable and sharable knowledge base. At the same time, an open knowledge circulation ecosystem is constructed, and the sorted books and notes are directionally matched to learners, researchers or educational institutions to maximize the resource utilization rate.

[0186] More specifically, the present invention has the following advantages:

[0187] (1) Users upload book images, and the system performs accurate recognition of mixed text. The core data source input is the digital image of the inner page of the book uploaded by the user, which contains both printed original text (book text) and handwritten annotations (note text). The system adopts an OCR architecture with deep feature decoupling, extracts general text features through a shared underlying convolutional network, and realizes parallel recognition channels for printed and handwritten texts at the high-order semantic level. In a preferred embodiment, image preprocessing is performed after the image is uploaded. The image preprocessing can integrate adaptive illumination correction and a U-Net-driven document repair module to perform pixel-level restoration for physical damages such as wrinkles, shadows, and perspective distortion, ensuring the robustness of feature extraction in low-quality shooting scenarios.

[0188] The OCR hybrid recognition model adopted dynamically allocates recognition weights through a gated feature separation mechanism. The printed text recognition branch (channel) uses a deformable convolutional network to capture geometric invariant features in regular layouts, and optimizes the text alignment accuracy by combining prior knowledge of character spacing; the handwritten text recognition branch (channel) introduces an adversarial generative data augmentation strategy to simulate diverse writing styles and ink penetration effects, and strengthens the context correlation modeling of cursive characters through an attention gated unit. The recognition results of the two types of texts establish cross-modal associations through spatial coordinate system mapping, and construct a tree-shaped semantic structure with printed paragraphs as the root nodes and handwritten annotations as the sub-nodes.

[0189] To address the complex interaction problems in the mixed-text scenario, the OCR mixed recognition model embeds a cross-modal attention mechanism, establishing a dynamic association between the printed query vector and the handwritten key-value pairs in the Transformer encoding layer to accurately locate the knowledge anchor points corresponding to handwritten notes. For example, it automatically binds the annotations for solving differential equations to the paragraphs in the "First-Order Linear Equations" chapter of the textbook and marks the reference relationships between the supplementary derivation steps and the original theorems. The system outputs structured text data, including character-level coordinate bounding boxes, region type labels, and cross-modal association confidence levels, providing fine-grained alignment information for subsequent value analysis. The optimized OCR solution improves the recognition accuracy of printed and handwritten characters in the mixed-text scenario, enhances the cross-modal association accuracy, and can automatically parse cross-reference logics such as "See Theorem 2.3" in handwritten notes, constructing a reliable data foundation for knowledge value quantification.

[0190] (2) Evaluate the value relevance of the knowledge systems of the handwritten note text content and the book text content.

[0191] The system establishes a cross-modal semantic alignment channel through the ISBN encoding to implement an accurate influence mechanism of note quality on book pricing. When the user uploads the book cover image, the system automatically locates and recognizes the ISBN encoding, connects to authoritative data sources such as the publisher's digital resource library and the National Library metadata system to obtain the standard electronic text, forming a benchmark knowledge framework for evaluating note value. During this process, the system uses the NLP analysis module to perform multi-level comparisons between the handwritten note content extracted by OCR and the original text: calculating the semantic similarity at the paragraph level to detect the completeness of knowledge point supplementation, using the NLP note quality assessment model to identify the innovation of the annotations, and verifying the correctness of the derivation process in combination with the subject knowledge graph. This is a comprehensive evaluation process, and its core indicators include both the logical gain value and the knowledge contribution degree.

[0192] If the handwritten note presents high-value content in core chapters, such as mathematical theorem proofs and historical event analyses - including but not limited to extended example analyses, interdisciplinary knowledge associations, and in-depth dissections of error-prone points, the system will automatically trigger a quality coefficient gain mechanism and introduce a non-linear adjustment factor into the basic pricing formula. For example, when it detects an original diagram of the supply and demand curve dynamic model in the notes of "Microeconomics", the quality coefficient can be increased, driving up the final selling price, while a simple knowledge point extraction note can only maintain the basic coefficient. This dynamic adjustment mechanism based on content value enables high-quality notes to break through the linear growth limit in the traditional pricing model and truly reflect their knowledge added value.

[0193] Finally, there is the structured data actively provided by users, including two dimensions: the physical state of the book and the subject attributes. The degree of damage is quantified through a three-level classification (no damage / slight damage / severe damage). Among them, slight damage is defined as page corners being folded or edges being worn but not affecting reading, while severe damage includes cases such as missing pages, large-area stains, or loose bindings. This classification needs to be cross-validated with the physical book photos uploaded. The subject classification strictly follows the 22 major category systems of the "Chinese Library Classification Law". For example, notes on the political theory for postgraduate entrance examinations are classified into category D (Politics, Law), and medical notes are classified into category R. This classification directly affects the setting of the subject benchmark price in the basic pricing. To improve the data entry efficiency, the system will pre-extract text keywords such as "calculus" and "gene mutation" during the OCR stage, recommend possible subject categories, and users can confirm or correct them through a drop-down menu or visual tags.

[0194] (3) The three types of input data form a complementary verification mechanism within the system. The comparison between the OCR text and the original text can detect the knowledge coverage of the user's notes. The degree of damage filled in by the user and the page missing areas identified by the OCR hybrid model corroborate each other, while the subject classification restricts the scope of the knowledge graph call. This multi-source data fusion mechanism effectively reduces the error impact of a single data source. For example, when the OCR results in missing text fragments due to scribbled handwriting, the integrity assessment model of knowledge points can be recalibrated through the key chapter information supplemented by the user, thus ensuring the robustness of the quality evaluation results.

[0195] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as these technical feature combinations do not conflict, they should be considered as the scope described in this specification. Moreover, the above embodiments only represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can be made, and these all belong to the protection scope of the present invention.

Claims

1. A method for digital recycling and sharing of book notes based on the collaboration of OCR and NLP, characterized in that Including: Construct an OCR hybrid recognition model and an NLP note quality assessment model; Obtain a book picture uploaded by a seller user, which contains book content and note content; Input the book picture into the OCR hybrid recognition model to obtain the OCR recognition coefficient, book text, note text, and book information including the book title and original price; Match the corresponding book category according to the book information, and incorporate it into the knowledge base after being tagged; Input the book text and note text into the NLP note quality assessment model to establish the value association between the note text and the book text, and obtain the note value coefficient; Generate a pricing result based on the OCR recognition coefficient, note value coefficient, and original price of the book; Attach the corresponding price to the book and put it on the shelf for display as a commodity.

2. The method for digitizing, recycling, and sharing book notes based on OCR-NLP collaboration according to claim 1, wherein The method further includes: constructing a damage degree judgment model based on deep learning; identifying the integrity of the book picture, which is output when the OCR hybrid recognition model identifies that the book picture is incomplete; inputting the incomplete book picture into the damage degree judgment model based on deep learning to obtain the damage coefficient; and the pricing result is also generated in combination with the damage coefficient.

3. The method for digital recycling and sharing of book notes based on OCR-NLP collaboration according to claim 1, characterized in that, The method further includes: recommending corresponding commodities according to the purchase intention of the buyer user; and entering the transaction process after obtaining the confirmation of the purchase intention of the buyer user.

4. The method for digitizing, recycling and sharing book notes based on OCR-NLP collaboration according to claim 1, characterized in that The method further includes: sending the matched corresponding book category to the seller user for confirmation or correction; when receiving the confirmation feedback from the seller user, adding the corresponding classification label; when receiving the correction feedback from the seller user, re-match the corresponding book category according to the corrected or supplemented book information of the seller user, and add the corresponding classification label.

5. The method for digital recycling and sharing of book notes based on OCR-NLP collaboration according to claim 1, wherein The method further includes: sending the pricing result to the seller user for confirmation; when receiving the feedback from the seller user confirming the pricing, attach the corresponding price to the book and put it on the shelf for display as a commodity; when receiving the feedback from the seller user having an objection to the pricing, provide a pricing adjustment range not exceeding the floating ratio threshold for the seller user to adjust the pricing, or re-price according to the seller user's selection, and let the seller user re-upload the book picture containing the note content to re-obtain the book picture for re-pricing.

6. The method for digitizing, recycling and sharing book notes based on OCR-NLP collaboration according to claim 1, wherein The method further includes: providing a book information supplementary input box to obtain the book information supplemented by the seller user, including the purchase time and damage degree.

7. The method for digital recycling and sharing of book notes based on OCR-NLP collaboration according to claim 1, characterized in that, When inputting the book picture into the OCR hybrid recognition model to obtain the book text, note text, and book information including the book title and original price of the book, the specific acquisition of the book text and note text is as follows: Feature extraction and decoupling, extracting general text features through a shared underlying convolutional network, and the feature pyramid network fusing multi-scale features; separating printed and handwritten features through a gating mechanism, allocating them to different recognition branches, and dynamically allocating recognition weights; Dual-branch parallel recognition, setting a printed text recognition branch and a handwritten text recognition branch to respectively recognize printed text and handwritten text; Cross-modal association and structured output, establishing cross-modal association between the recognition results of printed text and handwritten text through spatial coordinate system mapping, constructing a tree-like semantic structure with printed paragraphs as root nodes and handwritten annotations as child nodes and outputting, to obtain the book text and note text.

8. The method for digitizing, recycling and sharing book notes based on OCR-NLP collaboration according to claim 7, characterized in that, The printed text recognition branch generates a printed text sequence through deformable convolutional networks and CTC decoding optimization with increased character spacing constraints; the handwritten text recognition branch generates a handwritten text sequence through adversarial generative data augmentation and attention gating units.

9. The method for digital recycling and sharing of book notes based on OCR-NLP collaboration according to claim 1, wherein Inputting the book text and the note text into the NLP note quality assessment model to establish the value association between the note text and the book text and obtain the note value coefficient, specifically including the following: Text input: Input the book text and the note text into the NLP note quality assessment model. Semantic anchor alignment: Use the improved dynamic time warping algorithm, combined with the semantic constraints of the knowledge graph and character coordinate information, to locate the association anchors between the note text and the book text and output the semantic matching degree. Logical gain analysis: Call the subject-specific model to verify the logical rigor of the aligned note text content. For mathematics, verify the application conditions of the theorems; for literature, verify the plot evidence, and output the logical gain value. Knowledge topology evaluation: Map the note text after logical gain analysis to the corresponding subject knowledge graph, and quantify the contribution of the content of the note text to the corresponding subject knowledge graph through a graph attention network, and output the knowledge contribution value. Comprehensive scoring and value coefficient output: Construct a dynamic weight allocation network through a gated recurrent unit. The dynamic weight allocation network dynamically adjusts the weights of the semantic matching degree, logical gain value, and knowledge contribution degree according to the subject characteristics, and calculates the comprehensive note score through weighted calculation and normalizes it to output the note value coefficient.

10. A book note digital recycling and sharing system based on the collaboration of OCR and NLP, which is used to execute the method described in any one of claims 1 to 9 above, and is characterized in that, Including: A model construction module for constructing an OCR hybrid recognition model and an NLP note quality assessment model. An image acquisition module for acquiring a book image uploaded by the seller user and containing book content and note content. A text recognition module for inputting the book image into the OCR hybrid recognition model to obtain the OCR recognition coefficient, book text, note text, and book information including the book title and original price. A classification and induction module for matching the corresponding book classification according to the book information and incorporating it into the knowledge base after labeling. A note quality assessment module for inputting the book text and the note text into the NLP note quality assessment model to establish the value association between the note text and the book text and obtain the note value coefficient. A pricing module for generating a pricing result according to the OCR recognition coefficient, note value coefficient, and book original price. A listing module for listing the book with the corresponding price and displaying it as a commodity.

Citation Information

Cited By

  • Transform-based stroke trajectory prediction calligraphy evaluation method

    CN122157283A