Interactive processing methods and devices for physical books
By capturing and recognizing images of physical books, generating interactive pages for excerpts and annotations, and using server-side book mapping and retrieval, the problem of insufficient reading experience in existing technologies is solved, achieving efficient digitization of physical books and user-friendly interactive processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-26
AI Technical Summary
Existing digitization technologies for physical books have not yet effectively improved the reading experience, especially in terms of the lack of deep integration in image recognition and user interaction, making it difficult for users to perform efficient note-taking and annotation operations.
By capturing and recognizing images of pages from physical books, an interactive excerpt page is generated. The user's excerpt elements are obtained and uploaded to the server. The interactive comment page is rendered and comment data is collected. The server is used to perform book mapping retrieval to achieve related display of comments.
It enables efficient digitization of physical books, improves the user's experience in taking notes and annotating, and enhances the interactivity and comprehensibility of book content.
Smart Images

Figure CN122086285A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of image recognition technology, and in particular to an interactive processing method and apparatus for physical books. Background Technology
[0002] With the rapid development of image recognition technology, the digitization of physical books has become an important way to expand reading methods and improve reading efficiency. Among them, document scanning technology based on user terminals has been widely used, which can convert physical book pages into digital content. Users can read based on the digital book content. On this basis, with the continuous improvement of digitization level and users' expectations for reading experience, how to further improve the book reading experience has become the focus of attention for all parties. Summary of the Invention
[0003] This specification provides one or more embodiments of an interactive processing method for physical books, applied to an application. The method includes: capturing and recognizing images of pages from a physical book to obtain book data and generate an excerpt interaction page; acquiring the excerpt element corresponding to an excerpt command submitted by a user based on the excerpt interaction page and uploading it to a server; rendering the excerpt element to obtain a corresponding comment interaction page; collecting comment data input on the comment interaction page and uploading it to the server; and displaying the comment data in association with the book location information issued by the server; the book location information is obtained by performing a book mapping retrieval based on search data containing the excerpt element and the comment data.
[0004] This specification provides one or more embodiments of another interactive processing method for physical books, applied to a server. The method includes: receiving excerpt elements uploaded by an application; the excerpt elements being submitted based on an excerpt interaction page, which is generated based on book data obtained through image acquisition and image recognition of book pages from a physical book; acquiring comment data input by a user on a comment interaction page generated based on the excerpt elements; performing a book mapping retrieval based on search data containing the excerpt elements and the comment data to obtain book location information; and sending the book location information to the application for comment-related display.
[0005] This specification provides one or more embodiments of an interactive processing device for physical books, running on an application program. The device includes: a book data acquisition module configured to acquire and recognize images of pages from a physical book, obtain book data, and generate an excerpt interaction page based on the book data; a text excerpt module configured to acquire excerpt elements corresponding to excerpt commands submitted by a user based on the excerpt interaction page and upload them to a server; an annotation data submission module configured to render the excerpt elements and obtain corresponding annotation interaction pages, acquire annotation data input on the annotation interaction pages, and upload the annotation data; and an excerpt interface display module configured to display annotations associated with the book location information sent by the server and the annotation data; the book location information is obtained by book mapping retrieval based on at least one of the excerpt elements and the annotation data.
[0006] This specification provides one or more embodiments of another interactive processing device for physical books, running on a server. The device includes: an excerpt element receiving module configured to receive excerpt elements uploaded by an application; the excerpt elements are submitted based on an excerpt interaction page, which is generated based on book data obtained through image acquisition and image recognition of book pages of a physical book; a book mapping retrieval module configured to obtain comment data input by a user on a comment interaction page generated based on the excerpt elements, and perform a book mapping retrieval to obtain book location information based on retrieval data containing the excerpt elements and the comment data; and a book location information distribution module configured to distribute the book location information to the application for comment-related display.
[0007] This specification provides one or more embodiments of an interactive processing device for physical books, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: perform image acquisition and image recognition on pages of the physical book to obtain book data and generate an excerpt interactive page; acquire an excerpt element corresponding to an excerpt instruction submitted by a user based on the excerpt interactive page and upload it to a server; render the excerpt element to obtain a corresponding comment interactive page, acquire comment data input on the comment interactive page and upload it to the server; and display comment associations based on book location information issued by the server and the comment data; the book location information is obtained by performing book mapping retrieval based on retrieval data containing the excerpt element and the comment data.
[0008] This specification provides one or more embodiments of another interactive processing device for physical books, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: receive excerpt elements uploaded by an application; the excerpt elements are submitted based on an excerpt interaction page, which is generated based on book data obtained by image acquisition and image recognition of book pages of a physical book; acquire comment data input by a user on a comment interaction page generated based on the excerpt elements; perform a book mapping retrieval based on search data containing the excerpt elements and the comment data to obtain book location information; and send the book location information to the application for comment-related display.
[0009] This specification provides one or more embodiments of a computer-readable storage medium for storing computer-executable instructions. When executed, these instructions implement the following process: image acquisition and image recognition of pages from a physical book to obtain book data and generate an excerpt interaction page; obtaining the excerpt element corresponding to the excerpt instruction submitted by the user based on the excerpt interaction page and uploading it to a server; rendering the excerpt element to obtain a corresponding comment interaction page; collecting the comment data entered on the comment interaction page and uploading it to the server; and displaying the comment data in association with the book location information issued by the server, whereby the book location information is obtained through book mapping retrieval based on search data containing the excerpt element and the comment data.
[0010] This specification provides one or more embodiments of another computer-readable storage medium for storing computer-executable instructions, which, when executed, implement the following process: receiving excerpt elements uploaded by an application; the excerpt elements are submitted based on an excerpt interaction page, which is generated based on book data obtained by image acquisition and image recognition of book pages of a physical book; obtaining comment data input by a user on a comment interaction page generated based on the excerpt elements; performing a book mapping retrieval based on search data containing the excerpt elements and the comment data to obtain book location information; and sending the book location information to the application for comment-related display. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1A schematic diagram illustrating the implementation environment of an interactive processing method for a physical book provided in one or more embodiments of this specification; Figure 2 A flowchart illustrating an interactive processing method for physical books provided in one or more embodiments of this specification; Figure 3 A schematic diagram of the interactive interface of a first application provided for one or more embodiments of this specification; Figure 4 A schematic diagram of the interactive interface of a second application provided in one or more embodiments of this specification; Figure 5 A schematic diagram of the interactive interface of a third application provided in one or more embodiments of this specification; Figure 6 A schematic diagram of the interactive interface of a fourth application provided in one or more embodiments of this specification; Figure 7 A schematic diagram of the interactive interface of a fifth application provided in one or more embodiments of this specification; Figure 8 A flowchart illustrating an interactive processing method for physical books applied to a physical book excerpting and commenting scenario, provided for one or more embodiments of this specification; Figure 9 A flowchart illustrating another interactive processing method for physical books provided in one or more embodiments of this specification; Figure 10 A schematic diagram illustrating an embodiment of an interactive processing device for a physical book, provided in one or more embodiments of this specification. Figure 11 A schematic diagram of another embodiment of an interactive processing device for a physical book provided in one or more embodiments of this specification; Figure 12 A schematic diagram of the structure of an interactive processing device for a physical book provided in one or more embodiments of this specification; Figure 13 This is a schematic diagram of the structure of another interactive processing device for a physical book provided in one or more embodiments of this specification. Detailed Implementation
[0012] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0013] The interactive processing method for physical books provided in one or more embodiments of this specification is applicable to the implementation environment of interactive processing of physical books. (Refer to...) Figure 1 The implementation environment includes at least: Application 101, Server 102, Retrieval Model 103; The application 101 is used to capture and recognize images of pages from physical books to obtain book data and generate an excerpt interaction page based on the book data. It also acquires the excerpt elements submitted by the user on the excerpt interaction page and uploads them to the server 102. The application renders the excerpt elements to obtain an annotation interaction page, collects the annotation data entered by the user on the annotation interaction page and uploads it to the server 102, and receives the book location information sent by the server and displays it in association with the annotation data. The application 101 can run on a user terminal, which can be a mobile phone, personal computer, tablet computer, e-book reader, wearable device, device for information interaction based on AR (Augmented Reality) / VR (Virtual Reality), and laptop computer, etc.
[0014] The server 102 is used to receive the excerpted elements and annotation data uploaded by the application 101, perform book mapping retrieval based on the retrieval data containing these data and obtain book location information, and return the book location information to the application 101; the server 102 can run on one or more servers, a server cluster composed of several servers, or a cloud server of a cloud computing platform.
[0015] The retrieval model 103 is deployed on the server 102. It performs book mapping retrieval based on the excerpt elements and commentary data input by the server 102 and obtains book location information. The book location information is then output to the server 102 and returned to the application 101.
[0016] In this implementation environment, during the interaction between the user and the physical book, application 101 first performs image acquisition and recognition on the book page to obtain book data and generate an excerpt interaction page. Then, the excerpt elements submitted by the user on the excerpt interaction page are uploaded to server 102. Subsequently, the excerpt elements are rendered to obtain an annotation interaction page. The annotation data input by the user is collected and uploaded to server 102. Server 102 inputs the received retrieval data containing excerpt elements and annotation data into retrieval model 103 to perform book mapping retrieval, obtain book location information, and send it to application 101. Finally, application 101 displays the annotation association based on the book location information and annotation data, thereby realizing annotation and book mapping retrieval based on the physical book's page.
[0017] This specification provides one or more embodiments of an interactive processing method for physical books, as follows: Reference Figure 2 The interactive processing method for physical books provided in this embodiment can be applied to applications. The method specifically includes steps S202 to S208.
[0018] Step S202: Image acquisition and image recognition are performed on the pages of the physical book to obtain book data and generate an interactive excerpt page.
[0019] The book page mentioned in this embodiment refers to a page in a physical book. Specifically, a book page can be a table of contents page, a main text page, a reference page, or an appendix page. The function of a book page is to carry knowledge content for readers to read and cite. This is different from a book label page, which serves as physical packaging and overall identification. The book label page can be the cover page.
[0020] In practice, to convert the pages of a physical book into book data that can be used for interactive processing, the book pages are first captured to obtain book page images, and then image recognition is performed on the book page images to obtain book data. The book data is the data obtained after capturing and recognizing the book pages, and is used to generate an interactive excerpt page. Specifically, the book data can be book content data composed of the reading content of the book pages, such as text data, chart data, and / or formula data in the book pages, or book structure data composed of the layout information of the book pages, such as text and image layout relationship data and paragraph division data. Based on this, after obtaining the book data, an excerpt interaction page is generated. The excerpt interaction page is an interactive interface generated based on the book data, used by users to select and submit excerpt elements. The excerpt interaction page includes a display area for displaying the excerptable content of the book pages and is configured with interactive controls for users to select and submit excerpt elements.
[0021] For example, Figure 3 The interactive page shown includes an image upload control 301 and a capture control 302. Users can interact with the image upload control 301 to select images; specifically, images can be uploaded by clicking or dragging. Supported image formats include JPG and PNG, and the application can automatically recognize text. Users can also interact with the capture control 302 to activate the camera and directly capture images of book pages for recognition. After the user completes image upload or capture, the user terminal displays... Figure 4 The extracted element recognition page shown displays the original image at the top and a recognition progress bar at the bottom before the recognition is complete.
[0022] In practical applications, different acquisition methods may be used when capturing images of book pages. For example, a user may take a horizontal or vertical shot of the book page depending on the posture of holding the user's terminal, or they may choose to capture the entire page or only a local area from the perspective of the book page content. Different acquisition methods will lead to differences in the quality of the captured book page images. To improve the quality and accuracy of image acquisition, this embodiment provides an optional implementation method for image acquisition and image recognition of physical book pages, including: The framing features of the book page images are extracted, and the image acquisition mode is determined based on the obtained framing features; Configure the parameters according to the acquisition constraint parameters corresponding to the image acquisition mode, acquire book page images, and perform image recognition on the book page images to obtain book data.
[0023] Among them, the book page viewfinder image refers to the preview image provided by the application before image acquisition, used to analyze the current acquisition method; the viewfinder features refer to the features extracted from the book page viewfinder image to determine the current acquisition method, including shooting direction, integrity of the book page, ambient light intensity and / or distortion, and there is a one-to-one mapping relationship between viewfinder features and image acquisition mode; the acquisition constraint parameters refer to the acquisition parameters configured for the image acquisition component based on different image acquisition modes. Specifically, the acquisition constraint parameters can be optical parameters associated with the acquisition method, image processing parameters, or output format parameters.
[0024] Specifically, in the process of extracting framing features from the view image of a book page, a framing feature extraction network can be used. This network can be a lightweight convolutional neural network (CNN) that integrates the DOC (Document Orientation Correction) algorithm and the ROID (Region of Interest Detection) algorithm. By inputting the view image of the book page into the network, the DOC algorithm is used to analyze the main orientation of the image to correct the screen rotation caused by the user's grip posture and output the orientation features. Then, the ROID algorithm is used to detect the boundary regions based on the corrected view image and output the region features. Finally, the orientation features and region features are fused to obtain the framing features.
[0025] After obtaining the above-mentioned framing features, the corresponding image acquisition mode is determined based on the framing features, and the image acquisition component is configured according to the corresponding acquisition constraint parameters. The image acquisition component is then used to acquire images of book pages, and image recognition is performed on the book page images to obtain book data. For example, when the framing features represent vertical and local areas, the image acquisition mode is determined to be a local vertical mode. The optical parameters of the image acquisition component are adjusted according to the preset acquisition constraint parameters of this mode, and book page images are acquired and recognized.
[0026] In the specific execution process, to improve the reliability of book page images and thus the accuracy of subsequent image recognition, after configuring parameters according to the acquisition constraint parameters corresponding to the image acquisition mode and acquiring book page images, and before performing image recognition to obtain book data, image parameter detection and calculation can be performed on the book page images to determine whether the book page images meet the image parameter conditions. In an optional implementation provided in this embodiment, it further includes: Image parameter detection and calculation are performed on book page images to obtain image parameter evaluation values; If the image parameter evaluation value meets the corresponding image parameter conditions, perform image recognition on the book page image to obtain book data; if the image parameter evaluation value does not meet the corresponding image parameter conditions, generate the corresponding guidance instruction and perform secondary image acquisition on the book page.
[0027] The image parameter evaluation value refers to the image quality score obtained after evaluating the acquired book page images from multiple dimensions. Specifically, these multiple dimensions include sharpness, integrity, illumination intensity, and / or image geometry. The image parameter conditions refer to the preset image quality thresholds for the above multiple dimensions, which are the basis for determining whether a second image acquisition is needed. For example, the image parameter conditions could be that the sharpness evaluation value is greater than x and the illumination intensity evaluation value is within the threshold range. Only when the image parameter evaluation value meets the image parameter conditions will the operation of image recognition to obtain book data be performed on the book page images. In the specific execution process, image parameter detection and calculation are performed on the book page images to obtain image parameter evaluation values. Image parameter detection and calculation can be performed through an image quality assessment model. The image quality assessment network can be a lightweight convolutional neural network that integrates the NR-IQA (Comprehensive No-Reference Image Quality Assessment) algorithm. By inputting the book page images into the image quality assessment network, image parameter detection and calculation are performed to obtain image parameter evaluation values.
[0028] Furthermore, the image parameter evaluation value is judged. If the image parameter evaluation value meets the corresponding image parameter conditions, image recognition is performed on the book page image to obtain book data. If the corresponding image parameter conditions are not met, a corresponding guidance instruction is generated and a second image acquisition of the book page is performed. The guidance instruction refers to the multimodal operation instruction information generated by the system and displayed to the user when the image parameter evaluation value does not meet the standard, in order to guide the user to complete the second image acquisition. Specifically, the guidance instruction includes text guidance information, voice guidance information and / or image guidance information. For example, if the book page image is overexposed, the guidance instruction "The image is overexposed, please avoid direct sunlight" is generated.
[0029] In practice, in order to enhance the user's interactive experience with physical books and enable book data to be overlaid on the book page images in an interactive form to generate an excerpt interactive page, thereby achieving accurate excerpting of the content of physical books by users, the excerpt interactive page can be generated based on the book page images and book data in the upper layer of the interactive page. Specifically, in one optional implementation of this embodiment, the extract interaction page is generated in the following manner: The application displays book page images on its interactive page and establishes a coordinate mapping relationship between the book data and the corresponding image element regions in the book page images. Based on the coordinate mapping relationship, the book data is rendered in the upper layer of the interactive page to obtain the excerpt interactive page.
[0030] The coordinate mapping relationship refers to the positional relationship between the book data and the corresponding image element region in the book page image. Specifically, the coordinate mapping relationship is obtained by calculating the positional association between the book data and the corresponding image element in the book page image, so that any data in the book data can be mapped to the corresponding element region in the book page image, providing a basis for the subsequent generation of interactive excerpt pages.
[0031] In the specific execution process, when rendering the book data in the upper layer of the interactive page, the book data is drawn to the position corresponding to the image elements in the book page image based on the above coordinate mapping relationship, and the book data is made interactive, thereby forming an interactive page that is visually aligned with the book page image and supports direct user interaction, providing a foundation for users to make excerpts later.
[0032] It should be noted that, considering that the physical book images, user comments, and other related data involved in this specification may, to some extent, constitute user privacy, authorization from the user must be obtained before collecting such data to ensure that the data collection operation complies with relevant data management regulations. For example, user authorization can be granted during user registration or when collecting images. Specific methods of data authorization could include sending a data authorization reminder to the user, who can then confirm the reminder with an instruction to obtain data authorization; alternatively, data authorization could be obtained by signing a data authorization agreement to obtain authorization for data collection or data transmission. This embodiment does not limit the scope of the authorization.
[0033] Step S204: Obtain the extract element corresponding to the extract instruction submitted by the user based on the extract interaction page and upload it to the server.
[0034] In practice, based on the generated excerpt interaction page, the system first obtains the excerpt instructions submitted by the user on the page. Based on these instructions, it extracts or selects corresponding excerpt elements from the book data. After obtaining the excerpt elements, it uploads them to the server. The server then receives the uploaded excerpt elements from the application, providing a data foundation for subsequent book mapping and retrieval. The excerpt instructions refer to the user's interactive commands on the excerpt interaction page, used to select book data. Specifically, these commands can be executed by the user clicking through the book data on the page, or by using a selection box to select continuous or non-contiguous areas of book data. The excerpt elements refer to the corresponding book data extracted according to the aforementioned excerpt instructions.
[0035] In addition, after collecting the comment data entered on the comment interaction page, the excerpted elements can be uploaded together with the comment data. In this case, the step of obtaining the excerpted elements corresponding to the excerpted instructions submitted by the user based on the excerpt interaction page and uploading them to the server can be replaced by: obtaining the excerpted elements corresponding to the excerpted instructions submitted by the user based on the excerpt interaction page. Accordingly, the subsequent steps of rendering the excerpted elements and obtaining the corresponding comment interaction page, collecting the comment data entered on the comment interaction page and uploading it to the server can be replaced by: rendering the excerpted elements and obtaining the corresponding comment interaction page, collecting the comment data entered on the comment interaction page, and uploading the comment data and the excerpted elements to the server. Alternatively, the book page image obtained through image acquisition and image recognition can be uploaded to the server along with the excerpt element. In this case, the step of obtaining the excerpt element corresponding to the excerpt instruction submitted by the user based on the excerpt interaction page and uploading it to the server can be replaced by: obtaining the excerpt element corresponding to the excerpt instruction submitted by the user based on the excerpt interaction page, and uploading the excerpt element and the book page image to the server.
[0036] In the specific execution process, different interaction methods may be used when users interact with the excerpt interaction page. For example, some users prefer to interact with the excerpt interaction page by clicking, while others prefer to use a selection box. In this case, in order to intuitively display the user's interaction results and thus improve the user's interaction experience and the accuracy of the excerpt, the excerpt elements and the corresponding excerpt image element areas can be synchronously marked. In an optional implementation method provided in this embodiment, obtaining the excerpt element corresponding to the excerpt instruction submitted by the user based on the excerpt interaction page includes: Extract the extract elements from the book data based on at least one extract range corresponding to the extract instruction; Based on the coordinate mapping relationship, the corresponding image element region of the excerpt element in the book page image is determined, and the excerpt element and the image element region are marked synchronously.
[0037] Synchronous marking refers to marking the corresponding areas of the book page image, i.e., the areas of the extracted image elements, according to the coordinate mapping relationship while marking the extracted excerpt elements. This allows the user's interactive operations on the excerpt page to be intuitively mapped to the book page image. Specifically, synchronous marking can take the form of visual synchronous marking, such as highlighting elements or highlighting areas of elements.
[0038] For example, Figure 5The interactive page shown includes an excerpt element selection interface 501 and an excerpt element display interface 502. Users can select or confirm excerpt elements in the excerpt element selection interface 501. After the selection of excerpt elements is completed, the selected excerpt elements will be displayed in the excerpt element display interface 502.
[0039] Step S206: Render the extracted elements and obtain the corresponding comment interaction page, collect the comment data entered on the comment interaction page and upload it to the server.
[0040] The comment interaction page described in this embodiment refers to an interactive interface generated based on the aforementioned excerpt elements for inputting comment data. Specifically, the comment interaction page includes an excerpt display area that displays the excerpt elements and is configured with interactive controls that support multimodal input and comment submission controls to input and submit comment data for the excerpt elements. Among them, the interactive controls can be text boxes that support text input, recording controls that support voice input, or interactive controls that support image upload. Correspondingly, the comment data refers to the data created by the user in the comment interaction page through the interactive controls for commenting on or annotating the excerpt elements. The comment data includes text comment data, voice comment data, and / or image comment data.
[0041] In practice, after obtaining the extracted elements, the extracted elements are rendered and a comment interaction page is generated. Users input comment data through the interactive controls on the comment interaction page and submit the comment data through the comment submission control, which can be uploaded to the server. Correspondingly, the server receives the comment data submitted by the application, providing the basis for subsequent comment association display and book mapping retrieval.
[0042] For example, Figure 6 The interactive page shown can display the user's... Figure 5 The selected excerpt elements, along with the corresponding book information and matching degree, are displayed on the interactive page. The page may also include an annotation data input interface 601 and a voice annotation input control 602. Users can input annotations on the annotation data input interface 601 and interact with the voice annotation input control 602 to input voice annotations.
[0043] Alternatively, instead of uploading the annotation data after obtaining it, the annotation data can be stored in the application and a annotation request indicating that the user's annotation is complete can be sent to the server. Correspondingly, after receiving the annotation request, the server performs a book mapping search based on the excerpt elements. In this case, the steps of rendering the excerpt elements and obtaining the corresponding annotation interaction page, collecting the annotation data entered on the annotation interaction page, and uploading it to the server can be replaced with: rendering the excerpt elements and obtaining the corresponding annotation interaction page, collecting the annotation data entered on the annotation interaction page, and sending an annotation request corresponding to the annotation data to the server.
[0044] In practical applications, user comments on extracted elements often contain the user's emotional inclinations and cognitive intentions. For example, a text comment saying "very good" reflects the user's approval, while a hurried voice comment may imply the user's confusion. In this case, to better understand the user's input comments, recommended resources can be determined based on the comments. Specifically, sentiment tag recognition can be performed on the comments to obtain sentiment tags and determine recommended resources. In addition, the comments can be uploaded to the server, where sentiment tag recognition can be performed to obtain sentiment tags, determine recommended resources, and then distribute the recommended resources to the application.
[0045] Specifically, in one optional implementation of this embodiment, after the steps of collecting the comment data entered on the comment interaction page and uploading it to the server are performed, the method further includes: The annotation data is encoded to obtain annotation encoded data, and the annotation emotional features are extracted from the annotation encoded data to obtain annotation emotional features; The emotional features of the comments are mapped to emotional tags to obtain emotional tags, and recommended resources are determined based on the emotional tags.
[0046] In the specific execution process, the above-mentioned sentiment tag mapping and recommendation resource determination can be performed through the sentiment mapping model. Specifically, the annotation encoder in the sentiment mapping model can encode the annotation data. The annotation encoder can be a Transformer encoder that integrates the Attention mechanism. By inputting the annotation data into the Transformer encoder, the encoder uses the Attention mechanism to analyze the global dependency of the annotation data and understands and encodes the deep semantics and sentiment tendency of the annotation through the knowledge obtained by pre-training, thus obtaining the annotation encoded data. After obtaining the annotation coding data, the emotional features of the annotation coding data are extracted to obtain the emotional features of the annotation. In the above emotional feature extraction process, the emotional features can be extracted through the emotional feature extraction module. The emotional feature extraction module can be a multilayer perceptron (MLP) based on fully connected layers. By inputting the annotation coding data into the multilayer perceptron, the fully connected layers inside the multilayer perceptron perform multi-layer nonlinear transformations on the annotation coding data, gradually abstracting deep features that are highly related to the emotional judgment from the original annotation coding data, and generating the emotional features of the annotation based on the deep features. Furthermore, in the process of obtaining sentiment tags by mapping sentiment features to sentiment labels, sentiment label mapping can be performed through a sentiment mapping network. This sentiment mapping network can be a lightweight convolutional neural network integrating a Softmax output layer. By inputting the sentiment features of the comments into the sentiment mapping network, the feature transformation layer first integrates and adjusts the dimensions of the sentiment features. The Softmax output layer normalizes the transformed sentiment features, calculates and outputs a set of probability distributions corresponding to preset sentiment labels. Finally, the system selects the label with the highest probability value as the sentiment label that accurately summarizes the sentiment tendency of the comment. Based on this sentiment label, the system recommends corresponding resources to the user, such as related articles and videos. For example, if user A comments a passage in a serious literary work as "well written," after model analysis, the sentiment label of this comment is determined to be "highly appreciated." Based on this sentiment label and the serious literary genre, the system matches and recommends a set of literary works and literary criticism articles of the same style to the user, thereby evoking emotional resonance and further expanding the user's reading comprehension.
[0047] Optionally, the annotation data can be multimodal annotation data, such as including voice annotation data and image annotation data. In this case, in order to more comprehensively and accurately understand the user's multimodal emotions, the above-mentioned step of encoding the annotation data to obtain annotation encoded data and extracting emotional features from the annotation encoded data to obtain annotation emotional features can be replaced by: decoupling the multimodal annotation data to obtain each modality's annotation data, and encoding and extracting emotional features from each modality's annotation data to obtain the corresponding emotional features of each modality. Correspondingly, the step of mapping sentiment features to sentiment labels to obtain sentiment labels and determining recommended resources based on sentiment labels can be replaced by: mapping sentiment features to sentiment labels to obtain sentiment labels for each modality, performing weighted fusion of sentiment labels for each modality based on preset modality sentiment weights, and determining recommended resources based on the weighted sentiment labels obtained from the weighted fusion.
[0048] Step S208: Display the annotations in association with the book location information and annotation data sent by the server.
[0049] In this embodiment, book mapping retrieval refers to the process of starting from the pages of a physical book and performing an online search based on search data. Specifically, book mapping retrieval can be searching for the electronic version of a physical book, or searching for books associated with a physical book, such as different versions of books with the same book identifier, or searching for other books referenced on the pages of a physical book. Book location information refers to the book information obtained through book mapping retrieval, and / or the specific element location information corresponding to the current excerpt element, such as the page number, chapter number, and / or page position of the excerpt element. Optionally, book location information is obtained by performing book mapping retrieval based on search data containing excerpt elements and annotation data.
[0050] In practice, after receiving the book location information from the server, in order to intuitively associate the user-submitted annotation data and excerpts with the corresponding areas in the physical book, thereby improving the comprehensibility of the annotation data and the efficiency of knowledge review, the annotation association display can be completed based on the book location information and the annotation data.
[0051] In practical applications, to identify the intent and contextual semantics of excerpted elements in user-submitted comment data, thereby improving the accuracy of book mapping retrieval, book mapping retrieval can be performed based on retrieval data containing excerpted elements and comment data. Specifically, in one optional implementation of this embodiment, book mapping retrieval based on retrieval data containing excerpted elements and comment data includes: Search features are extracted from the excerpted elements and commentary data to obtain excerpt search features and commentary search features, and semantic features are extracted from the excerpted elements to obtain semantic features; Book location information is obtained by performing book mapping retrieval in the candidate book database based on semantic features, commentary retrieval features, and excerpt retrieval features.
[0052] Specifically, a semantic feature extraction network is used to perform semantic parsing on the excerpted elements and extract semantic features. At the same time, a retrieval feature extraction network is used to process the excerpted elements and commentary data respectively to obtain excerpt retrieval features and commentary retrieval features that represent their retrieval intent. Based on the semantic features, commentary retrieval features, and excerpt retrieval features, book mapping retrieval is performed in the candidate book database to obtain book location information. Among them, the semantic feature extraction network can be a large language model based on the BERT (Bidirectional Encoder Representations from Transformers) architecture, which identifies the contextual semantic information of the excerpted elements and outputs semantic features through its fully connected layers and attention mechanism.
[0053] Furthermore, in the above-mentioned book mapping retrieval process, in order to comprehensively evaluate the contribution of different features to the retrieval results and improve the robustness and accuracy of book mapping retrieval, book mapping retrieval can also be performed by configuring reliability weights and merging rankings. In one optional implementation method provided in this embodiment, book mapping retrieval is performed in the candidate book database based on semantic features, commentary retrieval features, and excerpt retrieval features to obtain book location information, including: Input semantic features, commentary retrieval features, and summary retrieval features into the retrieval model and obtain their respective retrieval result sets; Assign corresponding confidence weights to each search result set, and perform weighted fusion and sorting of the search result sets based on the confidence weights to obtain a mapping search list; determine the book location information based on the mapping search list.
[0054] The retrieval result set refers to the preliminary set of book location information obtained after using any one of the semantic features, commentary retrieval features, or excerpt retrieval features as query input and searching candidate books through a retrieval model. Specifically, each retrieval result set reflects the retrieval tendency from a single dimension. For example, a retrieval result set based on semantic features may be more biased towards chapters related to concepts, while a retrieval result set based on excerpt retrieval features may be more focused on paragraphs containing specific content. The mapping retrieval list refers to the list of results generated by weighting and merging multiple retrieval result sets according to their corresponding confidence weights and then reordering them, arranged in descending order of relevance scores. Specifically, the mapping retrieval list integrates retrieval information obtained from searches based on different features, thereby providing a more accurate and comprehensive candidate sequence of book location information.
[0055] Semantic features, commentary retrieval features, and excerpt retrieval features are each input as independent queries into the retrieval model. The retrieval model can be a dense retrieval model based on feature similarity calculation. For each input feature, the retrieval model performs a retrieval and returns a set of retrieval results sorted in descending order of similarity score. Then, a preset confidence weight is assigned to the obtained retrieval result set, and the retrieval result set is weighted, fused, and sorted based on the confidence weight to obtain a mapping retrieval list. Finally, the book location information is determined based on the generated retrieval mapping list.
[0056] In practical applications, due to the potentially large number of candidate books or the ambiguity of user excerpts and annotations, several results in the book-mapping retrieval list may have similarities. In such cases, to improve user satisfaction and retrieval accuracy, multiple candidate book location information can be displayed, and the book location information can be determined based on the user's selection. In one optional implementation of this embodiment, annotation association display is performed based on the book location information and annotation data sent by the server, including: If the server detects that multiple candidate book location information has been sent, the location information of each candidate book will be displayed on the interactive page of the application; the corresponding candidate book location information will be determined based on the selection command submitted by the user and displayed in conjunction with the annotation data.
[0057] In practice, when the server distributes location information for multiple candidate books, the application displays the location information for each candidate book on the interactive interface. Users browse and compare these candidate book location information and submit selection commands to determine the book location information. Then, the information is displayed in association with the annotation data. The annotation association display refers to binding the user-submitted annotation data with the book location information and displaying it uniformly on the interactive interface. This allows the user's annotation to be clearly associated with the specific source location in the physical book, thereby establishing a connection between digital annotations and physical books. For example, a label such as "Source: Das Kapital, page 129" can be displayed next to an annotation. Clicking this label allows users to further preview the context or jump to the electronic version page.
[0058] Furthermore, depending on the actual application scenario or user needs, the objects of the above-mentioned associated display are not limited to annotation data. Book location information can also be associated with excerpt elements. In this case, the step of displaying annotations based on the book location information and annotation data sent by the server can be replaced with displaying excerpts based on the book location information and excerpt elements sent by the server.
[0059] Furthermore, in practical applications, the same physical book may have multiple different published versions, and the text content of these versions is often very similar. In this case, it may be difficult to locate the physical book by performing book mapping retrieval based on the user's excerpts and annotations. To improve the accuracy of book mapping retrieval, book page images obtained through image acquisition and image recognition can also be uploaded to the server. The uploading of book page images can be performed before or after the operation of uploading excerpts, but before the book mapping retrieval.
[0060] After obtaining the book page images, a book mapping retrieval can be performed based on the retrieval data containing the book page images and excerpt elements. Specifically, in one optional implementation method provided in this embodiment, the following is included: Upload images of book pages obtained through image acquisition and image recognition to the server; Topological and structural features are extracted from book page images to obtain character topological features and page structural features; Based on page structure features, the associated pages of candidate books are determined. Feature matching is performed between character topology features, page structure features, and key element features of excerpted elements and the associated page features corresponding to the associated pages to obtain book location information.
[0061] Among them, character topological features refer to the geometric features extracted from book page images that describe the shape, size, and relative position of characters; page structure features refer to the features that characterize the layout and / or format of book pages. Related pages refer to pages selected from candidate books based on the similarity of page structure features. Specifically, related pages can be pages with the same page number as the page image of a book. By locating the related page with the corresponding page number in the candidate books based on the page number corresponding to the page image of the book, the retrieval efficiency can be improved.
[0062] Key element features refer to the features extracted from the excerpt elements that can represent the key content of the excerpt elements. Specifically, in the process of extracting key element features, key element features can be extracted through a key feature extraction network. The key feature extraction network can be a CNN integrating an attention mechanism. By inputting the excerpt elements into the key feature extraction network, the convolutional layers extract the local semantic features of the excerpt elements, and then the attention mechanism is used to weight and focus the local semantic features, finally obtaining key element features that can reflect the key content of the excerpt elements.
[0063] Specifically, in the process of book mapping retrieval, the topological and structural features of the book page images uploaded by the application are first extracted to obtain the character topological features and page structural features of the book page images; at the same time, the key element features of the excerpt elements are extracted; then, based on the page structural features, the associated pages of the candidate books are determined, and feature matching is performed with the associated pages corresponding to the associated pages according to the character topological features, page structural features, and key element features to obtain the book location information.
[0064] As mentioned above, book mapping retrieval can also be performed using a book mapping retrieval model. In this case, the book mapping retrieval model can include a multimodal feature extraction module, an associated page determination module, and a feature matching module. In the multimodal feature extraction module, the image feature extraction submodule first performs multi-layer convolution and pooling operations on the input book page image to extract character topological features and page structure features, respectively. At the same time, the Transformer encoder of the text feature extraction submodule performs semantic encoding on the input excerpt elements and focuses on key elements through the Attention mechanism to extract key element features. In the associated page determination module, the page structure matching subunit locates the associated page with the same page number in the candidate books based on the page number corresponding to the book page image. Furthermore, in the feature matching module, the feature fusion layer first performs multimodal feature fusion of character topology features, page structure features, and key element features to generate unified query features; then, through the similarity calculation layer, the similarity between the query features and the features of the associated pages of the associated pages is calculated, and the associated page with the highest matching degree is selected based on the similarity calculation result, thereby outputting the book location information.
[0065] In the specific implementation process, it can be further combined with the above-mentioned method of configuring reliability weights and merging and sorting, and / or the method of displaying multiple candidate book location information and determining the book location information according to the user's selection instruction to improve user interaction satisfaction and retrieval accuracy. In this case, a new implementation method is obtained, which configures reliability weights and merges and sorts them during retrieval, displays multiple candidate book information after retrieval and determines the book location information according to the user's selection instruction. The specific implementation process refers to the implementation method provided above.
[0066] It should be noted that book mapping retrieval can also be performed by combining book page images, excerpt elements, and commentary data. Based on this, the above-mentioned book mapping retrieval to obtain book location information based on retrieval data containing excerpt elements and commentary data can be replaced by: extracting topological and structural features from book page images to obtain character topological features and page structural features; determining the associated pages of candidate books based on page structural features; extracting retrieval features from excerpt elements and commentary data to obtain excerpt retrieval features and commentary retrieval features; and performing comprehensive feature matching between character topological features, page structural features, excerpt retrieval features, and commentary retrieval features and the associated page features corresponding to the associated pages to finally generate book location information.
[0067] Furthermore, in practical applications, to efficiently utilize existing reference relationships in book page images to locate the position of referenced content within the original book, reference element identification and reference location retrieval can be performed based on the obtained book page images to obtain book location information. In one optional implementation of this embodiment, the method further includes: The book page image is subjected to reference element identification and feature extraction to obtain reference features. The referenced book is retrieved based on the reference features to obtain the reference location information of the referenced book as the book location information.
[0068] Specifically, firstly, the book page image is recognized using OCR (Optical Character Recognition) technology; then, the recognized citation elements are feature-extracted to obtain citation features; finally, based on the citation features, a matching query is performed in candidate books to obtain citation location information, which is then sent to the application as book location information.
[0069] As mentioned above, in the process of book mapping retrieval based on reference elements, a reference location retrieval model can also be used. In this case, the reference location retrieval model can include a reference element detection module, a reference feature encoding module, and a matching query module. In the reference element detection module, an object detection network based on a convolutional neural network architecture scans the image to identify reference elements. In the reference feature encoding module, OCR and syntax parsing algorithms are used to identify and standardize the reference elements and encode them into reference features. In the index query and location module, the reference features are matched with candidate books to obtain book location information.
[0070] In practical use, after completing the excerpting and annotation of a physical book, users may wish to review, manage, or reuse it. To improve interaction efficiency, one optional implementation of this embodiment further includes: The associated storage list is displayed based on the associated interaction commands submitted by the user on the interaction page; the associated storage list stores associated data obtained by associating excerpted elements with book location information and / or annotation data; The associated storage is implemented in the following way: Establish the association between the excerpted elements and the book's location information and / or annotation data, and generate associated data; store the associated data in the corresponding associated storage list based on the book's location information.
[0071] In the specific execution process, the system constructs the association between the extracted elements and the book location information and / or annotation data, and generates associated data. Then, based on the book location information, the system stores the associated data into the corresponding associated storage list. Specifically, it can store the data through the book identifier in the book location information, and add the associated data to the associated storage list corresponding to that book identifier. If the book identifier appears for the first time, a new associated storage list is created for it. This book-based storage method not only facilitates subsequent aggregation, display, and management by book dimension, but also provides a data foundation for building a cross-book knowledge graph.
[0072] For example, Figure 7 The interactive page can display the number of excerpts, excerpt elements, annotation data, and book location information. The interactive page can include a view control 701, a delete control 702, an export control 703, and a create control 704. Users can interact with the view control 701 to view their saved excerpts and can also trigger the delete control 702 to delete specified content. Users can use the export control 703 on the search page to export excerpts according to a preset format and can also trigger the create control 704 to jump to... Figure 3 Perform image acquisition and image recognition.
[0073] In summary, the interactive processing method for physical books provided in this embodiment first performs image acquisition and image recognition on the pages of the physical book to obtain book data and generate an excerpt interaction page. Then, to determine the user's excerpt elements, the method obtains the excerpt elements corresponding to the excerpt instructions submitted by the user based on the excerpt interaction page and uploads them to the server. Furthermore, to support user annotation of the excerpt elements, the method renders the excerpt elements and obtains the corresponding annotation interaction page, collects the annotation data entered on the annotation interaction page, and uploads it to the server. Subsequently, the method displays the annotations in association with the book location information and annotation data sent by the server. This achieves efficient digital excerpting, in-depth content annotation, and accurate book location mapping based on the physical book pages, thereby improving the convenience of user excerpting interaction and excerpt data management. Furthermore, in the process of book mapping retrieval, in order to accurately locate the user's excerpted content and thus improve the accuracy of user interaction with physical books, book mapping retrieval can be performed based on the excerpted elements. At the same time, in order to avoid the limitations of semantic ambiguity, cross-book repetition, or insufficient contextual information that may exist in the excerpted elements, and to further improve the accuracy and contextual relevance of the retrieval, book page images can also be introduced. Book mapping retrieval can be performed based on excerpted elements and book page images to obtain book location information, thereby improving the accuracy of book mapping retrieval.
[0074] Steps S202 to S208 provided in this embodiment can be executed by an application. It should be noted that the steps S202 to S208 executed by the application and steps S902 to S906 executed by the server in the following embodiment can cooperate with each other during execution. Therefore, when reading this embodiment, please refer to the corresponding content of steps S902 to S906 provided in the following method embodiment, and when reading the following method embodiment, please refer to the corresponding content of steps S202 to S208 provided in this embodiment.
[0075] The following example uses an interactive processing method for physical books provided in this embodiment, applied to an application program, specifically in a scenario involving excerpting and commenting on physical books. Figure 8 The interactive processing method for physical books provided in this embodiment will be further explained below. Figure 8 The interactive processing method for physical books, applied to the scenario of excerpting and commenting on physical books, specifically includes the following steps: Step S802: Image acquisition and image recognition are performed on the content pages of the physical book to obtain book data.
[0076] Step S804: Generate an interactive excerpt page based on book data and book page images.
[0077] Step S806: Obtain the extract element corresponding to the extract instruction submitted by the user based on the extract interaction page.
[0078] Step S808: Upload the extracted elements to the server.
[0079] Step S812: Upload the book page image to the server.
[0080] Step S816: Render the extracted elements and obtain the corresponding comment interaction page, and collect the comment data entered on the comment interaction page.
[0081] Step S818: Upload the annotation data to the server.
[0082] Step S826: Receive book location information obtained by the server after performing book mapping retrieval based on annotation data, book page images, and excerpt elements.
[0083] Step S828: Display the annotations in association based on the book's location information and annotation data.
[0084] Steps S802 to S808, S812, S816, S818, S826, and S828 provided in this embodiment are executed by the application. It should be noted that the steps S802 to S808, S812, S816, S818, S826, and S828 executed by the application can cooperate with steps S810, S814, and S820 to S824 executed by the server in the following embodiment. Therefore, when reading this embodiment, please refer to the corresponding content of steps S810, S814, and S820 to S824 provided in the following method embodiment. When reading the following method embodiment, please refer to the corresponding content of steps S802 to S808, S812, S816, S818, S826, and S828 provided in this embodiment.
[0085] It should be noted that any one or more of steps S802 to S808, S812, S816, S818, S826, and S828 can be combined with any one or more of steps S202 to S208 to form a new implementation method according to the needs of implementation and deployment; in addition, steps S802 to S808, S812, S816, S818, S826, and S828 can be combined according to the actual deployment needs. Choose any one or more technical features and combine them with any one or more technical features provided in steps S202 to S208 above to form a new implementation method; or, any one or more technical features in steps S802 to S808, S812, S816, S818, S826, and S828 can be replaced with any one or more technical features provided in steps S202 to S208 above to form a new implementation method according to the actual deployment needs, which will not be elaborated here.
[0086] One or more embodiments of another interactive processing method for physical books provided in this specification are as follows: Reference Figure 9 The interactive processing method for physical books provided in this embodiment can be applied to the server side. The method specifically includes steps S902 to S906.
[0087] Step S902: Receive the excerpt element uploaded by the application; the excerpt element is submitted based on the excerpt interaction page; the excerpt interaction page is generated based on book data obtained by image acquisition and image recognition of the book pages of the physical book.
[0088] The book page mentioned in this embodiment refers to a page in a physical book. Specifically, a book page can be a table of contents page, a main text page, a reference page, or an appendix page. The function of a book page is to carry knowledge content for readers to read and cite. This is different from a book label page, which serves as physical packaging and overall identification. The book label page can be the cover page.
[0089] The excerpt interaction page refers to an interactive interface generated based on book data, used by users to select and submit excerpt elements. The excerpt interaction page includes a display area for showing the excerptable content of the book pages and is configured with interactive controls for users to select and submit excerpt elements.
[0090] The extraction command refers to the operation command initiated by the user through interaction on the above-mentioned extraction interaction page to select book data. Specifically, the extraction command can be that the user interacts with the book data on the extraction interaction page by clicking, or the user interacts by selecting by box selection to select continuous or non-contiguous areas of book data; wherein, the extraction element refers to the corresponding book data extracted according to the above-mentioned extraction command.
[0091] In practice, to convert the pages of a physical book into book data that can be used for interactive processing, the application first captures images of the pages of the physical book to obtain book page images, and then performs image recognition on the book page images to obtain book data. The book data is the data obtained after capturing and recognizing the images of the book pages, and is used to generate excerpt interactive pages. Specifically, the book data can be book content data composed of the reading content of the book pages, such as text data, chart data and / or formula data in the book pages, or book structure data composed of the layout information of the book pages, such as text and image layout relationship data and paragraph division data. Based on this, after obtaining the book data, an excerpt interaction page is generated. The excerpt interaction page is an interactive interface generated based on the book data, used by users to select and submit excerpt elements. The excerpt interaction page includes a display area for displaying the excerptable content of the book pages and is configured with interactive controls for users to select and submit excerpt elements.
[0092] For example, Figure 3 The interactive page shown includes an image upload control 301 and a capture control 302. Users can interact with the image upload control 301 to select images; specifically, images can be uploaded by clicking or dragging. Supported image formats include JPG and PNG, and the application can automatically recognize text. Users can also interact with the capture control 302 to activate the camera and directly capture images of book pages for recognition. After the user completes image upload or capture, the user terminal displays... Figure 4 The extracted element recognition page shown displays the original image at the top and a recognition progress bar at the bottom before the recognition is complete.
[0093] In practical applications, different acquisition methods may be used when capturing images of book pages. For example, a user may take a horizontal or vertical shot of the book page depending on the posture of holding the user's terminal, or they may choose to capture the entire page or only a local area from the perspective of the book page content. Different acquisition methods will lead to differences in the quality of the captured book page images. To improve the quality and accuracy of image acquisition, this embodiment provides an optional implementation method for image acquisition and image recognition of physical book pages, including: The framing features of the book page images are extracted, and the image acquisition mode is determined based on the obtained framing features; Configure the parameters according to the acquisition constraint parameters corresponding to the image acquisition mode, acquire book page images, and perform image recognition on the book page images to obtain book data.
[0094] Among them, the book page viewfinder image refers to the preview image provided by the application before image acquisition, used to analyze the current acquisition method; the viewfinder features refer to the features extracted from the book page viewfinder image to determine the current acquisition method, including shooting direction, integrity of the book page, ambient light intensity and / or distortion, and there is a one-to-one mapping relationship between viewfinder features and image acquisition mode; the acquisition constraint parameters refer to the acquisition parameters configured for the image acquisition component based on different image acquisition modes. Specifically, the acquisition constraint parameters can be optical parameters associated with the acquisition method, image processing parameters, or output format parameters.
[0095] Specifically, in the process of extracting framing features from the view image of a book page, a framing feature extraction network can be used. This network can be a lightweight convolutional neural network (CNN) that integrates the DOC (Document Orientation Correction) algorithm and the ROID (Region of Interest Detection) algorithm. By inputting the view image of the book page into the network, the DOC algorithm is used to analyze the main orientation of the image to correct the screen rotation caused by the user's grip posture and output the orientation features. Then, the ROID algorithm is used to detect the boundary regions based on the corrected view image and output the region features. Finally, the orientation features and region features are fused to obtain the framing features.
[0096] After obtaining the above-mentioned framing features, the corresponding image acquisition mode is determined based on the framing features, and the image acquisition component is configured according to the corresponding acquisition constraint parameters. The image acquisition component is then used to acquire images of book pages, and image recognition is performed on the book page images to obtain book data. For example, when the framing features represent vertical and local areas, the image acquisition mode is determined to be a local vertical mode. The optical parameters of the image acquisition component are adjusted according to the preset acquisition constraint parameters of this mode, and book page images are acquired and recognized.
[0097] In the specific execution process, to improve the reliability of book page images and thus the accuracy of subsequent image recognition, after configuring parameters according to the acquisition constraint parameters corresponding to the image acquisition mode and acquiring book page images, and before performing image recognition to obtain book data, image parameter detection and calculation can be performed on the book page images to determine whether the book page images meet the image parameter conditions. In an optional implementation provided in this embodiment, it further includes: Image parameter detection and calculation are performed on book page images to obtain image parameter evaluation values; If the image parameter evaluation value meets the corresponding image parameter conditions, perform image recognition on the book page image to obtain book data; if the image parameter evaluation value does not meet the corresponding image parameter conditions, generate the corresponding guidance instruction and perform secondary image acquisition on the book page.
[0098] The image parameter evaluation value refers to the image quality score obtained after evaluating the acquired book page images from multiple dimensions. Specifically, these multiple dimensions include sharpness, integrity, illumination intensity, and / or image geometry. The image parameter conditions refer to the preset image quality thresholds for the above multiple dimensions, which are the basis for determining whether a second image acquisition is needed. For example, the image parameter conditions could be that the sharpness evaluation value is greater than x and the illumination intensity evaluation value is within the preset threshold. Only when the image parameter evaluation value meets the image parameter conditions will the operation of image recognition to obtain book data be performed on the book page images. In the specific execution process, image parameter detection and calculation are performed on the book page images to obtain image parameter evaluation values. Image parameter detection and calculation can be performed through an image quality assessment model. The image quality assessment network can be a lightweight convolutional neural network that integrates the NR-IQA (Comprehensive No-Reference Image Quality Assessment) algorithm. By inputting the book page images into the image quality assessment network, image parameter detection and calculation are performed to obtain image parameter evaluation values.
[0099] Furthermore, the image parameter evaluation value is judged. If the image parameter evaluation value meets the corresponding image parameter conditions, image recognition is performed on the book page image to obtain book data. If the corresponding image parameter conditions are not met, a corresponding guidance instruction is generated and a second image acquisition of the book page is performed. The guidance instruction refers to the multimodal operation instruction information generated by the system and displayed to the user when the image parameter evaluation value does not meet the standard, in order to guide the user to complete the second image acquisition. Specifically, the guidance instruction includes text guidance information, voice guidance information and / or image guidance information. For example, if the book page image is overexposed, the guidance instruction "The image is overexposed, please avoid direct sunlight" is generated.
[0100] In practice, in order to enhance the user's interactive experience with physical books and enable book data to be overlaid on the book page images in an interactive form to generate an excerpt interactive page, thereby achieving accurate excerpting of the content of physical books by users, the excerpt interactive page can be generated based on the book page images and book data in the upper layer of the interactive page. Specifically, in one optional implementation of this embodiment, the extract interaction page is generated in the following manner: The application displays book page images on its interactive page and establishes a coordinate mapping relationship between the book data and the corresponding image element regions in the book page images. Based on the coordinate mapping relationship, the book data is rendered in the upper layer of the interactive page to obtain the excerpt interactive page.
[0101] The coordinate mapping relationship refers to the positional relationship between the book data and the corresponding image element region in the book page image. Specifically, the coordinate mapping relationship is obtained by calculating the positional association between the book data and the corresponding image element in the book page image, so that any data in the book data can be mapped to the corresponding element region in the book page image, providing a basis for the subsequent generation of interactive excerpt pages.
[0102] In the specific execution process, when rendering the book data in the upper layer of the interactive page, the book data is drawn to the position corresponding to the image elements in the book page image based on the above coordinate mapping relationship, and the book data is made interactive, thereby forming an interactive page that is visually aligned with the book page image and supports direct user interaction, providing a foundation for users to make excerpts later.
[0103] It should be noted that, considering that the physical book images, user comments, and other related data involved in this specification may, to some extent, constitute user privacy, authorization from the user must be obtained before collecting such data to ensure that the data collection operation complies with relevant data management regulations. For example, user authorization can be granted during user registration or when collecting images. Specific methods of data authorization could include sending a data authorization reminder to the user, who can then confirm the reminder with an instruction to obtain data authorization; alternatively, data authorization could be obtained by signing a data authorization agreement to obtain authorization for data collection or data transmission. This embodiment does not limit the scope of the authorization.
[0104] Based on the generated excerpt interaction page, the application first obtains the excerpt instructions submitted by the user on the excerpt interaction page. Based on these instructions, it extracts or selects the corresponding excerpt elements from the book data. After obtaining the excerpt elements, it can upload them to the server. Correspondingly, the server receives the excerpt elements uploaded by the application, providing a data foundation for subsequent book mapping retrieval. Here, the excerpt instruction refers to the operation command initiated by the user through interaction on the aforementioned excerpt interaction page to select book data. Specifically, the excerpt instruction can be that the user interacts with the book data on the excerpt interaction page by clicking, or it can be that the user interacts by selecting a box to select continuous or non-contiguous areas of book data. The excerpt element refers to the corresponding book data extracted according to the aforementioned excerpt instructions.
[0105] Alternatively, the book page images obtained through image acquisition and image recognition, along with the excerpt elements, can be uploaded to the server. In this case, the step of receiving the excerpt elements uploaded by the application can be replaced by receiving the book page images and excerpt elements uploaded by the application.
[0106] In the specific execution process, different interaction methods may be used when users interact with the excerpt interaction page. For example, some users prefer to interact with the excerpt interaction page by clicking, while others prefer to use a selection box. In this case, in order to intuitively display the user's interaction results and thus improve the user's interaction experience and the accuracy of the excerpt, the excerpt elements and the corresponding excerpt image element areas can be synchronously marked. In an optional implementation method provided in this embodiment, obtaining the excerpt element corresponding to the excerpt instruction submitted by the user based on the excerpt interaction page includes: Extract the extract elements from the book data based on at least one extract range corresponding to the extract instruction; Based on the coordinate mapping relationship, the corresponding image element region of the excerpt element in the book page image is determined, and the excerpt element and the image element region are marked synchronously.
[0107] Synchronous marking refers to marking the corresponding areas of the book page image, i.e., the areas of the extracted image elements, according to the coordinate mapping relationship while marking the extracted excerpt elements. This allows the user's interactive operations on the excerpt page to be intuitively mapped to the book page image. Specifically, synchronous marking can take the form of visual synchronous marking, such as highlighting elements or highlighting areas of elements.
[0108] For example, Figure 5 The interactive page shown includes an excerpt element selection interface 501 and an excerpt element display interface 502. Users can select or confirm excerpt elements in the excerpt element selection interface 501. After the selection of excerpt elements is completed, the selected excerpt elements will be displayed in the excerpt element display interface 502.
[0109] Step S904: Obtain the comment data entered by the user on the comment interaction page; perform book mapping retrieval based on the search data containing the excerpt element and the comment data to obtain book location information; the comment interaction page is generated based on the excerpt element.
[0110] The comment interaction page described in this embodiment refers to an interactive interface generated based on the aforementioned excerpt elements for inputting comment data. Specifically, the comment interaction page includes an excerpt display area that displays the excerpt elements and is configured with interactive controls that support multimodal input and comment submission controls to input and submit comment data for the excerpt elements. Among them, the interactive controls can be text boxes that support text input, recording controls that support voice input, or interactive controls that support image upload. Correspondingly, the comment data refers to the data created by the user in the comment interaction page through the interactive controls for commenting on or annotating the excerpt elements. The comment data includes text comment data, voice comment data, and / or image comment data.
[0111] For example, Figure 6 The interactive page shown can display the user's... Figure 5 The selected excerpt elements, along with the corresponding book information and matching degree, are displayed on the interactive page. The page may also include an annotation data input interface 601 and a voice annotation input control 602. Users can input annotations on the annotation data input interface 601 and interact with the voice annotation input control 602 to input voice annotations.
[0112] Alternatively, instead of uploading the annotation data, the annotation data can be stored in the application and a annotation request indicating that the user's annotation is complete can be sent to the server. Correspondingly, after receiving the annotation request, the server performs a book mapping search based on the excerpt elements. In this case, the operation of obtaining the annotation data entered by the user on the annotation interaction page generated based on the excerpt elements can be replaced by obtaining the annotation request submitted by the user on the annotation interaction page generated based on the excerpt elements.
[0113] Accordingly, the operation of obtaining book location information by performing book mapping retrieval based on search data containing excerpt elements and commentary data can be replaced by obtaining book location information by performing book mapping retrieval based on search data containing commentary requests and excerpt elements.
[0114] In practical applications, user comments on extracted elements often contain the user's emotional inclinations and cognitive intentions. For example, a text comment saying "very good" reflects the user's approval, while a hurried voice comment may imply the user's confusion. In this case, to better understand the user's input comments, recommended resources can be determined based on the comments. Specifically, sentiment tag recognition can be performed on the comments to obtain sentiment tags and determine recommended resources. In addition, the comments can be uploaded to the server, where sentiment tag recognition can be performed to obtain sentiment tags, determine recommended resources, and then distribute the recommended resources to the application.
[0115] Specifically, in one optional implementation of this embodiment, after the steps of collecting the comment data entered on the comment interaction page and uploading it to the server are performed, the method further includes: The annotation data is encoded to obtain annotation encoded data, and the annotation emotional features are extracted from the annotation encoded data to obtain annotation emotional features; The emotional features of the comments are mapped to emotional tags to obtain emotional tags, and recommended resources are determined based on the emotional tags.
[0116] In the specific execution process, the above-mentioned sentiment tag mapping and recommendation resource determination can be performed through the sentiment mapping model. Specifically, the annotation encoder in the sentiment mapping model can encode the annotation data. The annotation encoder can be a Transformer encoder that integrates the Attention mechanism. By inputting the annotation data into the Transformer encoder, the encoder uses the Attention mechanism to analyze the global dependency of the annotation data and understands and encodes the deep semantics and sentiment tendency of the annotation through the knowledge obtained by pre-training, thus obtaining the annotation encoded data. After obtaining the annotation coding data, the emotional features of the annotation coding data are extracted to obtain the emotional features of the annotation. In the above emotional feature extraction process, the emotional features can be extracted through the emotional feature extraction module. The emotional feature extraction module can be a multilayer perceptron (MLP) based on fully connected layers. By inputting the annotation coding data into the multilayer perceptron, the fully connected layers inside the multilayer perceptron perform multi-layer nonlinear transformations on the annotation coding data, gradually abstracting deep features that are highly related to the emotional judgment from the original annotation coding data, and generating the emotional features of the annotation based on the deep features. Furthermore, in the process of obtaining sentiment tags by mapping sentiment features to sentiment labels, sentiment label mapping can be performed through a sentiment mapping network. This sentiment mapping network can be a lightweight convolutional neural network integrating a Softmax output layer. By inputting the sentiment features of the comments into the sentiment mapping network, the feature transformation layer first integrates and adjusts the dimensions of the sentiment features. The Softmax output layer normalizes the transformed sentiment features, calculates and outputs a set of probability distributions corresponding to preset sentiment labels. Finally, the system selects the label with the highest probability value as the sentiment label that accurately summarizes the sentiment tendency of the comment. Based on this sentiment label, the system recommends corresponding resources to the user, such as related articles and videos. For example, if user A comments a passage in a serious literary work as "well written," after model analysis, the sentiment label of this comment is determined to be "highly appreciated." Based on this sentiment label and the serious literary genre, the system matches and recommends a set of literary works and literary criticism articles of the same style to the user, thereby evoking emotional resonance and further expanding the user's reading comprehension.
[0117] Optionally, the annotation data can be multimodal annotation data, such as including voice annotation data and image annotation data. In this case, in order to more comprehensively and accurately understand the user's multimodal emotions, the above-mentioned step of encoding the annotation data to obtain annotation encoded data and extracting emotional features from the annotation encoded data to obtain annotation emotional features can be replaced by: decoupling the multimodal annotation data to obtain each modality's annotation data, and encoding and extracting emotional features from each modality's annotation data to obtain the corresponding emotional features of each modality. Correspondingly, the step of mapping sentiment features to sentiment labels to obtain sentiment labels and determining recommended resources based on sentiment labels can be replaced by: mapping sentiment features to sentiment labels to obtain sentiment labels for each modality, performing weighted fusion of sentiment labels for each modality based on preset modality sentiment weights, and determining recommended resources based on the weighted sentiment labels obtained from the weighted fusion.
[0118] In this embodiment, book mapping retrieval refers to the process of starting from the pages of a physical book and performing an online search based on search data. Specifically, book mapping retrieval can be searching for the electronic version of a physical book, or searching for books associated with a physical book, such as different versions of books with the same book identifier, or searching for other books referenced on the pages of a physical book. Book location information refers to the book information obtained through book mapping retrieval, and / or the specific element location information corresponding to the current excerpt element, such as the page number, chapter number, and / or page position of the excerpt element. Optionally, book location information is obtained by performing book mapping retrieval based on search data containing excerpt elements and annotation data.
[0119] In practice, after receiving the book location information from the server, in order to intuitively associate the user-submitted annotation data and excerpts with the corresponding areas in the physical book, thereby improving the comprehensibility of the annotation data and the efficiency of knowledge review, the annotation association display can be completed based on the book location information and the annotation data.
[0120] In practical applications, to identify the intent and contextual semantics of excerpted elements in user-submitted comment data, thereby improving the accuracy of book mapping retrieval, book mapping retrieval can be performed based on retrieval data containing excerpted elements and comment data. Specifically, in one optional implementation of this embodiment, book mapping retrieval based on retrieval data containing excerpted elements and comment data includes: Search features are extracted from the excerpted elements and commentary data to obtain excerpt search features and commentary search features, and semantic features are extracted from the excerpted elements to obtain semantic features; Book location information is obtained by performing book mapping retrieval in the candidate book database based on semantic features, commentary retrieval features, and excerpt retrieval features.
[0121] Specifically, a semantic feature extraction network is used to perform semantic parsing on the excerpted elements and extract semantic features. At the same time, a retrieval feature extraction network is used to process the excerpted elements and commentary data respectively to obtain excerpt retrieval features and commentary retrieval features that represent their retrieval intent. Based on the semantic features, commentary retrieval features, and excerpt retrieval features, book mapping retrieval is performed in the candidate book database to obtain book location information. Among them, the semantic feature extraction network can be a large language model based on the BERT (Bidirectional Encoder Representations from Transformers) architecture, which identifies the contextual semantic information of the excerpted elements and outputs semantic features through its fully connected layers and attention mechanism.
[0122] Furthermore, in the above-mentioned book mapping retrieval process, in order to comprehensively evaluate the contribution of different features to the retrieval results and improve the robustness and accuracy of book mapping retrieval, book mapping retrieval can also be performed by configuring reliability weights and merging rankings. In one optional implementation method provided in this embodiment, book mapping retrieval is performed in the candidate book database based on semantic features, commentary retrieval features, and excerpt retrieval features to obtain book location information, including: Input semantic features, commentary retrieval features, and summary retrieval features into the retrieval model and obtain their respective retrieval result sets; Assign corresponding confidence weights to each search result set, and perform weighted fusion and sorting of the search result sets based on the confidence weights to obtain a mapping search list; determine the book location information based on the mapping search list.
[0123] The retrieval result set refers to the preliminary set of book location information obtained after using any one of the semantic features, commentary retrieval features, or excerpt retrieval features as query input and searching candidate books through a retrieval model. Specifically, each retrieval result set reflects the retrieval tendency from a single dimension. For example, a retrieval result set based on semantic features may be more biased towards chapters related to concepts, while a retrieval result set based on excerpt retrieval features may be more focused on paragraphs containing specific content. The mapping retrieval list refers to the list of results generated by weighting and merging multiple retrieval result sets according to their corresponding confidence weights and then reordering them, arranged in descending order of relevance scores. Specifically, the mapping retrieval list integrates retrieval information obtained from searches based on different features, thereby providing a more accurate and comprehensive candidate sequence of book location information.
[0124] Semantic features, commentary retrieval features, and excerpt retrieval features are each input as independent queries into the retrieval model. The retrieval model can be a dense retrieval model based on feature similarity calculation. For each input feature, the retrieval model performs a retrieval and returns a set of retrieval results sorted in descending order of similarity score. Then, a preset confidence weight is assigned to the obtained retrieval result set, and the retrieval result set is weighted, fused, and sorted based on the confidence weight to obtain a mapping retrieval list. Finally, the book location information is determined based on the generated retrieval mapping list.
[0125] In practical applications, due to the potentially large number of candidate books or the ambiguity of user excerpts and annotations, several results in the book-mapping retrieval list may have similarities. In such cases, to improve user satisfaction and retrieval accuracy, multiple candidate book location information can be displayed, and the book location information can be determined based on the user's selection. In one optional implementation of this embodiment, annotation association display is performed based on the book location information and annotation data sent by the server, including: If the server detects that multiple candidate book location information has been sent, the location information of each candidate book will be displayed on the interactive page of the application; the corresponding candidate book location information will be determined based on the selection command submitted by the user and displayed in conjunction with the annotation data.
[0126] In practice, when the server distributes location information for multiple candidate books, the application displays the location information for each candidate book on the interactive interface. Users browse and compare these candidate book location information and submit selection commands to determine the book location information. Then, the information is displayed in association with the annotation data. The annotation association display refers to binding the user-submitted annotation data with the book location information and displaying it uniformly on the interactive interface. This allows the user's annotation to be clearly associated with the specific source location in the physical book, thereby establishing a connection between digital annotations and physical books. For example, a label such as "Source: Das Kapital, page 129" can be displayed next to an annotation. Clicking this label allows users to further preview the context or jump to the electronic version page.
[0127] Furthermore, depending on the actual application scenario or user needs, the objects of the above-mentioned associated display are not limited to annotation data. Book location information can also be associated with excerpt elements. In this case, the step of displaying annotations based on the book location information and annotation data sent by the server can be replaced with displaying excerpts based on the book location information and excerpt elements sent by the server.
[0128] Furthermore, in practical applications, the same physical book may have multiple different published versions, and the text content of these versions is often very similar. In this case, it may be difficult to locate the physical book by performing book mapping retrieval based on the user's excerpts and annotations. To improve the accuracy of book mapping retrieval, book page images obtained through image acquisition and image recognition can also be uploaded to the server. The uploading of book page images can be performed before or after the operation of uploading excerpts, but before the book mapping retrieval.
[0129] After obtaining the book page images, a book mapping retrieval can be performed based on the retrieval data containing the book page images and excerpt elements. Specifically, in one optional implementation method provided in this embodiment, the following is included: Upload images of book pages obtained through image acquisition and image recognition to the server; Topological and structural features are extracted from book page images to obtain character topological features and page structural features; Based on page structure features, the associated pages of candidate books are determined. Feature matching is performed between character topology features, page structure features, and key element features of excerpted elements and the associated page features corresponding to the associated pages to obtain book location information.
[0130] Among them, character topological features refer to the geometric features extracted from book page images that describe the shape, size, and relative position of characters; page structure features refer to the features that characterize the layout and / or format of book pages. Related pages refer to pages selected from candidate books based on the similarity of page structure features. Specifically, related pages can be pages with the same page number as the page image of a book. By locating the related page with the corresponding page number in the candidate books based on the page number corresponding to the page image of the book, the retrieval efficiency can be improved.
[0131] Key element features refer to the features extracted from the excerpt elements that can represent the key content of the excerpt elements. Specifically, in the process of extracting key element features, key element features can be extracted through a key feature extraction network. The key feature extraction network can be a CNN integrating an attention mechanism. By inputting the excerpt elements into the key feature extraction network, the convolutional layers extract the local semantic features of the excerpt elements, and then the attention mechanism is used to weight and focus the local semantic features, finally obtaining key element features that can reflect the key content of the excerpt elements.
[0132] Specifically, in the process of book mapping retrieval, the topological and structural features of the book page images uploaded by the application are first extracted to obtain the character topological features and page structural features of the book page images; at the same time, the key element features of the excerpt elements are extracted; then, based on the page structural features, the associated pages of the candidate books are determined, and feature matching is performed with the associated pages corresponding to the associated pages according to the character topological features, page structural features, and key element features to obtain the book location information.
[0133] As mentioned above, book mapping retrieval can also be performed using a book mapping retrieval model. In this case, the book mapping retrieval model can include a multimodal feature extraction module, an associated page determination module, and a feature matching module. In the multimodal feature extraction module, the image feature extraction submodule first performs multi-layer convolution and pooling operations on the input book page image to extract character topological features and page structure features, respectively. At the same time, the Transformer encoder of the text feature extraction submodule performs semantic encoding on the input excerpt elements and focuses on key elements through the Attention mechanism to extract key element features. In the associated page determination module, the page structure matching subunit locates the associated page with the same page number in the candidate books based on the page number corresponding to the book page image. Furthermore, in the feature matching module, the feature fusion layer first performs multimodal feature fusion of character topology features, page structure features, and key element features to generate unified query features; then, through the similarity calculation layer, the similarity between the query features and the features of the associated pages of the associated pages is calculated, and the associated page with the highest matching degree is selected based on the similarity calculation result, thereby outputting the book location information.
[0134] In the specific implementation process, it can be further combined with the above-mentioned method of configuring reliability weights and merging and sorting, and / or the method of displaying multiple candidate book location information and determining the book location information according to the user's selection instruction to improve user interaction satisfaction and retrieval accuracy. In this case, a new implementation method is obtained, which configures reliability weights and merges and sorts them during retrieval, displays multiple candidate book information after retrieval and determines the book location information according to the user's selection instruction. The specific implementation process refers to the implementation method provided above.
[0135] It should be noted that book mapping retrieval can also be performed by combining book page images, excerpt elements, and commentary data. Based on this, the above-mentioned book mapping retrieval to obtain book location information based on retrieval data containing excerpt elements and commentary data can be replaced by: extracting topological and structural features from book page images to obtain character topological features and page structural features; determining the associated pages of candidate books based on page structural features; extracting retrieval features from excerpt elements and commentary data to obtain excerpt retrieval features and commentary retrieval features; and performing comprehensive feature matching between character topological features, page structural features, excerpt retrieval features, and commentary retrieval features and the associated page features corresponding to the associated pages to finally generate book location information.
[0136] Furthermore, in practical applications, to efficiently utilize existing reference relationships in book page images to locate the position of referenced content within the original book, reference element identification and reference location retrieval can be performed based on the obtained book page images to obtain book location information. In one optional implementation of this embodiment, the method further includes: The book page image is subjected to reference element identification and feature extraction to obtain reference features. The referenced book is retrieved based on the reference features to obtain the reference location information of the referenced book as the book location information.
[0137] Specifically, firstly, the book page image is recognized using OCR (Optical Character Recognition) technology; then, the recognized citation elements are feature-extracted to obtain citation features; finally, based on the citation features, a matching query is performed in candidate books to obtain citation location information, which is then sent to the application as book location information.
[0138] As mentioned above, in the process of book mapping retrieval based on reference elements, a reference location retrieval model can also be used. In this case, the reference location retrieval model can include a reference element detection module, a reference feature encoding module, and a matching query module. In the reference element detection module, an object detection network based on a convolutional neural network architecture scans the image to identify reference elements. In the reference feature encoding module, OCR and syntax parsing algorithms are used to identify and standardize the reference elements and encode them into reference features. In the index query and location module, the reference features are matched with candidate books to obtain book location information.
[0139] In practical use, after completing the excerpting and annotation of a physical book, users may wish to review, manage, or reuse it. To improve interaction efficiency, one optional implementation of this embodiment further includes: The associated storage list is displayed based on the associated interaction commands submitted by the user on the interaction page; the associated storage list stores associated data obtained by associating excerpted elements with book location information and / or annotation data; The associated storage is implemented in the following way: Establish the association between the excerpted elements and the book's location information and / or annotation data, and generate associated data; store the associated data in the corresponding associated storage list based on the book's location information.
[0140] In the specific execution process, the system constructs the association between the extracted elements and the book location information and / or annotation data, and generates associated data. Then, based on the book location information, the system stores the associated data into the corresponding associated storage list. Specifically, it can store the data through the book identifier in the book location information, and add the associated data to the associated storage list corresponding to that book identifier. If the book identifier appears for the first time, a new associated storage list is created for it. This book-based storage method not only facilitates subsequent aggregation, display, and management by book dimension, but also provides a data foundation for building a cross-book knowledge graph.
[0141] For example, Figure 7 The interactive page can display the number of excerpts, excerpt elements, annotation data, and book location information. The interactive page can include a view control 701, a delete control 702, an export control 703, and a create control 704. Users can interact with the view control 701 to view their saved excerpts and can also trigger the delete control 702 to delete specified content. Users can use the export control 703 on the search page to export excerpts according to a preset format and can also trigger the create control 704 to jump to... Figure 3 Perform image acquisition and image recognition.
[0142] Step S906: Send the book location information to the application for annotation association display.
[0143] Among them, the annotation association display refers to binding the user-submitted annotation data with the book's location information and displaying it uniformly in the interactive interface, so that the user's annotation can be clearly associated with the specific source location in the physical book, thereby establishing a connection between digital annotations and physical books. For example, next to an annotation data, a label "Source: Das Kapital, page 129" can be displayed. Clicking on the label can further preview the context or jump to the electronic version page.
[0144] Furthermore, depending on the actual application scenario or user needs, the objects of the above-mentioned associated display are not limited to annotation data. Book location information can also be associated with excerpt elements. In this case, the step of displaying annotations based on the book location information and annotation data sent by the server can be replaced with displaying excerpts based on the book location information and excerpt elements sent by the server.
[0145] The following example uses the interactive processing method for physical books provided in this embodiment, applied to the server side, as an example of its application in a physical book excerpting and commenting scenario, combined with... Figure 8 The interactive processing method for physical books provided in this embodiment will be further explained below. Figure 8 The interactive processing method for physical books, applied to the scenario of excerpting and commenting on physical books, specifically includes the following steps: Step S810: Receive the excerpt elements uploaded by the application.
[0146] Step S814: Receive the book page image uploaded by the application.
[0147] Step S820: Receive the comment data uploaded by the application.
[0148] Step S822: Based on the annotation data, book page images, and excerpt elements, perform book mapping retrieval and obtain book location information.
[0149] Step S824: Send the book location information to the application.
[0150] It should be noted that any one or more of steps S810, S814, and S820 to S824 can be combined with any one or more of steps S902 to S906 to form a new implementation method according to the needs of implementation and deployment. In addition, any one or more technical features in steps S810, S814, and S820 to S824 can be selected and combined with any one or more technical features provided in steps S902 to S906 to form a new implementation method according to the actual deployment needs. Alternatively, any one or more technical features in steps S810, S814, and S820 to S824 can also be replaced with any one or more technical features provided in steps S902 to S906 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.
[0151] The following is an embodiment of an interactive processing device for physical books provided in this specification: In the above embodiments, an interactive processing method for physical books is provided. Correspondingly, an interactive processing device for physical books is also provided, which runs on an application program. The following description is in conjunction with the accompanying drawings.
[0152] Reference Figure 10 This illustration shows a schematic diagram of an embodiment of an interactive processing device for physical books provided in this embodiment.
[0153] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0154] This embodiment provides an interactive processing device for physical books, running on an application program. The device includes: The book data acquisition module 1002 is configured to perform image acquisition and image recognition on the pages of physical books, obtain book data, and generate an interactive excerpt page; The text extraction module 1004 is configured to obtain the extraction element corresponding to the extraction instruction submitted by the user based on the extraction interaction page and upload it to the server. The comment data submission module 1006 is configured to render the extracted elements and obtain the corresponding comment interaction page, collect the comment data input on the comment interaction page, and upload it to the server. The excerpt display module 1008 is configured to display annotations in association with the book location information sent by the server and the annotation data; the book location information is obtained by book mapping retrieval based on retrieval data containing the excerpt elements and the annotation data.
[0155] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0156] Another embodiment of the interactive processing device for physical books provided in this specification is as follows: In the above embodiments, another interactive processing method for physical books is provided, and correspondingly, another interactive processing device for physical books is also provided, which will be described below with reference to the accompanying drawings.
[0157] Reference Figure 11 This illustration shows a schematic diagram of another embodiment of the interactive processing device for physical books provided in this embodiment.
[0158] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0159] This embodiment provides an interactive processing device for physical books, running on a server. The device includes: The excerpt element receiving module 1102 is configured to receive excerpt elements uploaded by the application; the excerpt elements are submitted based on the excerpt interaction page; the excerpt interaction page is generated based on book data obtained by image acquisition and image recognition of book pages of physical books; The book mapping retrieval module 1104 is configured to obtain the annotation data entered by the user on the annotation interaction page, and perform book mapping retrieval based on the retrieval data containing the excerpt element and the annotation data to obtain book location information; the annotation interaction page is generated based on the excerpt element. The book location information sending module 1106 is configured to send the book location information to the application for annotation association display.
[0160] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0161] The following is an embodiment of an interactive processing device for physical books provided in this specification: Corresponding to the above-described interactive processing method for physical books, based on the same technical concept, one or more embodiments of this specification also provide an interactive processing device for physical books, which is used to execute the above-described interactive processing method for physical books. Figure 12 This is a schematic diagram of the structure of an interactive processing device for a physical book, provided for one or more embodiments of this specification.
[0162] This embodiment provides an interactive processing device for physical books, including: like Figure 12As shown, device 1200 mainly consists of a communication interface 1202, a user interface 1204, a processor 1206, and a data storage 1208. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1210. The communication interface 1202 enables device 1200 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1202 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1202 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1202 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1202 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1204 includes receiving user input and providing output to the user. Therefore, user interface 1204 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1204 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1204 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 1200 may support remote access from other devices via communication interface 1202 or another physical interface (not shown). User interface 1204 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1204 may also be configured as a display device for rendering or displaying text fragments.
[0163] Processor 1206 may include one or more general-purpose processors and / or dedicated processors. Data storage 1208 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1206. Data storage 1208 may include removable and non-removable components.
[0164] Processor 1206 is capable of executing program instructions 1218 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 1208 to perform the various functions described herein. Data storage 1208 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1200, enable device 1200 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1218 by processor 1206 may result in processor 1206 using data 1212. For example, program instructions 1218 may include an operating system 1222 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1200 and one or more application programs 1220 (e.g., a browser, social application, or game application). Similarly, data 1212 may include operating system data 1216 and application data 1214. Operating system data 1216 is primarily accessible to operating system 1222, while application data 1214 is primarily accessible to one or more application programs 1220. Application data 1214 may reside in a file system visible or hidden to the user of device 1200. Application 1220 may communicate with operating system 1212 via one or more application programming interfaces (APIs). These APIs facilitate application 1220 reading and / or writing application data 1214, transmitting or receiving information via communication interface 1202, receiving or displaying information on user interface 1204, etc. In some terms, application 1220 may be simply referred to as "app". Furthermore, application 1220 may be downloaded to device 1200 through one or more online app stores or app markets. However, applications may also be installed on device 1200 in other ways, such as through a web browser or a physical interface on device 1200 (e.g., a USB port).
[0165] In one specific embodiment, the interactive processing device for physical books includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the interactive processing device for physical books, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Image acquisition and image recognition are performed on pages of physical books to obtain book data and generate interactive excerpt pages; Obtain the extraction element corresponding to the extraction instruction submitted by the user based on the extraction interaction page and upload it to the server; The extracted elements are rendered to obtain the corresponding comment interaction page, and the comment data entered on the comment interaction page is collected and uploaded to the server. The annotations are displayed in association with the book location information issued by the server and the annotation data; the book location information is obtained by book mapping retrieval based on the retrieval data containing the excerpt elements and the annotation data.
[0166] Another embodiment of the interactive processing device for physical books provided in this specification is as follows: Corresponding to the other interactive processing method for physical books described above, based on the same technical concept, one or more embodiments of this specification also provide another interactive processing device for physical books, which is used to execute the other interactive processing method for physical books provided above. Figure 13 This is a schematic diagram of the structure of another interactive processing device for a physical book provided in one or more embodiments of this specification.
[0167] This embodiment provides an interactive processing device for physical books, including: like Figure 13As shown, device 1300 mainly consists of a communication interface 1302, a user interface 1304, a processor 1306, and a data storage 1308. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1310. The communication interface 1302 enables device 1300 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1302 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1302 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1302 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1302 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1304 includes receiving user input and providing output to the user. Therefore, user interface 1304 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1304 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1304 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 1300 may support remote access from other devices via communication interface 1302 or another physical interface (not shown). User interface 1304 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1304 may also be configured as a display device for rendering or displaying text fragments.
[0168] Processor 1306 may include one or more general-purpose processors and / or special-purpose processors. Data storage 1308 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1306. Data storage 1308 may include removable and non-removable components.
[0169] Processor 1306 is capable of executing program instructions 1318 (e.g., compiled or uncompiled program logic and / or machine code) stored in data store 1308 to perform the various functions described herein. Data store 1308 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1300, enable device 1300 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1318 by processor 1306 may result in processor 1306 using data 1312. For example, program instructions 1318 may include an operating system 1322 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1300 and one or more application programs 1320 (e.g., a browser, social application, or game application). Similarly, data 1312 may include operating system data 1316 and application data 1314. Operating system data 1316 is primarily accessible to operating system 1322, while application data 1314 is primarily accessible to one or more application programs 1320. Application data 1314 may reside in a file system visible or hidden to the user of device 1300. Application 1320 may communicate with operating system 1312 via one or more application programming interfaces (APIs). These APIs facilitate application 1320 reading and / or writing application data 1314, transmitting or receiving information via communication interface 1302, receiving or displaying information on user interface 1304, etc. In some terms, application 1320 may be simply referred to as an "app". Furthermore, application 1320 may be downloaded to device 1300 through one or more online app stores or app markets. However, applications may also be installed on device 1300 in other ways, such as through a web browser or a physical interface on device 1300 (e.g., a USB port).
[0170] In one specific embodiment, the interactive processing device for physical books includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the interactive processing device for physical books, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Receive excerpt elements uploaded by the application; the excerpt elements are submitted based on the excerpt interaction page; the excerpt interaction page is generated based on book data obtained by image acquisition and image recognition of book pages of physical books; The system obtains the annotation data entered by the user on the annotation interaction page, and performs a book mapping retrieval based on the search data containing the excerpt element and the annotation data to obtain book location information; the annotation interaction page is generated based on the excerpt element. The book's location information is sent to the application for associated annotation display.
[0171] This specification provides an embodiment of a computer-readable storage medium as follows: Corresponding to the interactive processing method for physical books described above, and based on the same technical concept, one or more embodiments of this specification also provide a computer-readable storage medium.
[0172] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, implement the following process: Image acquisition and image recognition are performed on pages of physical books to obtain book data and generate interactive excerpt pages; Obtain the extraction element corresponding to the extraction instruction submitted by the user based on the extraction interaction page and upload it to the server; The extracted elements are rendered to obtain the corresponding comment interaction page, and the comment data entered on the comment interaction page is collected and uploaded to the server. The annotations are displayed in association with the book location information issued by the server and the annotation data; the book location information is obtained by book mapping retrieval based on the retrieval data containing the excerpt elements and the annotation data.
[0173] It should be noted that the embodiments of a computer-readable storage medium described in this specification and the embodiments of an interactive processing method for a physical book described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0174] Another embodiment of a computer-readable storage medium provided in this specification is as follows: Corresponding to the other interactive processing method for physical books described above, and based on the same technical concept, one or more embodiments of this specification also provide another computer-readable storage medium.
[0175] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, implement the following process: Receive excerpt elements uploaded by the application; the excerpt elements are submitted based on the excerpt interaction page; the excerpt interaction page is generated based on book data obtained by image acquisition and image recognition of book pages of physical books; The system obtains the annotation data entered by the user on the annotation interaction page, and performs a book mapping retrieval based on the search data containing the excerpt element and the annotation data to obtain book location information; the annotation interaction page is generated based on the excerpt element. The book's location information is sent to the application for associated annotation display.
[0176] It should be noted that the embodiments of another computer-readable storage medium described in this specification and the embodiments of another interactive processing method for physical books described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0177] This specification provides an example of a computer program product as follows: Corresponding to the interactive processing method for physical books described above, and based on the same technical concept, one or more embodiments of this specification also provide a computer program product.
[0178] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: Image acquisition and image recognition are performed on pages of physical books to obtain book data and generate interactive excerpt pages; Obtain the extraction element corresponding to the extraction instruction submitted by the user based on the extraction interaction page and upload it to the server; The extracted elements are rendered to obtain the corresponding comment interaction page, and the comment data entered on the comment interaction page is collected and uploaded to the server. The annotations are displayed in association with the book location information issued by the server and the annotation data; the book location information is obtained by book mapping retrieval based on the retrieval data containing the excerpt elements and the annotation data.
[0179] It should be noted that the embodiments of a computer program product described in this specification and the embodiments of an interactive processing method for a physical book described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0180] Another example of a computer program product provided in this specification is as follows: Corresponding to the other interactive processing method for physical books described above, and based on the same technical concept, one or more embodiments of this specification also provide another computer program product.
[0181] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: Receive excerpt elements uploaded by the application; the excerpt elements are submitted based on the excerpt interaction page; the excerpt interaction page is generated based on book data obtained by image acquisition and image recognition of book pages of physical books; The system obtains the annotation data entered by the user on the annotation interaction page, and performs a book mapping retrieval based on the search data containing the excerpt element and the annotation data to obtain book location information; the annotation interaction page is generated based on the excerpt element. The book's location information is sent to the application for associated annotation display.
[0182] It should be noted that the embodiments of another computer program product described in this specification and the embodiments of another physical book interaction processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0183] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, please refer to each other. Each embodiment focuses on describing the differences from other embodiments. For example, the device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments are all similar to the method embodiments, so the descriptions are relatively simple. For reading the relevant content of the device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments, please refer to the description of the method embodiments.
[0184] While one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many steps, and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims. This specification uses specific terms to describe embodiments of this specification. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0185] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0186] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement to the methodology cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0187] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0188] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0189] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0190] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0191] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0192] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0193] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0194] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0195] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0196] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0197] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising at least one…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0198] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0199] The above description is merely an embodiment of this document and is not intended to limit the scope of this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.
Claims
1. An interactive processing method of a physical book, applied to an application program, the method comprising: image acquisition and image recognition on a book page of the physical book to obtain book data and generate an excerpt interactive page; acquiring an excerpt element corresponding to an excerpt instruction submitted by a user based on the excerpt interactive page and uploading to a server; rendering the excerpt element and obtaining a corresponding annotation interactive page, collecting annotation data input in the annotation interactive page and uploading to the server; based on the book positioning information issued by the server and the annotation data, performing annotation association display; the book positioning information is obtained by book mapping retrieval according to search data containing the excerpt element and the annotation data.
2. The interactive processing method of the physical book according to claim 1, after the step of acquiring the excerpt element corresponding to the excerpt instruction submitted by the user based on the excerpt interactive page and uploading to the server is executed, and before the step of rendering the excerpt element and obtaining the corresponding annotation interactive page, collecting the annotation data input in the annotation interactive page and uploading to the server is executed, the method further comprises: uploading the book page image obtained by the image acquisition and image recognition to the server; receiving the book positioning information obtained by the book mapping retrieval issued by the server; the book mapping retrieval is performed according to the book page image and the excerpt element contained in the search data.
3. The interactive processing method of the physical book according to claim 2, the book mapping retrieval is implemented in the following way: topological feature extraction and structural feature extraction are performed on the book page image to obtain character topological features and page structure features; based on the page structure features, determine the associated pages of the candidate book, and perform feature matching according to the character topological features, the page structure features, the key element features of the excerpt element and the associated page features corresponding to the associated pages to obtain the book positioning information.
4. The interactive processing method of the physical book according to claim 2, the book mapping retrieval is implemented in the following way: reference element recognition and feature extraction are performed on the book page image to obtain reference features, and reference book retrieval is performed according to the reference features to obtain reference positioning information of the reference book as the book positioning information.
5. The interactive processing method of the physical book according to claim 1, the book mapping retrieval according to the search data containing the excerpt element and the annotation data comprises: performing search feature extraction on the excerpt element and the annotation data to obtain excerpt search features and annotation search features, and performing semantic feature extraction on the excerpt element to obtain semantic features; based on the semantic features, the annotation search features and the excerpt search features, perform book mapping retrieval in the candidate book to obtain the book positioning information.
6. The interactive processing method of the physical book according to claim 5, the book mapping retrieval based on the semantic features, the annotation search features and the excerpt search features in the candidate book to obtain the book positioning information comprises: inputting the semantic features, the comment retrieval features and the excerpt retrieval features into a retrieval model and obtaining respective corresponding retrieval result sets; assigning respective corresponding confidence weights to the retrieval result sets and performing weighted fusion and sorting on the retrieval result sets based on the confidence weights to obtain a mapping retrieval list; determining the book positioning information according to the mapping retrieval list.
7. The entity book interaction processing method according to claim 5, wherein the comment association display based on the book positioning information issued by the server and the comment data comprises: if multiple candidate book positioning information issued by the server is detected, displaying each candidate book positioning information on the interactive page of the application program; based on the selection instruction submitted by the user, determining corresponding candidate book positioning information as the book positioning information and performing comment association display with the comment data.
8. The entity book interaction processing method according to claim 1, after the operation of collecting the comment data input on the comment interactive page and uploading to the server, further comprising: performing encoding processing on the comment data to obtain comment encoding data, and performing sentiment feature extraction on the comment encoding data to obtain comment sentiment features; performing sentiment label mapping on the comment sentiment features to obtain sentiment labels, and determining recommended resources according to the sentiment labels.
9. The entity book interaction processing method according to claim 1, wherein the excerpt interactive page is generated in the following manner: displaying a book page image on the interactive page of the application program, and establishing a coordinate mapping relationship between the book data and corresponding image element regions in the book page image; based on the coordinate mapping relationship, rendering the book data on an upper layer of the interactive page to obtain the excerpt interactive page.
10. The entity book interaction processing method according to claim 9, wherein the obtaining of the excerpt element corresponding to the excerpt instruction submitted by the user based on the excerpt interactive page comprises: performing element extraction on the book data according to at least one excerpt range corresponding to the excerpt instruction to obtain the excerpt element; based on the coordinate mapping relationship, determining an excerpt image element region corresponding to the excerpt element in the book page image, and synchronously marking the excerpt element and the excerpt image element region.
11. The entity book interaction processing method according to claim 1, wherein the image collection and image recognition of the book page of the entity book comprises: performing taking feature extraction on a book page taking image, and determining an image collection mode based on the obtained taking features; performing parameter configuration according to collection constraint parameters corresponding to the image collection mode and collecting a book page image of the book page, and performing image recognition on the book page image to obtain the book data.
12. The method of claim 11, after the operation of parameter configuring and collecting the book page image according to the collection constraint parameter corresponding to the image collection mode, and before the operation of obtaining the book data by image recognition on the book page image, further comprising: detecting and calculating image parameters of the book page image to obtain an image parameter evaluation value; if the image parameter evaluation value meets a corresponding image parameter condition, performing the operation of obtaining the book data by image recognition on the book page image; if the image parameter evaluation value does not meet the corresponding image parameter condition, generating a corresponding guide instruction and performing secondary image collection on the book page.
13. The method of claim 1, further comprising: displaying an associated storage list according to an associated interaction instruction submitted by the user on an interaction page; the associated storage list storing associated data obtained by associating the excerpt element with the book positioning information and / or the annotation data; wherein the association is achieved in the following manner: constructing an association between the excerpt element and the book positioning information and / or the annotation data and generating the associated data; and storing the associated data in a corresponding associated storage list based on the book positioning information.
14. An interactive processing method of a physical book, applied to a server, the method comprising: receiving an excerpt element uploaded by an application; the excerpt element being submitted based on an excerpt interaction page, the excerpt interaction page being generated based on book data obtained by image collection and image recognition on a book page of a physical book; obtaining annotation data input by a user on an annotation interaction page generated based on the excerpt element, performing book mapping retrieval based on retrieval data containing the excerpt element and the annotation data to obtain book positioning information; and issuing the book positioning information to the application for annotation association display.
15. The method of claim 14, wherein the book mapping retrieval based on the retrieval data containing the excerpt element and the annotation data to obtain the book positioning information comprises: performing topological feature extraction and structural feature extraction on a book page image uploaded by the application to obtain character topological features and page structural features; determining an associated page of a candidate book based on the page structural features, and performing feature matching based on the character topological features, the page structural features, and key element features of the excerpt element and associated page features of the associated page to obtain the book positioning information.
16. An interactive processing device of a physical book, running in an application, the device comprising: a book data acquisition module configured to perform image collection and image recognition on a book page of a physical book to obtain book data and generate an excerpt interaction page; and a text excerpt module configured to obtain an excerpt element corresponding to an excerpt instruction submitted by a user based on the excerpt interaction page and upload the excerpt element to a server. The comment data submission module is configured to render the excerpt element, obtain a corresponding comment interactive page, collect comment data input in the comment interactive page, and upload the comment data to the server; The excerpt interface display module is configured to display the comment in association with the comment data based on book positioning information issued by the server; the book positioning information is obtained by book mapping retrieval based on retrieval data containing the excerpt element and the comment data. 17.An entity book interactive processing apparatus, running on a server, comprising: An excerpt element receiving module configured to receive an excerpt element uploaded by an application program; The excerpt element is submitted based on an excerpt interactive page generated based on book data obtained by image collection and image recognition on a book page of an entity book; A book mapping retrieval module is configured to obtain comment data input by a user in a comment interactive page generated based on the excerpt element, and obtain book positioning information by book mapping retrieval based on retrieval data containing the excerpt element and the comment data; A book positioning information issuing module is configured to issue the book positioning information to the application program for comment association display. 18.An entity book interactive processing device, comprising: A processor; and a memory configured to store computer executable instructions which, when executed, cause the processor to: collect and recognize images of a book page of an entity book, obtain book data, and generate an excerpt interactive page; obtain an excerpt element corresponding to an excerpt instruction submitted by a user based on the excerpt interactive page and upload the excerpt element to a server; render the excerpt element, obtain a corresponding comment interactive page, collect comment data input in the comment interactive page, and upload the comment data to the server; display the comment in association with the comment data based on book positioning information issued by the server; The book positioning information is obtained by book mapping retrieval based on retrieval data containing the excerpt element and the comment data. 19.An entity book interactive processing device, comprising: A processor; and a memory configured to store computer executable instructions which, when executed, cause the processor to: receive an excerpt element uploaded by an application program; The excerpt element is submitted based on an excerpt interactive page generated based on book data obtained by image collection and image recognition on a book page of an entity book; obtain comment data input by a user in a comment interactive page, and obtain book positioning information by book mapping retrieval based on retrieval data containing the excerpt element and the comment data; the comment interactive page is generated based on the excerpt element; issue the book positioning information to the application program for comment association display. 20.A computer readable storage medium for storing computer executable instructions which, when executed, implement the steps of the method of claim 1 or 14.