Browser reading mode configuration method and device

By applying OCR recognition technology in the browser, analyzing and generating reading mode layout templates, the problem that traditional browser reading mode cannot be deeply analyzed and personalized adjustments is solved, and a more efficient and personalized reading experience is achieved.

CN120011666APending Publication Date: 2025-05-16SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510059503.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The traditional browser reading mode lacks the ability to deeply analyze and personalize the adjustment of text content, and cannot effectively solve the user's reading difficulties in information overload.

Method used

Through OCR recognition technology, the reading mode layout template is analyzed and generated, the directory structure is dynamically obtained and created, key nodes are analyzed and coordinates are recorded, and the page content is realized in-depth analysis and personalized display.

Benefits of technology

It improves the ability to parse website pages with different technologies and layouts, enhances the reading experience, and can better adapt to and adapt to users' reading habits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011666A_ABST
    Figure CN120011666A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of Web application programs, and particularly provides a browser reading mode configuration method and device, which is based on OCR (Optical Character Recognition) identification and comprises the following steps: S1, after a reading mode is started, waiting for page content loading completion through an onload event of a window object; s2, analyzing and obtaining the content of the page; s3, generating a text content picture required by OCR (Optical Character Recognition); s4, recognizing the text content picture to obtain a text; s5, creating a reading mode layout template; s6, dynamically acquiring directory content by using JavaScript and creating a directory structure; s7, dynamically generating a link in the reading mode panel; and S8, adding reading progress storage, and monitoring and intercepting the http request to realize management of request resources. Compared with the prior art, the reading experience can be improved, and the system browser can better adapt to and adapt to the reading habit of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Web application programs, and specifically provides a browser reading mode configuration method and device. Background Art

[0002] With the rapid development of the Internet, web content has become increasingly rich and complex. Users are faced with the problem of information overload when browsing the web, which leads them to need a more efficient and comfortable way to read online content.

[0003] Optical character recognition (OCR) technology is a technology that can convert text in an image into machine-encoded text. It is widely used in document digitization, automated data input, and other fields. Combining OCR technology can apply a new approach to capture page content and effectively conduct in-depth analysis of text content.

[0004] Although traditional browser reading modes can simplify page layouts and remove advertisements and unnecessary elements, they generally lack the ability to conduct in-depth analysis and personalized adjustment of text content. Therefore, how to solve the traditional browser reading mode is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention

[0005] The present invention aims at solving the above-mentioned deficiencies of the prior art and provides a browser reading mode configuration method with strong practicality.

[0006] A further technical task of the present invention is to provide a browser reading mode configuration device that is reasonably designed, safe and applicable.

[0007] The technical solution adopted by the present invention to solve its technical problem is:

[0008] A browser reading mode configuration method, based on OCR recognition, has the following steps:

[0009] S1. After turning on the reading mode, wait for the page content to be loaded through the onload event of the window object;

[0010] S2, analyzing and obtaining the content of the page;

[0011] S3, generating text content images required for OCR recognition;

[0012] S4, using online or local OCR to recognize the text content image to obtain the text;

[0013] S5. Create a reading mode layout template;

[0014] S6. If the OCR recognized content contains keywords, use JavaScript to dynamically obtain the directory content and create a directory structure;

[0015] S7, OCR parses key nodes and records coordinates, and dynamically generates links in the reading mode panel;

[0016] S8. Add reading progress storage, monitor and intercept http requests to manage request resources.

[0017] Furthermore, in step S1, first, a NodeLayer overlay is dynamically generated using HTML5 and JavaScript, and the layer displays a prompt message "Entering reading mode", and the animation effect is defined by the @keyframes rule of CSS3.

[0018] Furthermore, in step S2, the page content is analyzed and obtained, and when the text is recognized, the DOM tree is parsed, and a weighted score is calculated according to the text features, and the corresponding weight is assigned to the parent DOM node. Finally, the node with the highest weight is found to be the text node.

[0019] Further, in step S3, Canvas provides an API for HTML5, uses the GPU capability to capture page elements to generate canvas elements, and partially intercepts the page coordinates of the text node obtained in step S2;

[0020] Use the toDataURL method of Canvas to convert the format of the obtained canvas element to generate base64-encoded image data.

[0021] Further, in step S4, the online OCR recognition uses RagFlow or OpenCV; the local OCR recognition is provided by the tesseract.js library, and the model inference is performed on the browser side using WebAssembly technology;

[0022] Perform image preprocessing for OCR recognition. After OCR recognition, the key text and coordinate information of the main content can be obtained;

[0023] In the user preferences, choose to use online OCR recognition or offline OCR recognition, and update the offline OCR recognition function version.

[0024] Furthermore, in step S5, a reading mode layout template is dynamically generated using HTML and CSS based on the text content recognized by OCR. The reading mode board will replace the original NodeLayer overlay layer for display, and hide the display of the original page content.

[0025] Further, in step S6, when the user selects a directory, the directory DOM node in the hidden page is obtained according to the directory node information coordinates parsed by OCR to perform a trigger operation, and finally realize the page anchor jump or page link jump;

[0026] The table of contents is displayed vertically on the left side of the reading mode panel and the current chapter item is highlighted, or you can use the drop-down list at the bottom of the text to select and switch, and use the fetch API to preload the content of the next chapter.

[0027] Furthermore, in step S7, a link is dynamically generated in the reading mode panel, and an event is bound using javascript, so that when the link is clicked, the link of the original page is triggered and the new page content is loaded;

[0028] When loading, a new page container will be added to the current reading panel and switched;

[0029] When the user clicks the chapter-turning button, the corresponding page loading and processing is triggered. When maintaining a node coordinate mapping table, the corresponding page loading and processing is triggered according to the user operation.

[0030] Furthermore, in step S8, when adding reading progress storage, in reading mode, the user's page turning, chapter switching, and pull-up and pull-down operations are monitored, and the current link, chapter, and position are recorded and stored in the localstorage space of the browser;

[0031] When monitoring and intercepting http requests to manage requested resources, set up black and white lists to filter and manage the request content.

[0032] A browser reading mode configuration device, comprising: at least one memory and at least one processor;

[0033] The at least one memory is used to store a machine-readable program;

[0034] The at least one processor is used to call the machine-readable program to execute a browser reading mode configuration method.

[0035] Compared with the prior art, the browser reading mode configuration method and device of the present invention have the following outstanding beneficial effects:

[0036] The present invention can more accurately and efficiently parse website pages that use different technologies and formats to display content, improve reading experience, and facilitate system browsers to better adapt to and adapt to users' reading habits. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0038] Attached Figure 1 The invention is a flowchart of a method for configuring a browser reading mode. DETAILED DESCRIPTION

[0039] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention is further described in detail below in conjunction with specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] A best embodiment is given below:

[0041] like Figure 1 As shown, a browser reading mode configuration method in this embodiment has the following steps:

[0042] S1. After turning on the reading mode, wait for the page content to be loaded through the onload event of the window object;

[0043] When the user enables reading mode in the browser, a NodeLayer overlay is first dynamically generated using HTML5 and JavaScript. This overlay displays the prompt message "Entering reading mode". At the same time, the animation effect is defined through the @keyframes rule of CSS3, so that the transition layer appears or disappears with animation effects.

[0044] S2, analyzing and obtaining the content of the page;

[0045] Text recognition is one of the most important functions for realizing reading mode. It can be realized by using algorithms such as just-read and Readability.

[0046] The logic of the recognition algorithm is to parse the DOM tree and calculate the weighted score based on the text features, where the text features include text length, text link density, punctuation marks, etc., and assign corresponding weights to the parent DOM nodes. Finally, the node with the highest weight is found to be the text node.

[0047] S3, generating text content images required for OCR recognition;

[0048] Canvas is an API provided by HTML5. It can use the GPU to capture page elements to generate canvas elements and partially intercept the page coordinates of the text node obtained in the previous step.

[0049] Use the toDataURL method of Canvas to convert the format of the obtained canvas element and generate base64-encoded image data.

[0050] S4, using online or local OCR to recognize the text content image to obtain the text;

[0051] Online OCR recognition uses RagFlow or OpenCV and other OCR engines to build services;

[0052] Local OCR recognition is provided by the tesseract.js library, using WebAssembly technology to perform model inference on the browser side;

[0053] In order to improve the accuracy of OCR, OCR recognition can perform image preprocessing, such as denoising, binarization, contrast adjustment, etc.

[0054] After OCR recognition, the key text and coordinate information of the main content can be obtained;

[0055] You can choose to use online OCR recognition or offline OCR recognition in user preferences, and update the version of offline OCR recognition function.

[0056] S5. Create a reading mode layout template;

[0057] Based on the text content recognized by OCR, HTML and CSS are used to dynamically generate a reading mode layout template. The template can be a single-column or multi-column layout, using horizontal or vertical page turning, depending on user preference, to ensure that the text is clear and easy to read.

[0058] The reading mode dashboard will replace the original NodeLayer overlay for display, and hide the original page content;

[0059] The reading mode panel uses CSS3 media queries to achieve responsive design, ensuring that the template displays consistently on different devices.

[0060] Provides a variety of preset themes and color schemes. Users can freely choose according to their preferences in the reading preferences, and use CSS variables to switch themes;

[0061] It also allows users to customize font style, font size, line spacing, etc. to enhance the personalized experience;

[0062] S6. If the OCR recognized content contains keywords, use JavaScript to dynamically obtain the directory content and create a directory structure;

[0063] When the user selects a directory, the directory DOM node in the hidden page will be obtained according to the directory node information coordinates parsed by OCR to perform a trigger operation, and finally realize the page anchor jump or page link jump;

[0064] The contents of the table of contents can be displayed vertically on the left side of the reading mode panel and the current chapter item can be highlighted, or you can use the drop-down list at the bottom of the text to select and switch;

[0065] The system can use the fetch API to preload the content of the next chapter to improve page turning speed and reading experience.

[0066] S7, OCR parses key nodes and records coordinates, and dynamically generates links in the reading mode panel;

[0067] OCR analyzes key nodes such as the previous chapter or the next chapter and records the coordinates. In the reading mode panel, a link to the previous chapter or the next chapter is dynamically generated, and JavaScript is used to bind the event so that when the link is clicked, the link to the previous chapter or the next chapter of the original page is triggered and the new page content is loaded;

[0068] When loading, a new page container will be added to the current reading panel and switched;

[0069] After the user clicks the chapter-turning button, the corresponding page loading and processing is triggered. A node coordinate mapping table is maintained to trigger the corresponding page loading and processing according to the user operation.

[0070] S8. Add reading progress storage, monitor and intercept http requests to manage request resources;

[0071] When adding reading progress storage, listen to the user's page turning, chapter switching, pull-up and pull-down operations, and record the current link, chapter, and position, and store them in the browser's localstorage space, so that you can continue reading according to the last progress when you open the page next time.

[0072] Listen to and intercept http requests to manage requested resources, and set up black and white lists based on this. However, when the page you open contains ads from third-party platforms or other requested resources that users don't like, you can set up black and white list addresses through reading preferences to filter and manage request content.

[0073] Based on the above method, a browser reading mode configuration device in this embodiment includes: at least one memory and at least one processor;

[0074] The at least one memory is used to store a machine-readable program;

[0075] The at least one processor is used to call the machine-readable program to execute a browser reading mode configuration method.

[0076] The above-mentioned specific implementations are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementations. Any technical solutions that conform to the above-mentioned specific implementations of the present invention and any appropriate changes or substitutions made by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.

[0077] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for configuring a browser reading mode, characterized in that: Based on OCR recognition, there are the following steps: S1. After turning on the reading mode, wait for the page content to be loaded through the onload event of the window object; S2, analyzing and obtaining the content of the page; S3, generating text content images required for OCR recognition; S4, using online or local OCR to recognize the text content image to obtain the text; S5. Create a reading mode layout template; S6. If the OCR recognized content contains keywords, use JavaScript to dynamically obtain the directory content and create a directory structure; S7, OCR parses key nodes and records coordinates, and dynamically generates links in the reading mode panel; S8. Add reading progress storage, monitor and intercept http requests to manage requested resources.

2. A browser reading mode configuration method according to claim 1, characterized in that: In step S1, first, a NodeLayer overlay is dynamically generated using HTML5 and JavaScript, and the layer displays a prompt message "Entering reading mode", and the animation effect is defined by the @keyframes rule of CSS3.

3. A browser reading mode configuration method according to claim 2, characterized in that: In step S2, the page content is analyzed and obtained, and when the text is recognized, the DOM tree is parsed, and a weighted score is calculated according to the text features, and the corresponding weight is assigned to the parent DOM node. Finally, the node with the highest weight is found to be the text node.

4. A browser reading mode configuration method according to claim 3, characterized in that: In step S3, Canvas provides an API for HTML5, uses the GPU capability to capture page elements to generate canvas elements, and partially intercepts the page coordinates of the text node obtained in step S2; Use the toDataURL method of Canvas to convert the format of the obtained canvas element to generate base64-encoded image data.

5. A browser reading mode configuration method according to claim 4, characterized in that: In step S4, online OCR recognition uses RagFlow or OpenCV; the local OCR recognition is provided by the tesseract.js library, and the model inference is performed on the browser side using WebAssembly technology; Perform image preprocessing for OCR recognition. After OCR recognition, the key text and coordinate information of the main content can be obtained; In the user preferences, choose to use online OCR recognition or offline OCR recognition, and update the offline OCR recognition function version.

6. A browser reading mode configuration method according to claim 5, characterized in that: In step S5, based on the text content recognized by OCR, a reading mode layout template is dynamically generated using HTML and CSS. The reading mode dashboard will replace the original NodeLayer overlay layer for display, and hide the original page content.

7. A browser reading mode configuration method according to claim 6, characterized in that: In step S6, when the user selects a directory, the directory DOM node in the hidden page is obtained according to the directory node information coordinates parsed by OCR to perform a trigger operation, and finally realize the page anchor jump or page link jump; The table of contents is displayed vertically on the left side of the reading mode panel and the current chapter item is highlighted, or you can use the drop-down list at the bottom of the text to select and switch, and use the fetch API to preload the content of the next chapter.

8. A browser reading mode configuration method according to claim 7, characterized in that: In step S7, a link is dynamically generated in the reading mode panel, and an event is bound using javascript, so that when the link is clicked, the link of the original page is triggered and the new page content is loaded; When loading, a new page container will be added to the current reading panel and switched; When the user clicks the chapter-turning button, the corresponding page loading and processing is triggered. When maintaining a node coordinate mapping table, the corresponding page loading and processing is triggered according to the user operation.

9. A browser reading mode configuration method according to claim 8, characterized in that: In step S8, when adding reading progress storage, in reading mode, the user's page turning, chapter switching, and pull-up and pull-down operations are monitored, and the current link, chapter, and position are recorded and stored in the local storage space of the browser; When monitoring and intercepting http requests to manage requested resources, set up black and white lists to filter and manage the request content.

10. A browser reading mode configuration device, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 9.