Screen reading method and device, electronic equipment and readable storage medium

By receiving user input, determining the target summary generation mode, and using an AI model to process the screen-read text, the summary information is generated and played, solving the problem of long reading time for long articles and achieving the effect of quickly obtaining key information.

CN121907957APending Publication Date: 2026-04-21VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2026-02-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing screen reading applications require users to spend a considerable amount of time listening to long texts, resulting in low information acquisition efficiency and failing to meet the need for quickly obtaining core information.

Method used

By receiving user input, the target summary generation mode is determined, and the text resources are processed based on the mode to generate summary text information. An AI model is used for semantic understanding and logical organization, and finally the summary information is played back by voice.

Benefits of technology

By compressing the full-text reading time from about 15 minutes to a shorter time, and avoiding the redundant reception of invalid information, users can efficiently obtain key information within a limited time, thus improving information acquisition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121907957A_ABST
    Figure CN121907957A_ABST
Patent Text Reader

Abstract

The invention discloses a screen reading method and device, electronic equipment and a readable storage medium, and belongs to the technical field of electronic equipment. The method comprises the steps of receiving first input under the condition that a display screen reads text resources; in response to the first input, determining a target abstract generation mode and a target text range corresponding to the target abstract generation mode; the target abstract generation mode comprises at least one of abstract generation according to text paragraphs, abstract generation according to text fragments and abstract generation according to full texts, and the text fragments are determined based on a text content structure of a screen reading text resource; based on the target abstract generation mode, processing the text resources in the target text range to obtain abstract text information of the text resources; and playing the abstract text information through voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic equipment technology, specifically relating to a screen reading method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] With the rapid development of the mobile internet, electronic devices have become the primary medium for people to access information, and various types of content, such as news reports and professional articles, are increasingly widely disseminated on electronic devices. To meet users' needs for accessing information in multiple scenarios, screen-to-speech applications are now commonly found on electronic devices. These applications can automatically read the content displayed on the screen and read it aloud in full. This feature greatly facilitates users' daily use.

[0003] However, despite the excellent performance of existing screen reading applications in multitasking, the current screen reading mode often requires users to spend about 15 minutes listening to long texts with thousands or even tens of thousands of words, resulting in low information acquisition efficiency. Summary of the Invention

[0004] The purpose of this application is to provide a screen reading method, apparatus, electronic device, and readable storage medium that can improve information acquisition efficiency.

[0005] In a first aspect, embodiments of this application provide a screen reading method, including: When displaying text resources for reading aloud on the screen, receive the first input; In response to the first input, a target summary generation mode and a target text range corresponding to the target summary generation mode are determined; the target summary generation mode includes at least one of summarizing by text paragraph, summarizing by text fragment, and summarizing by full text, wherein the text fragment is determined based on the text content structure of the screen-reading text resource; Based on the target summary generation mode, text resources within the target text range are processed to obtain summary text information of the text resources; The audio plays a summary of the text information.

[0006] Secondly, embodiments of this application provide a screen reading device, including: The receiving module is used to receive the first input when the text resource is read aloud on the display screen; The determination module is used to determine the target summary generation mode and the target text range corresponding to the target summary generation mode in response to the first input; the target summary generation mode includes at least one of summarizing by text paragraph, summarizing by text fragment, and summarizing by full text, wherein the text fragment is determined based on the text content structure of the screen-reading text resource; The processing module is used to process text resources within the target text range based on the target summary generation mode to obtain summary text information of the text resources; The playback module is used to play summary text information via voice.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in the first aspect.

[0009] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface, the communication interface and the processor being coupled together, the processor being used to run programs or instructions to implement the steps of the method as described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the method as described in the first aspect.

[0011] This embodiment determines the target summary generation mode and the target text range corresponding to the target summary generation mode based on the received first input. Then, it processes the text information within the target text range based on the target summary generation mode to obtain summary text information, which is then played back by voice. This achieves the condensation of text information, compressing the original long full-text reading time into a short time, avoiding the redundant reception of invalid information, and enabling users to efficiently obtain key information within a limited time, thus improving information acquisition efficiency. Attached Figure Description

[0012] Figure 1 A flowchart illustrating a screen reading method provided in this application embodiment; Figure 2 A schematic diagram of a playback interface provided in an embodiment of this application; Figure 3 This application provides an embodiment of an interface for synchronously playing and displaying summary text information; Figure 4 A schematic diagram of a question-and-answer interface provided in an embodiment of this application; Figure 5 A schematic diagram of another question-and-answer interface provided in an embodiment of this application; Figure 6This is a schematic diagram illustrating the display of original text location information provided in an embodiment of this application; Figure 7 This is a schematic diagram illustrating the display of different original text content using different styles, provided as an embodiment of this application. Figure 8 This is a schematic diagram of the structure of a screen reading device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0014] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.

[0015] To meet users' needs for information access in various scenarios, screen-to-speech applications are now widely available on electronic devices. These applications can automatically read the content displayed on the screen and read it aloud in full. This feature greatly facilitates users' daily use.

[0016] However, although existing screen reading applications perform well in multitasking, they currently use a full-text reading mode. For long texts with thousands or even tens of thousands of words, users often need to spend about 15 minutes to listen to them, resulting in low information acquisition efficiency. This is not compatible with the fast pace of life and work today and cannot meet users' needs for quickly obtaining core information.

[0017] Therefore, embodiments of this application provide a screen reading method, apparatus, electronic device, computer-readable storage medium, and product.

[0018] The screen reading method, apparatus, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0019] The screen reading method provided in this application can be applied to students, working professionals, readers who enjoy reading, such as those who like to read novels or essays, and self-media workers.

[0020] For example, for students, when reading long academic papers or extracurricular readings, the solutions provided in this application can help them quickly grasp the research framework and core conclusions of the articles, saving time for pre-class preparation or literature screening.

[0021] For example, for working professionals, when dealing with lengthy industry reports or competitor analyses, the solutions provided in this application can quickly help them understand the core data and conclusions of the report in a short period of time, providing efficient support for meeting preparation.

[0022] For example, for readers who enjoy reading novels or essays, the solutions provided in this application can help them quickly review the plot or core ideas of the text during short breaks.

[0023] For example, for self-media creators, when reading a large number of long reference articles in the self-media creation scenario, the solution provided in this application can quickly extract key information and greatly improve creation efficiency.

[0024] Figure 1 The flowchart illustrates a screen reading method provided in this application embodiment. This screen reading method can be applied to electronic devices. In practical applications, electronic devices may include, but are not limited to, mobile phones, tablets, notebook computers, personal digital assistants, etc.

[0025] like Figure 1 As shown, the screen reading method may include the following steps S110-S140.

[0026] S110. When the text resource is read aloud on the display screen, receive the first input.

[0027] S120. In response to the first input, determine the target summary generation mode and the target text range corresponding to the target summary generation mode.

[0028] The target summary generation mode includes at least one of the following: summary generation by text paragraph, summary generation by text fragment, and summary generation by full text. The text fragment is determined based on the text content structure of the screen-reading text resource.

[0029] S130. Based on the target summary generation mode, process the text resources within the target text range to obtain the summary text information of the text resources.

[0030] S140. Speech playback of summary text information.

[0031] This embodiment determines the target summary generation mode and the target text range corresponding to the target summary generation mode based on the received first input. Then, it processes the text information within the target text range based on the target summary generation mode to obtain summary text information, which is then played back by voice. This achieves the condensation of text information, compressing the original long full-text reading time into a short time, avoiding the redundant reception of invalid information, and enabling users to efficiently obtain key information within a limited time, thus improving information acquisition efficiency.

[0032] The above steps are explained in detail below: In S110, the screen-reading text resource is the text displayed on the screen that the voice assistant can read aloud. This text can be editable text converted from images, screenshots, etc. For example, in some embodiments, the content in an image or screenshot can be recognized and converted into text using Optical Character Recognition (OCR) technology to obtain the screen-reading text resource.

[0033] The screen-read text resource may include one or more paragraphs, or one or more chapters, depending on the user's needs. This embodiment does not limit the specific content of the screen-read text resource; for example, it may be news reports, professional articles, novels, etc.

[0034] In some embodiments, when the summarization function is activated, the screen-reading text resource can be acquired and displayed. Alternatively, the screen-reading text resource can be acquired and displayed upon receiving user input such as voice, gestures, or touch. For example, when the summarization function is activated, the text currently displayed on the screen can be used as the screen-reading text resource.

[0035] The first input is used to determine the target summary generation mode of the screen-reading text resource. The target summary generation mode can be a mode selected by the user for generating summaries. For example, the first input can include voice input or gesture input. For instance, the user can input the target summary generation mode by voice or by gesture, with different summary generation modes corresponding to different gestures.

[0036] In some embodiments, when the electronic device displays a summary generation mode selection interface, the first input can also be the user's selection of a target summary generation mode. The electronic device can then determine the user-selected summary generation mode as the target summary generation mode. For example, after acquiring screen-reading text resources, the electronic device can display a summary generation mode selection interface. This interface can include multiple candidate summary generation modes for the user to choose from. For instance, it can include three candidate summary generation modes: summarizing by text paragraphs, summarizing by text segments, and summarizing by the entire text. The user can select a suitable summary generation mode as needed, and the electronic device can then determine the user-selected summary generation mode as the target summary generation mode.

[0037] In some embodiments, the electronic device can also automatically determine the target summary generation mode based on the user's historical habits. For example, the electronic device can acquire and analyze the summary generation modes used by the user within a historical time period to determine the user's preferred summary generation mode, and determine the target summary generation mode based on the user's preferred summary generation mode. For example, if the user has the highest probability of using the full-text summary generation mode within a historical time period, then the full-text summary generation mode can be defaulted to as the target summary generation mode.

[0038] In some embodiments, after the electronic device adaptively determines the target summary generation mode, it can also display the target summary generation mode for user confirmation, so that the adopted summary generation mode can meet the user's needs as much as possible and improve the user experience.

[0039] In S120, exemplarily, the electronic device determines the target summary generation mode and the corresponding target text range based on the user's first input, providing a reliable foundation for subsequently generating accurate summaries. This embodiment determines the target summary generation mode and the corresponding target text range based on user input, improving the accuracy of the summary generation mode and thus enabling the generated summary text information to better meet the user's actual needs.

[0040] The target summary generation mode may include at least one of the following: summary generation by text paragraph, summary generation by text segment, and summary generation by full text. The target summary generation mode may include one or more of these modes.

[0041] Summarizing by paragraph means generating a separate summary for each independent paragraph, allowing users to quickly grasp the key information of each paragraph. Summarizing by fragment means generating a separate summary for each text fragment. These fragments are determined based on the text content structure of the screen-reading text resource; for example, they could be a chapter, a section, or fragments divided in other ways. Summarizing by fragment allows users to quickly understand the key information of each fragment. Summarizing the entire text means extracting the essence of the entire article and generating a unified summary covering the core themes, enabling users to quickly grasp the main points of the entire text.

[0042] This embodiment supports multiple summary generation modes. Users can select the appropriate summary generation mode as needed to generate corresponding summary text to meet different scenario requirements.

[0043] In some embodiments, the target summary generation mode may include summarizing by text paragraph and summarizing by the whole text. In this way, users can quickly understand the key information of each paragraph and also quickly understand the core theme of the whole text, thus meeting a variety of user needs.

[0044] The target text range can be part or all of the screen-read text resource, depending on the target summary generation mode. For example, if the target summary generation mode is to generate summaries by text paragraph, the target text range can be each paragraph. If the target summary generation mode is to generate summaries by text fragments, the target text range can be text fragments such as chapters or sections. If the target summary generation mode is to generate summaries by the entire text, the target text range can be the entire text.

[0045] In S130, in this embodiment, the electronic device can process the text resources within the target text range corresponding to the target summary generation mode in the screen-reading text resources to obtain summary text information of the text resources. The text resources here may include information such as titles, punctuation, text content, and logical connectors. For example, the electronic device can input the extracted text resources within the target text range into an Artificial Intelligence (AI) model. Using the AI ​​model, based on the target summary generation mode, it can perform semantic understanding, logical organization, and core information extraction on the text resources within the target text range to obtain summary text information. The AI ​​model can be, for example, a local lightweight model, a cloud-based high-precision model, or a combination of both.

[0046] For example, when generating summaries by text paragraph, the AI ​​model focuses on the key information of a single paragraph; when generating summaries by text fragments, the AI ​​model takes into account the logical connections and core viewpoints within the text fragments; and when generating summaries by the full text, the AI ​​model focuses on the main idea and key data of the entire text.

[0047] In some embodiments, users can also preset the length of the summary text, for example, it can be set to short, medium, or long, where short, medium, and long correspond to different word counts. In some embodiments, different target summary generation modes can also correspond to different lengths. For example, when generating summaries by text paragraphs, the length of the summary text is short; when generating summaries by text fragments, the length of the summary text is medium; and when generating summaries by the full text, the length of the summary text is long.

[0048] In some embodiments, the AI ​​model can process text resources within the target text range based on the target summary generation mode and summary length to obtain summary text information of the corresponding length.

[0049] In S140, for example, the AI ​​model can convert the generated summary text information into a speech signal, which can then be played through an audio player on an electronic device.

[0050] In some embodiments, the electronic device may play the summary text information using default tone, speed, etc. In some embodiments, users may also set the playback speed, tone, timbre, etc. of the summary text information to meet their diverse needs.

[0051] This embodiment leverages AI technology to condense lengthy text content on the screen, extracting summary information and compressing the original 15-minute full-text reading time into a shorter period. This allows users to quickly grasp the core message and avoid redundant information intake. Whether in fragmented scenarios such as commuting or work breaks, or when there is a need to quickly filter information, it helps users efficiently obtain key content within a limited time, significantly reducing the time cost of reading long texts and improving information acquisition efficiency.

[0052] In some embodiments, the screen reading method may further include the following steps: While the summary text information is being played aloud, the summary text information is displayed simultaneously.

[0053] In some embodiments, taking the screen-read text resource as "White Paper on Pet Phased Feeding and Nutritional Needs in City A, 2025" as an example, such as... Figure 2As shown, the electronic device 200 can, by default, only play the summary text information of the screen-read text resource. The playback interface 201 includes a play button 202, a back button 203, and a forward button 204. When the user clicks the play button 202, the summary text information can be played. When the user wants to listen to a previously played message again, they can click the back button 203. When the user wants to fast forward, they can click the forward button 204.

[0054] In some embodiments, the playback interface 201 may also include a summary display control 205, which can be enabled when the user wants the electronic device 200 to display summary text information at the same time.

[0055] In some embodiments, when a user clicks the summary display control 205, such as Figure 3 As shown, the electronic device 200 can not only play the summary text information 302, but also display the summary text information 302.

[0056] In some embodiments, when displaying summary text information 302, the electronic device 200 can also differentiate the display of summary text information. For example, summary text information that has already been played can be highlighted, while summary text information that has not been played can be displayed normally. Of course, different font sizes, fonts, etc., can also be used to differentiate between the displayed summary text information that has been played and the summary text information that has not been played. Figure 3 For example, the summary text information that has been played is displayed in bold, while the summary text information that has not been played is displayed normally, making it convenient for users to listen and read at the same time.

[0057] In some embodiments, users can also switch the summary generation mode at any time to regenerate the summary.

[0058] In this embodiment, during the semantic playback of summary text information, the corresponding summary text information can be displayed synchronously, making it convenient for users to listen and read at the same time.

[0059] Taking the target summary generation mode, which includes summarizing by text paragraph or by text segment, as an example, the summary text information of the text resource includes overall summary text information and partial summary text information. The overall summary text information is used to summarize the core theme of the screen-reading text resource, while the partial summary text information is used to summarize the key information of each text segment. In this way, users can quickly grasp the key content of the whole text and quickly understand the key information of each paragraph, meeting various user needs.

[0060] In some embodiments, the screen reading method may further include the following steps: Synchronously display the index of the original text location of the text resource corresponding to the summary text information.

[0061] The original text location index of a text resource is the original text location index corresponding to the abstract text information. For example, when generating a summary by text paragraph or by text segment, in addition to displaying the abstract text information on the screen, the original text location index of the corresponding text resource can also be displayed simultaneously. The original text location index of a text resource can include at least one of the following: paragraph number, chapter title.

[0062] By synchronously displaying the index of the original text resource corresponding to the summary text information, users can not only quickly understand the key content of a certain part of the text information, but also quickly and intuitively determine the location of the text information in the original text, which facilitates subsequent searches and improves search efficiency.

[0063] In some embodiments, during the playback of the summary text information by voice, or after the playback of the summary text information has ended, the screen reading method may further include the following steps: Receive a second input, which indicates query information related to the text content of the screen-read text resource; In response to the second input, the system generates a response to the query based on the text content of the text resource read aloud on the screen. Output response information.

[0064] For example, the second input can be voice input, gesture input, text input, etc. Based on the second input, the electronic device can determine the user's query information. In some embodiments, when a user inputs query information via voice, considering that the user's voice may carry a dialect, the user's voice can be first parsed and corrected to obtain the query information. That is, in this embodiment, the voice input of the question-and-answer function supports dialect recognition and semantic error correction. Query information related to the text content of the screen-reading text resource may include, for example, "What is the author's core viewpoint on this policy?" or "What are the specific details of the cases mentioned in the text?"

[0065] Specifically, electronic devices can parse user queries, determine the corresponding text resources, call AI models to perform targeted retrieval and analysis of the corresponding text resources, generate response information for the queries, and output the response information through the electronic devices.

[0066] For example, response information can be played through an audio player, displayed as text, or output through a combination of playback and display. The specific output method can be dynamically determined by the user based on the scenario, or the electronic device can automatically identify the user's current scenario and select an appropriate output method based on that scenario. Alternatively, the electronic device can dynamically determine the output method based on the user's usage habits.

[0067] In some embodiments, such as Figure 3 As shown, the playback interface 201 may also include a Q&A button 206. After clicking the Q&A button 206, the user can ask personalized questions based on the summary text information.

[0068] For example, after receiving a user's click input on the question-and-answer button 206, the electronic device can display... Figure 4 The question-and-answer interface 401 shown allows users to input inquiry information 402. The electronic device can convert the received inquiry information 402 into text and display it on the interface 401. In some embodiments, the inquiry information 402 may include "How do you understand the upgrade from partner to family member?". After receiving the inquiry information 402, the electronic device can parse it to determine the corresponding text resource, and then call an AI model to analyze the text resource, for example, outputting... Figure 4 The response information 403 shown is for user reference. Figure 4 Taking the display of response information 403 via text as an example, in some embodiments, the response information 403 can also be played synchronously, providing users with multiple output methods to meet their needs in different scenarios.

[0069] In some embodiments, when a user clicks the Q&A button 206, the electronic device may pause the playback of the summary text information.

[0070] This embodiment, building upon the summary reading function, allows users to directly ask questions about viewpoints of interest or existing doubts within long texts. The electronic device can then invoke an AI model to generate and output responses based on the AI ​​model's understanding of the entire text. This eliminates the need for users to switch to a search engine for additional queries, effectively overcoming the limitation of existing screen reading functions that cannot interactively answer questions about content. This allows users to directly obtain targeted information while listening or reading, deepening their understanding of long texts and further enhancing the convenience of interaction.

[0071] In some embodiments, after "in response to the second input, generating response information for the query based on the text content of the screen-reading text resource," the screen-reading method may further include the following steps: Receive a third input, which indicates query information related to the original location of the response information; In response to the third input, based on the response information, full-text search and semantic matching are performed in the screen-reading text resource to determine the original text location information of the response information in the screen-reading text resource; Displays the original text location information.

[0072] For example, the third input is used to inquire about the original text location of the response information. The third input may include, but is not limited to, voice input, gesture input, and text input. The inquiry information related to the original text location of the response information may include, for example, "In which paragraph is the market data mentioned in the text?" or "In which chapter is the author's refuted viewpoint located?"

[0073] For example, an electronic device can determine query information related to the original text location of the response information based on a third input, and then invoke an AI model based on this query information. The AI ​​model performs full-text retrieval and semantic matching on the screen-reading text resource to determine the original text location information of the response information within the screen-reading text resource, and then displays it. In practical applications, there can be one or more original text location information related to the response information.

[0074] For example, when parsing screen-reading text resources, the AI ​​model can establish a mapping relationship between viewpoints or summary text information and the original content. Subsequently, based on this mapping relationship and combined with the user's inquiry information, it can accurately locate the chapters and paragraphs related to the inquiry information.

[0075] catch Figure 4 The query information shown, such as Figure 5 As shown, if a user continues to ask, "In which paragraph of the article is this viewpoint located?", the electronic device can receive and display the query information 501, "In which paragraph of the article is this viewpoint located?". Based on the user's query information 501, the electronic device can invoke an AI model to determine the location of the query information 501 in the original text and display it. For example, Figure 5 The original text location information 502 and 503 can be displayed, indicating that the above viewpoint appears in both the first paragraph of Chapter 1 and the fourth paragraph of Chapter 6.

[0076] In this embodiment, users can inquire about the location of a certain viewpoint or information in the original text. Based on the user's inquiry, the electronic device can quickly locate the viewpoint or information in the original text and display it to the user. This effectively solves the cumbersome problem of requiring users to manually search for the relevant content in related technologies, allowing users to directly jump to the original text location for precise or detailed reading, and further improving the efficiency of users in utilizing long text content.

[0077] In some embodiments, the original text location information is at least one, and the screen reading method may further include the following steps: Receive a fourth input, which is used to determine the target original text location information from the original text location information; In response to the fourth input, if the original text of the text resource is read aloud on the display screen, the user is redirected to the target original text content corresponding to the target original text location information.

[0078] The target original text location information can be one or more. For example, the fourth input can include, but is not limited to, voice input, gesture input, and touch input. For instance, in some embodiments, the target original text location information can be determined from the original text location information by touch, such as by clicking on the target original text location information.

[0079] For example, after receiving the user's fourth input, the electronic device can directly jump to the target original text content corresponding to the target original text location information for the user to view.

[0080] For example, such as Figure 5 As shown, when the user clicks on the original text location information 502, as... Figure 6 As shown, the electronic device can jump to the first paragraph of the first chapter and display the original text of the first paragraph of the first chapter 601.

[0081] In some embodiments, the original text 601 in the first paragraph of the first chapter can be displayed differently from other content outside the first paragraph of the first chapter. For example, the original text 601 in the first paragraph of the first chapter can be highlighted or displayed in other ways to facilitate users to fully understand the distribution of a certain viewpoint or information in the whole text and the contextual logic.

[0082] In some embodiments, users can click on different original text location information as needed to easily switch between and view information from multiple locations.

[0083] This embodiment can automatically jump to the target text content corresponding to the target text location information selected by the user, without the need for manual segment-by-segment searching, which greatly improves the efficiency of finding specific information.

[0084] Taking the example of having at least two original text location information, in some embodiments, displaying the original text location information includes: Display a list of locations containing at least two original text location information in a preset area of ​​the screen; the preset area may be the top of the screen or the side of the screen.

[0085] For example, this embodiment can display a list of locations for each original text position at the top or side of the screen. For instance, different color highlights, underlines, or borders can be used to differentiate the location lists for different original text positions, making it easier for users to view. For example, paragraph 2 of Chapter 3 is highlighted in yellow, and paragraph 4 of Chapter 6 is highlighted in blue. Users can click on the corresponding location in the list to jump to the corresponding paragraph and have that section highlighted.

[0086] By displaying a location list that shows the original text's location information, users can quickly access the corresponding original text information, greatly improving the efficiency of obtaining specific information.

[0087] Taking the example that there are at least two original text location information, in some embodiments, the screen reading method may further include the following steps: When the original text of a text resource is read aloud on the display screen, different display styles are used to display the original text content corresponding to different original text location information.

[0088] For example, different colors can be used to highlight, underline, border, or bold the original text content corresponding to different original text location information.

[0089] like Figure 7 As shown, the electronic device can display the original text 601 of the first paragraph of Chapter 1 and the original text 602 of the fourth paragraph of Chapter 6 using different display styles. For example, the original text 601 of the first paragraph of Chapter 1 can be displayed in bold, and the original text 602 of the fourth paragraph of Chapter 6 can be displayed in italics. This makes it easier for users to distinguish the content in different locations.

[0090] In this embodiment, when the response information is associated with multiple locations, the electronic device can use different display styles to display the original text content corresponding to different original text locations of the response information, making it convenient for users to distinguish and view the target content at different locations.

[0091] This embodiment combines summary reading, AI Q&A, and opinion positioning. On the one hand, it allows users to quickly grasp the key information of long articles, which is in line with the fast-paced life and work of modern people. On the other hand, when users capture useful information or opinions while listening, if they want to see the specific location of the content in the original text, they can use AI Q&A to achieve real-time interactive Q&A. This not only meets users' needs for in-depth exploration of specific opinions and details, but also saves the extra operation of switching to a search engine, making information interaction more direct and accurate. Furthermore, combined with the location function, when users need to view the original context of a certain viewpoint, they do not need to manually search paragraph by paragraph. They can simply trigger the location function through a question and answer to quickly jump to the corresponding chapter or paragraph. This design not only reduces the difficulty of reading long texts, but also makes it easier for users to carefully read, extract, or re-analyze key information, deeply explore the value of long texts, and solve the pain points of "not being able to remember when listening and not being able to find when looking for" in the traditional reading mode. It comprehensively solves the problems of low efficiency, weak interaction, and difficulty in location in the processing of long texts by existing screen reading functions, providing users with a more efficient, accurate, and convenient way to process long text information and improving the user experience.

[0092] For example, if students have doubts about the derivation process of a certain theory, they can directly ask "What is the basis for the derivation of this formula?" using the AI ​​Q&A function. The AI ​​model will then provide a targeted answer based on the original text. When discussing a point in a paper with classmates, they can also use AI Q&A to have the AI ​​simulate different perspectives and help expand their thinking. At the same time, they can use the point location function to quickly find the specific location of the point in the original text, making it easier to accurately cite the original arguments in the discussion.

[0093] For example, if working professionals have doubts about the feasibility of a market strategy mentioned in a report, they can use the AI ​​Q&A function to obtain an analysis of the report's content. If they want to explore the similarities and differences between the strategy and other cases, they can also use the AI ​​Q&A function to have AI list relevant viewpoints for reference, and use the viewpoint positioning function to quickly locate the detailed description of the strategy in the report, making it convenient to directly display the original text for in-depth discussion in meetings.

[0094] For example, readers who enjoy reading novels or essays can use the AI ​​Q&A function to obtain analysis of the original plot; if they want to discuss the symbolic meaning of a certain plot with other readers, the AI ​​Q&A function can simulate different interpretations, and the viewpoint positioning function can find the specific location of the plot in the text, making it easier to accurately cite the details of the original text when sharing with others.

[0095] For example, for self-media workers, the AI ​​Q&A function can answer questions related to creation and inspire creative ideas. The viewpoint positioning function makes it easy to cite original texts as creative materials, thereby improving creative efficiency.

[0096] It should be noted that the model training method provided in this application embodiment can be executed by a model training device or a processing module within that model training device for executing the model training method. This application embodiment uses the execution of the model training method by a model training device as an example to illustrate the model training device provided in this application embodiment.

[0097] Figure 8 This is a schematic diagram of the structure of a screen reading device provided in an embodiment of this application.

[0098] like Figure 8 As shown, the screen reading device 800 may include: The receiving module 801 is used to receive the first input when the text resource is read aloud on the display screen; The determination module 802 is used to determine the target summary generation mode and the target text range corresponding to the target summary generation mode in response to the first input; the target summary generation mode includes at least one of summarizing by text paragraph, summarizing by text segment, and summarizing by full text, wherein the text segment is determined based on the text content structure of the screen-reading text resource; The processing module 803 is used to process text resources within the target text range based on the target summary generation mode to obtain summary text information of the text resources; The playback module 804 is used for voice playback of summary text information.

[0099] This embodiment determines the target summary generation mode and the target text range corresponding to the target summary generation mode based on the received first input. Then, it processes the text information within the target text range based on the target summary generation mode to obtain summary text information, which is then played back by voice. This achieves the condensation of text information, compressing the original long full-text reading time into a short time, avoiding the redundant reception of invalid information, and enabling users to efficiently obtain key information within a limited time, thus improving information acquisition efficiency.

[0100] In some possible implementations of the embodiments of this application, the screen reading device 800 may include: The display module is used to synchronously display the summary text information when the summary text information is played by voice.

[0101] In some possible implementations of the embodiments of this application, the target summary generation mode includes summary generation by text paragraph or summary generation by text fragment; The display module is also used to synchronously display the original text location index of the text resource corresponding to the summary text information; the original text location index of the text resource includes at least one of the following: paragraph number and chapter title.

[0102] In some possible implementations of the embodiments of this application, during the process of playing the summary text information by voice, or after the summary text information has finished playing, the receiving module 801 is further configured to receive a second input, the second input being used to indicate query information related to the text content of the screen-reading text resource; The screen reading device 800 may also include: The generation module is used to respond to the second input by generating response information for the query based on the text content of the screen-read text resource; The output module is used to output response information.

[0103] In some possible implementations of the embodiments of this application, the receiving module 801 is further configured to receive a third input after the generating module generates response information for the query information based on the text content of the screen-reading text resource in response to the second input. The third input is used to indicate the query information related to the original position of the response information. The determination module is also used to respond to a third input, and based on the response information, to perform full-text search and semantic matching in the screen-reading text resource to determine the original text location information of the response information in the screen-reading text resource; The display module is also used to display the original text location information.

[0104] In some possible implementations of the embodiments of this application, the original text location information is at least one. The receiving module is also used to receive a fourth input, which is used to determine the target original text location information from the original text location information; The screen reading device 800 may also include: The jump module is used to respond to the fourth input and, when the original text of the text resource is read aloud on the display screen, jump to the target original text content corresponding to the target original text location information.

[0105] In some possible implementations of the embodiments of this application, the original text location information is at least two; The display module is specifically used for: When the original text of a text resource is read aloud on the display screen, different display styles are used to display the original text content corresponding to different original text location information.

[0106] In some possible implementations of the embodiments of this application, the target summary generation mode includes summarizing by full text and summarizing by text fragments. The summary text information of the text resource includes overall summary text information and partial summary text information. The overall summary text information is used to summarize the core theme of the screen-reading text resource, and the partial summary text information is used to summarize the key information of each text fragment.

[0107] The screen reading device in this application embodiment can be a device or a component in an electronic device, such as an integrated circuit or a chip. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0108] The electronic device in this application embodiment can be a terminal with an operating system. The operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.

[0109] The screen reading device provided in this application embodiment can achieve... Figures 1-7 The various processes in the screen reading method embodiment can achieve the same technical effect, and will not be described again here to avoid repetition.

[0110] like Figure 9 As shown, this application embodiment also provides an electronic device 900, including a processor 901 and a memory 902. The memory 902 stores programs or instructions that can run on the processor 901. When the program or instructions are executed by the processor 901, they implement the various steps of the above-described screen reading method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0111] It should be noted that the electronic devices in the embodiments of this application include the mobile terminals and non-mobile terminals mentioned above.

[0112] Figure 10 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.

[0113] The electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.

[0114] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 10 The structure of the electronic device 1000 shown does not constitute a limitation on the electronic device 1000. The electronic device 1000 may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be described in detail here.

[0115] The user input unit 1007 is used to receive a first input when the display unit 1006 displays text resources on the screen. The processor 1010 is configured to respond to a first input to determine a target summary generation mode and a target text range corresponding to the target summary generation mode; the target summary generation mode includes at least one of summarizing by text paragraph, summarizing by text fragment, and summarizing by full text, wherein the text fragment is determined based on the text content structure of the screen-read text resource; based on the target summary generation mode, the text resource within the target text range is processed to obtain summary text information of the text resource; and the summary text information is played aloud.

[0116] This embodiment determines the target summary generation mode and the target text range corresponding to the target summary generation mode based on the received first input. Then, it processes the text information within the target text range based on the target summary generation mode to obtain summary text information, which is then played back by voice. This achieves the condensation of text information, compressing the original long full-text reading time into a short time, avoiding the redundant reception of invalid information, and enabling users to efficiently obtain key information within a limited time, thus improving information acquisition efficiency.

[0117] In some possible implementations of the embodiments of this application, the display unit 1006 is used to synchronously display the summary text information when the summary text information is played by voice.

[0118] In some possible implementations of the embodiments of this application, the target summary generation mode includes summary generation by text paragraph or summary generation by text fragment; The display unit 1006 is also used to synchronously display the original text location index of the text resource corresponding to the summary text information; the original text location index of the text resource includes at least one of the following: paragraph number and chapter title.

[0119] In some possible implementations of the embodiments of this application, the user input unit 1007 is further configured to receive a second input during the process of the processor 1010 playing the summary text information by voice, or after the summary text information has finished playing. The second input is used to indicate query information related to the text content of the screen-reading text resource. The processor 1010 is configured to, in response to a second input, generate response information for the query based on the text content of the screen-read text resource, and output the response information.

[0120] In some possible implementations of the embodiments of this application, the user input unit 1007 is further configured to receive a third input after the processor 1010 generates response information for the query information based on the text content of the screen-reading text resource in response to the second input. The third input is used to indicate the query information related to the original position of the response information. The processor 1010 is also configured to, in response to a third input, perform full-text retrieval and semantic matching in the screen-reading text resource based on the response information, and determine the original text location information of the response information in the screen-reading text resource; The display unit 1006 is also used to display the original text location information.

[0121] In some possible implementations of the embodiments of this application, the original text location information is at least one; The user input unit 1007 is also used to receive a fourth input, which is used to determine the target original text location information from the original text location information; The processor 1010 is also configured to, in response to a fourth input, jump to the target original text content corresponding to the target original text location information when the original text of the text resource is read aloud on the display screen of the display unit 1006.

[0122] In some possible implementations of the embodiments of this application, the original text location information is at least two; Display unit 1006 is specifically used for: Display a list of locations containing at least two original text location information in a preset area of ​​the screen; the preset area may be the top of the screen or the side of the screen.

[0123] In some possible implementations of the embodiments of this application, there are at least two original text location information. The display unit 1006 is also used to display the original text content corresponding to different original text location information using different display styles when the original text of the text resource is read aloud on the display screen.

[0124] In some possible implementations of the embodiments of this application, the target summary generation mode includes summarizing by full text and summarizing by text fragments. The summary text information of the text resource includes overall summary text information and partial summary text information. The overall summary text information is used to summarize the core theme of the screen-reading text resource, and the partial summary text information is used to summarize the key information of each text fragment.

[0125] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, joysticks, etc., which will not be described in detail here.

[0126] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1009 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0127] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.

[0128] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described screen reading method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0129] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0130] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described screen reading method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0131] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0132] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the screen reading method embodiment described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0133] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0135] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A screen reading method, characterized in that, include: When displaying text resources for reading aloud on the screen, receive the first input; In response to the first input, a target summary generation mode and a target text range corresponding to the target summary generation mode are determined; the target summary generation mode includes at least one of summarizing by text paragraph, summarizing by text segment, and summarizing by full text, wherein the text segment is determined based on the text content structure of the screen-reading text resource; Based on the target summary generation mode, the text resources within the target text range are processed to obtain the summary text information of the text resources; The summary text information is played back via voice.

2. The method according to claim 1, characterized in that, The method further includes: When the summary text information is played by voice, the summary text information is displayed synchronously.

3. The method according to claim 1, characterized in that, The target summary generation mode includes summary generation by text paragraph or summary generation by text fragment, and the method further includes: The original text location index of the text resource corresponding to the summary text information is displayed synchronously; the original text location index of the text resource includes at least one of the following: paragraph number and chapter title.

4. The method according to any one of claims 1-3, characterized in that, During the playback of the summary text information, or after the playback of the summary text information has ended, the method further includes: Receive a second input, which indicates query information related to the text content of the screen-reading text resource; In response to the second input, a response to the query is generated based on the text content of the screen-read text resource. Output the response information.

5. The method according to claim 4, characterized in that, After generating response information for the query based on the text content of the screen-read text resource in response to the second input, the method further includes: Receive a third input, the third input being used to indicate query information related to the original position of the response information; In response to the third input, based on the response information, full-text search and semantic matching are performed in the screen-reading text resource to determine the original text location information of the response information in the screen-reading text resource; Display the original text location information.

6. The method according to claim 5, characterized in that, The original text location information is at least one, and the method further includes: Receive a fourth input, the fourth input being used to determine the target original text location information from the original text location information; In response to the fourth input, if the original text of the screen-reading text resource is displayed, the user is redirected to the target original text content corresponding to the target original text location information.

7. The method according to claim 5, characterized in that, The original text location information includes at least two elements; The display of the original text location information includes: A location list containing at least two of the original text location information is displayed in a preset area of ​​the screen; the preset area includes the top of the screen or the side of the screen.

8. The method according to claim 5, characterized in that, The original text location information includes at least two elements, and the method further includes: When displaying the original text of the screen-reading text resource, different display styles are used to display the original text content corresponding to different original text location information.

9. The method according to any one of claims 1-3, characterized in that, The target summary generation mode includes summarizing the full text and summarizing the text segments. The summary text information of the text resource includes overall summary text information and partial summary text information. The overall summary text information is used to summarize the core theme of the screen-reading text resource, and the partial summary text information is used to summarize the key information of each text segment.

10. A screen reading device, characterized in that, include: The receiving module is used to receive the first input when the text resource is read aloud on the display screen; A determining module is configured to, in response to the first input, determine a target summary generation mode and a target text range corresponding to the target summary generation mode; the target summary generation mode includes at least one of summarizing by text paragraphs, summarizing by text segments, and summarizing by the whole text, wherein the text segments are determined based on the text content structure of the screen-reading text resource; The processing module is used to process the text resources within the target text range based on the target summary generation mode to obtain the summary text information of the text resources; The playback module is used to play the summary text information via voice.