A text-to-speech method and apparatus

By converting text on the mobile phone interface into audio and overlaying it with a text-to-speech control panel, the difficulty of text reading in scenarios where reading is inconvenient is solved, enabling convenient and real-time text reading interaction and improving the user experience.

CN115240636BActive Publication Date: 2025-12-12HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110485195.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-20
Filing Date
2021-04-30
Publication Date
2025-12-12
Estimated Expiration
2041-04-30

AI Technical Summary

Technical Problem

In situations such as bumpy rides or dim lighting, it is difficult to read articles on a mobile phone, as existing technologies struggle to provide convenient text reading and real-time interaction.

Method used

The text on the user interface is converted into audio data and read aloud by electronic devices. A reading control panel is overlaid on the interface to display the sentence being read in real time. It supports flexible control at the sentence level, including touch operation and progress control.

Benefits of technology

It enables convenient text reading in scenarios where reading is inconvenient, improves user interactivity and user experience, reduces waiting time, and enhances the real-time nature and flexibility of the reading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240636B_ABST
    Figure CN115240636B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a text reading method and device, and relate to the technical field of electronics, which can read text on an application interface of an electronic device, and prompt a user in real time about a sentence being read through a panel superimposed on the application interface, and switch the sentence being read in real time according to an indication of the user. The scheme comprises: an electronic device displays a first user interface; a first operation of a user is received; in response to the first operation, a first content of the first user interface is acquired; a second user interface is displayed, and content displayed by the second user interface comprises text in the first content; wherein the second user interface shields part of a display area of the first user interface; a first sentence in the content of the second user interface is read, and identification information of the first sentence being read in the second user interface is displayed. Embodiments of the present application are used for text reading.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to the Chinese Patent Application No. 202110425978.2, filed on April 20, 2021, and entitled "Method for Controlling Double-queue Buffering Online Playing", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] Embodiments of the present application relate to the technical field of electronics, and in particular to a text reading method and device. BACKGROUND

[0003] Audiobooks are a common demand of users. Previously, people mainly listened to audiobook programs through radios. After the rise of mobile Internet, people began to obtain various information through mobile phones. For example, reading various articles through mobile phones is a usage habit of everyone. However, in the scene of walking, driving or taking a bus, etc., or in the scene of eye fatigue, dim light, etc., people will encounter great difficulties in reading articles through mobile phones. At this time, if the article can be read to the user through the mobile phone, the user's pain points will be solved. SUMMARY

[0004] Embodiments of the present application provide a text reading method and device, which can convert the text on the user interface displayed by the electronic device into playable audio data to realize text reading, and can also prompt the user in real time about the sentence being read through the reading control panel displayed on the user interface, and switch the reading sentence at any time according to the user's indication.

[0005] To achieve the above-mentioned purpose, the embodiments of the present application adopt the following technical solutions:

[0006] On the one hand, the embodiments of the present application provide a method for reading the displayed content, which can be applied to an electronic device. The method comprises: an electronic device displays a first user interface; receiving a first operation of a user; in response to the first operation, obtaining a first content of the first user interface; displaying a second user interface, the content displayed by the second user interface including text in the first content; wherein the second user interface shields part of the display area of the first user interface; reading a first sentence in the content of the second user interface, and displaying the identification information of the first sentence being read in the corresponding text of the second user interface.

[0007] In this scheme, the electronic device can read the text on the first user interface, and display the identification information through the second user interface superimposed on part of the display area of the first user interface, so as to dynamically and intuitively prompt the user about the content being read, which is good in user interaction and good in user experience. For example, the first user interface can be the interface of a target application, and the second user interface can be the interface corresponding to the reading control panel.

[0008] In a possible design, before displaying the second user interface, the method further includes: identifying one or more sentences from the acquired first content of the first user interface.

[0009] That is, the electronic device identifies and divides the text in the acquired first content of the first user interface in the granularity of a sentence.

[0010] In another possible design, the method further includes: detecting, by the electronic device, a second operation of the user; and in response to the second operation, controlling, by the electronic device, the electronic device to read a second sentence corresponding to the second operation based on a text position or text content corresponding to the second operation or read from a position of the second sentence.

[0011] In this solution, the user can instruct the electronic device to read a sentence at an arbitrary position as expected, thereby achieving flexible, fine, and accurate reading control in the granularity of a sentence.

[0012] In another possible design, the method further includes: displaying, by the electronic device, identification information of a second sentence being read in the second user interface corresponding to the text.

[0013] That is, as the reading sentence is switched, the electronic device can display identification information of the sentence being read in real time, explicitly, and dynamically, which is more interactive with the user.

[0014] In another possible design, the method further includes: the second operation is a touch operation of the user acting on a text position corresponding to the second sentence displayed in the second user interface; and based on the detected touch operation, the method further includes: confirming, by the electronic device, to acquire a voice corresponding to the second sentence and read the voice.

[0015] In this way, the user can instruct the electronic device to read a sentence on the second user interface by a touch operation on a position of the sentence expected to be read.

[0016] In another possible design, the method further includes: displaying, by the electronic device, the second user interface based on the identified one or more sentences, wherein each of the one or more sentences displayed in the second user interface corresponds to a control; and based on the detected touch operation, the method further includes: confirming, by the electronic device, to acquire a voice corresponding to the second sentence and read the voice, specifically including: triggering, by the electronic device, the control corresponding to the second sentence based on the touch operation corresponding to the second operation; and in response to the control event, acquiring, by the electronic device, the voice of the second sentence corresponding to the control and reading the voice.

[0017] In this solution, the electronic device can determine a sentence specified by the user through a control event and read the sentence, thereby achieving flexible, fine, and accurate reading control in the granularity of a sentence.

[0018] In another possible design, the method further includes: the electronic device indexing the recognized one or more sentences; and the method further includes: the second user interface displaying content further including a play control, which can be used to control the progress of the reading or the speed of the reading. The play control includes a progress bar, which is matched with the index of the one or more sentences, and is controlled based on the granularity of the recognized sentences. The play control includes a down control, which is used to control the current reading progress to switch to the next sentence; and / or, the play control includes an up control, which is used to control the current reading progress to switch to the previous sentence.

[0019] In this way, when the user drags the progress bar, or instructs the electronic device to read the next sentence or the previous sentence, the electronic device can determine the sentence to be read by matching the index of the sentence.

[0020] In another possible design, the method further includes: the electronic device detecting a third operation of the user; and in response to the third operation, adjusting the size of the window corresponding to the second user interface.

[0021] That is, the user can instruct the electronic device to adjust the size of the window corresponding to the second user interface, for example, to minimize the second user interface, or to switch the second user interface between half-screen display and full-screen display, etc.

[0022] In another possible design, the method further includes: the electronic device detecting a fourth operation of the user; and in response to the fourth operation, minimizing the window corresponding to the second user interface, and continuing to read the current text content without interruption; after the window corresponding to the second user interface is minimized, the method further includes: restoring the second user interface by the card corresponding to the reading function, and displaying the current reading text content and progress.

[0023] That is, after the window corresponding to the second user interface is minimized, the electronic device can continue to read the text without interruption. And the second user interface disappears, and the second user interface can be restored by the corresponding card.

[0024] In another possible design, the play control includes a refresh control, and the method further includes: the electronic device, in response to the refresh control, stopping reading the current text, and re-acquiring the second content of the first user interface for reading.

[0025] The second content can be the same as or different from the first content. The electronic device can re-acquire the content on the current first user interface in response to the operation of the user instructing the refresh, and read the text in the re-acquired content.

[0026] In another possible design, the method further includes: the electronic device sending the identified one or more sentences to the server respectively for text-to-speech processing; receiving and caching the speech obtained by the server after processing; reading the speech obtained by the server in real time, or caching the speech obtained by the server in sequence according to the identified one or more sentences.

[0027] That is, the electronic device and the server perform text-to-speech conversion and caching in sequence according to sentences.

[0028] In another possible design, the method further includes: the electronic device caching the speech obtained by the server through two queues; the first queue is used to cache one or more speech packets of a sentence currently being processed by the server; and the second queue is used to cache speech corresponding to one or more sentences that have been processed by the server; and reading the first sentence in the content of the second user interface includes: if the speech currently received from the server is the speech of the first sentence, obtaining the speech corresponding to the first sentence from the first queue to read; or if the speech currently received from the server is not the speech of the first sentence, obtaining the speech corresponding to the first sentence from the second queue to read.

[0029] In this way, each speech packet is stored in the first queue after being synthesized, and the first queue enables the user to quickly hear the reading sound after waiting for one speech packet to be synthesized, so that the user waits for a short time before the reading starts, and the time for starting the reading can be reduced to the order of ten milliseconds, and the real-time performance of the reading is better. In addition, because the audio synthesis speed is fast, the sentences cached in the second queue generally precede the sentences being played, and when the user drags the progress bar, if the progress after the dragging falls into the second queue, the effect of immediate playing can be achieved without waiting, and the real-time responsiveness is strong.

[0030] In another possible design, before the electronic device displays the second user interface, the method further includes: the electronic device processing the obtained first content of the first user interface to remove non-text information, and obtaining the text in the first content; and displaying the processed text on the second user interface.

[0031] That is, the electronic device removes the non-text information that is not suitable for reading, and then displays the non-text information on the second user interface, so that the content of the reading is consistent with the content displayed on the second user interface.

[0032] In another possible design, before the electronic device displays the second user interface, the method further includes: the electronic device dividing the processed text into one or more sentences according to special punctuation marks, to obtain the identified one or more sentences.

[0033] That is, the electronic device divides the text into sentences before displaying the text on the second user interface.

[0034] In another possible design, the method further includes: in response to the voice instruction of the user, the electronic device controls the progress of the reading or the speed of the reading.

[0035] In this way, the user can control the progress or the speed of the reading of the text by using the voice instruction.

[0036] In another possible design, the method further includes: the electronic device detects a fifth operation of the user; and in response to the fifth operation of the user, the electronic device scrolls the text content displayed on the second user interface and continues the reading of the current text content without interruption.

[0037] In this solution, the text content displayed on the second user interface can be scrolled to facilitate the user to view or switch the sentence to be read.

[0038] In another aspect, an apparatus for reading displayed content is provided. The apparatus is included in an electronic device. The apparatus has the function of implementing the electronic device behavior in any of the methods in the above aspects and possible designs, so that the electronic device performs the method for reading displayed content performed by the electronic device in any of the above aspects and possible designs. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above function. For example, the apparatus can include a display unit, a processing unit, and a reading unit, etc.

[0039] In another aspect, an apparatus for reading displayed content is provided. The apparatus is included in an electronic device. The apparatus has the function of implementing the electronic device behavior in any of the methods in the above aspects and possible designs, so that the electronic device performs the method for reading displayed content performed by the electronic device in any of the above aspects and possible designs. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above function. For example, the apparatus can include a display unit, a processing unit, and a reading unit, etc.

[0040] In another aspect, an apparatus for reading displayed content is provided. The apparatus is included in an electronic device. The apparatus has the function of implementing the electronic device behavior in any of the methods in the above aspects and possible designs, so that the electronic device performs the method for reading displayed content performed by the electronic device in any of the above aspects and possible designs. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above function. For example, the apparatus can include a display unit, a processing unit, and a reading unit, etc.

[0041] In another aspect, an apparatus for reading displayed content is provided. The apparatus is included in an electronic device. The apparatus has the function of implementing the electronic device behavior in any of the methods in the above aspects and possible designs, so that the electronic device performs the method for reading displayed content performed by the electronic device in any of the above aspects and possible designs. The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above function. For example, the apparatus can include a display unit, a processing unit, and a reading unit, etc.

[0042] In yet another aspect, an embodiment of the present application provides a computer program product, which, when running on a computer, causes the computer to perform the method of reading aloud the content displayed by the electronic device in any possible design of the aspects above.

[0043] In yet another aspect, an embodiment of the present application provides a chip system applied to an electronic device. The chip system comprises one or more interface circuits and one or more processors; the interface circuit and the processor are interconnected through a circuit; the interface circuit is configured to receive a signal from a memory of the electronic device and send a signal to the processor, the signal comprising computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device performs the method of reading aloud the content displayed in any possible design of the aspects above.

[0044] The beneficial effects of the other aspects above can be referred to the description of the beneficial effects of the method aspects, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 A hardware structure schematic diagram of an electronic device provided by an embodiment of the present application;

[0046] Figure 2 A reading method flowchart provided by an embodiment of the present application;

[0047] Figure 3 An interface schematic diagram provided by an embodiment of the present application;

[0048] Figure 4 Another interface schematic diagram provided by an embodiment of the present application;

[0049] Figure 5 A group of interface schematic diagrams provided by an embodiment of the present application;

[0050] Figure 6 Another interface schematic diagram provided by an embodiment of the present application;

[0051] Figure 7 Another interface schematic diagram provided by an embodiment of the present application;

[0052] Figure 8A Another group of interface schematic diagrams provided by an embodiment of the present application;

[0053] Figure 8B Another group of interface schematic diagrams provided by an embodiment of the present application;

[0054] Figure 8C Another group of interface schematic diagrams provided by an embodiment of the present application;

[0055] Figure 8DAnother set of interface schematic diagrams provided by the embodiments of the present application;

[0056] Figure 8E A drag schematic diagram of a drag point provided by the embodiments of the present application;

[0057] Figure 9A Another set of interface schematic diagrams provided by the embodiments of the present application;

[0058] Figure 9B Another set of interface schematic diagrams provided by the embodiments of the present application;

[0059] Figure 9C Another set of interface schematic diagrams provided by the embodiments of the present application;

[0060] Figure 10 Another set of interface schematic diagrams provided by the embodiments of the present application;

[0061] Figure 11 Another set of interface schematic diagrams provided by the embodiments of the present application;

[0062] Figure 12 Another set of interface schematic diagrams provided by the embodiments of the present application;

[0063] Figure 13A Another set of interface schematic diagrams provided by the embodiments of the present application;

[0064] Figure 13B Another set of interface schematic diagrams provided by the embodiments of the present application;

[0065] Figure 14 A module interaction schematic diagram provided by the embodiments of the present application;

[0066] Figure 15 A cache schematic diagram provided by the embodiments of the present application;

[0067] Figure 16 Another cache schematic diagram provided by the embodiments of the present application;

[0068] Figure 17 Another cache schematic diagram provided by the embodiments of the present application;

[0069] Figure 18 A logic block diagram of text reading provided by the embodiments of the present application;

[0070] Figure 19 Another logic block diagram of text reading provided by the embodiments of the present application;

[0071] Figure 20 A module interaction timing diagram provided by the embodiments of the present application;

[0072] Figure 21A reading method flowchart provided in an embodiment of the present application;

[0073] Figure 22 An interface diagram of another text reading scheme provided in an embodiment of the present application;

[0074] Figure 23 A structural schematic diagram of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0075] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; the "and / or" in the present application only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0076] Hereinafter, the terms "first" and "second" are used only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more features. In the description of the embodiments, unless otherwise specified, the meaning of "multiple" is two or more than two.

[0077] In the embodiments of the present application, the words such as "exemplarily" or "for example" are used to represent as an example, illustration or description. Any embodiment or design scheme described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words such as "exemplarily" or "for example" are intended to present the relevant concept in a specific manner.

[0078] In view of the pain point that people are not convenient to read articles on electronic devices in some scenarios, the embodiments of the present application provide a text reading method, which can convert the text on the user interface of the electronic device into audio data for reading, and can also prompt the user in real time the position of the reading content, and switch the reading sentence at any time according to the user's indication.

[0079] For example, the electronic device can be a mobile terminal such as a mobile phone, a tablet, a wearable device (e.g., a smart watch), a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The embodiments of the present application do not limit the specific type of the electronic device.

[0080] Exemplary, Figure 1 A structural schematic diagram of the electronic device 100 is shown. The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 160, a speaker 160A, a receiver 160B, a microphone 160C, a headset interface 160D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0081] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated in one or more processors.

[0082] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of instruction fetching and instruction execution.

[0083] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can store instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instructions or data again, it can be directly called from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system.

[0084] In some embodiments, the antenna 1 of the electronic device 100 is coupled with the mobile communication module 150, and the antenna 2 is coupled with the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices such as a text to speech (TTS) cloud server through wireless communication technology. Thus, the electronic device 100 can send text to the TTS cloud server and obtain audio data converted from the text from the TTS cloud server.

[0085] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected with the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0086] The display screen 194 is configured to display images, videos, and the like. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diodes (QLED), or the like. In some embodiments, the electronic device 100 can include one or N display screens 194, where N is a positive integer greater than 1. For example, the display screen 194 can be configured to display an application interface, a reading control panel, a reading control card, or the like.

[0087] The internal memory 121 can be configured to store computer-executable program codes including instructions. The processor 110 performs various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program (e.g., a sound playing function, an image playing function, or the like) required by a function, and the like. The data storage area can store data (e.g., audio data, a phonebook, or the like) created during the use of the electronic device 100, and the like. In addition, the internal memory 121 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one of a magnetic disk storage device, a flash memory device, a universal flash storage (UFS), or the like.

[0088] The electronic device 100 can implement an audio function through the audio module 160, the speaker 160A, the receiver 160B, the microphone 160C, the earphone interface 160D, and the application processor, or the like. For example, music playing, recording, or the like.

[0089] The audio module 160 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 160 can also be configured to encode and decode an audio signal. In some embodiments, the audio module 160 can be disposed in the processor 110, or some functional modules of the audio module 160 can be disposed in the processor 110. For example, the audio module 160 can obtain a voice instruction of a user, and read a text to the user.

[0090] The speaker 160A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or a hands-free call through the speaker 160A.

[0091] The receiver 160B, also referred to as an "earpiece", is configured to convert an audio electrical signal into a sound signal. When the electronic device 100 is engaged in a call or a voice message, a user can listen to a voice through the receiver 160B by bringing the receiver 160B close to an ear.

[0092] The microphone 160C, also referred to as a "microphone", "voice microphone", is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, a user can make a sound through the mouth close to the microphone 160C, and input the sound signal into the microphone 160C. The electronic device 100 can be provided with at least one microphone 160C. In some other embodiments, the electronic device 100 can be provided with two microphones 160C, in addition to collecting sound signals, noise reduction functions can also be realized. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 160C, in addition to collecting sound signals, noise reduction, sound source identification, directional recording and other functions can also be realized.

[0093] The touch sensor 180K, also referred to as a "touch panel". The touch sensor 180K can be provided on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also referred to as a "touch screen". The touch sensor 180K is configured to detect a touch operation acting on or near the touch sensor 180K. The touch sensor 180K can transmit the detected touch operation to the application processor to determine the type of touch event. The display screen 194 can provide visual output related to the touch operation. In some other embodiments, the touch sensor 180K can also be provided on the surface of the electronic device 100, which is different from the position of the display screen 194. For example, the touch screen can detect the user's touch operation, so as to trigger the phone to start the text reading function or perform the reading control.

[0094] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software or a combination of software and hardware.

[0095] In the embodiments of the present application, the display screen 194 can display the interface of the target application, and the processor 110 controls the audio module 160 to play the audio data corresponding to the text on the currently displayed target application interface by running the instructions stored in the internal memory 121, thereby realizing the text reading function. The display screen 194 can also display a reading control panel and the like to facilitate user interaction and playback control. Moreover, the processor 110 can respond to the touch operation of the user on any sentence on the reading control panel to control the electronic device 100 to read from the sentence. Moreover, with the change of the playback progress such as fast forward or fast backward, the reading control panel can highlight the currently read sentence in real time, and the user experience is good.

[0096] In the following, the electronic device is taken as a mobile phone with the structure as shown in Figure 1 The text reading method provided by the embodiments of the present application will be described below. As shown in Figure 2 The method can include:

[0097] 201. The mobile phone starts the text reading function.

[0098] When the user wants the mobile phone to read the text on the currently displayed user interface, the user can instruct the mobile phone to start the text reading function, so as to read the text on the user interface for the user. The text reading function can also have a specific name, such as Xiaoyi reading and the like. Among them, the user interface that the user wants the mobile phone to read can be referred to as the first user interface. In the following embodiments, the user interface can be referred to as interface.

[0099] It can be understood that any user interface displayed by the mobile phone can correspond to a certain application on the mobile phone, such as the desktop displayed by the mobile phone, which can also be understood as the interface of the desktop management application. That is, the first user interface is the interface of a certain target application. When the user wants the mobile phone to read the text on the currently displayed interface of the target application, the user can instruct the mobile phone to start the text reading function.

[0100] In some embodiments of the present application, the target application can be any application on the mobile phone, which can be a native system application or a third-party application, and the specific type of the target application is not limited by the embodiments of the present application. For example, the target application can be a novel application, a browser application or a news application and the like. That is, the mobile phone can read the text displayed on the interface of any application, and can read the content seen by the user based on any application interface. That is, the mobile phone can read the text on any user interface.

[0101] In other embodiments, the target application is a preset first specific application, such as an application related to an article, or an application with usually more text content on the interface. For example, the target application can be a news application, a public number or a Zhihu The phone can read aloud text from the user interface of a specific pre-defined application.

[0102] In other embodiments, the target application is an application other than a preset second specific application. For example, the phone has an application blacklist that includes the second specific application (e.g., a financial application unsuitable for text-to-speech), and the interface of the second specific application is not allowed to be read aloud. The phone can then read aloud the text on the user interface of applications other than the second specific application.

[0103] In other embodiments, the mobile phone can read aloud text on the user interface that conforms to preset rules. This application does not limit the specific content of these preset rules. For example, the preset rules may include a number of characters on the user interface that is greater than or equal to a preset value.

[0104] In the following embodiments, the user interface of the target application may also be referred to simply as the interface of the target application or the application interface.

[0105] There are several ways a mobile phone can activate the text-to-speech function. For example, the phone can activate the text-to-speech function in response to a user's voice command. See [example...] for an example. Figure 3 After the phone detects the user's voice command "Xiaoyi Xiaoyi," it wakes up the voice assistant; then, after detecting the user's voice command "Read it aloud," the voice assistant wakes up the reading module and starts the text reading function. This application embodiment does not limit the specific content of the user's voice command to start the text reading function.

[0106] For example, a mobile phone can display a text-to-speech control on the application interface. In response to a user's action on this control, the phone can activate the text-to-speech function. This control prompts the user that the phone can read the text on the interface aloud, and allows the user to quickly activate the text-to-speech function. In some embodiments, the phone displays a floating text-to-speech ball on the application interface; this floating ball is the text-to-speech control. The phone activates the text-to-speech function after detecting a user's tap on the floating ball. In other embodiments, the phone displays, on the interface of a first specific application capable of text-to-speech, a text-to-speech control. Figure 4 The capsule-shaped reading control 401 shown.

[0107] For example, the phone only displays the text-to-speech control on the application interface after detecting a user's preset voice command. In response to the user's action on this control, the phone initiates the text-to-speech function.

[0108] For another example, the mobile phone can display a reading control on the interface after the smart screen reading function is enabled (e.g., the smart screen reading function is enabled after detecting a long press operation of the user on the interface of an application). The mobile phone starts the text reading function in response to an operation of the user on the reading control.

[0109] For another example, the mobile phone starts the text reading function after detecting a preset shortcut gesture of the user (e.g., a three-finger downward gesture or a gesture in the air, etc.).

[0110] It should be noted that the mobile phone can also start the text reading function in response to other touch operations, voice instructions or gestures of the user, and the specific manner of triggering the mobile phone to start the text reading function is not limited in the embodiments of the present application. The operation of the user indicating the mobile phone to start the text reading function can be referred to as a first operation.

[0111] In some embodiments, if a certain application does not support the text reading function, the mobile phone can prompt the user that the interface cannot be read or the function cannot be read after detecting the operation of the user indicating the start of the function, see the examples in (a)-(b) in Figure 5

[0112] 202. The mobile phone reads the text on the currently displayed interface of the application.

[0113] After the mobile phone starts the text reading function, the mobile phone acquires the text on the currently displayed interface of the application and reads the text for the user. For example, after the mobile phone starts the text reading function, the mobile phone first acquires the interface content of the target application (i.e., acquires the first content of the first user interface). The interface content includes information of each view node on the interface, such as text and picture information, etc. Then, the mobile phone acquires the text on the interface of the application from the acquired interface content.

[0114] In some embodiments, the mobile phone acquires the text on the interface of the application, including the text part of the interface of the application that has been loaded (the text part that has been loaded can be one screen, such as the content displayed on the current interface, or more than one screen, such as the content of the subsequent pages in addition to the content displayed on the current interface, which is not limited), and excluding the text part of the interface of the application that has not been loaded, which will not be continuously acquired without a refresh operation. In other embodiments, the mobile phone acquires the text on the interface of the application, including the text part of the interface of the application that has been loaded, and also including the text part of the interface of the application that has not been loaded and will be continuously acquired during the reading process.

[0115] After the mobile phone starts the text reading function, the mobile phone acquires the text on the interface of the application, and reads the text for the user, see Figure 6 ​In response to the user's operation in (a), the phone displays the reading control panel 601 on the application interface. At this time, the reading control panel displayed on the phone can also be referred to as a second user interface. The second user interface is different from the first user interface corresponding to the application interface to be read, and the second user interface at least blocks part of the display area of the first user interface. In some embodiments, the window size of the reading control panel is smaller than the window size of the application interface, so that the phone can conveniently allow the user to understand the application interface and the target application corresponding to the text being read while displaying the reading control panel. For example, after the phone starts the text reading function, the reading control panel can be displayed in half of the screen. The reading control panel can be displayed in the lower half of the screen. The bottom of the reading control panel can be aligned with the bottom of the application interface. The reading control panel can be used for interaction with the user to facilitate the user to view the reading content, or to perform reading control and the like.

[0116] If the current application interface has no readable text after the phone starts the text reading function, refer to Figure 7 In the example, the phone can prompt the user that the current interface has no readable text.

[0117] If the current application interface has readable text after the phone starts the text reading function, the reading control panel 601 can display the text on the application interface. The text can include one or more sentences. The text can be identified by sentence, and one sentence is divided into one paragraph.

[0118] In some embodiments, after the phone starts the text reading function, the phone can prompt the user about the total reading time of the text on the current application interface by various ways such as voice broadcast or text display according to the number of readable texts on the application interface to be read. For example, the phone can display the total reading time prompt information 602 on the reading control panel 601: "The current text is expected to be read for less than one minute", or "The current text is expected to be read for xx minutes", or "The current text is expected to be read for more than 30 minutes", and the like. In an implementation scheme, the time prompt information 602 can be displayed on the first line of the reading control panel, so as to facilitate the user to quickly and intuitively know the total time required for reading. During the subsequent reading process, the phone can also prompt the user about the reading time and / or the remaining reading time and the like on the reading control panel.

[0119] The phone can also prompt the user about the current reading state by various ways such as voice broadcast or text display. For example, the reading state prompt information 603 can be displayed on the reading control panel 601. The reading state prompt information 603 can include the reading state such as reading (or reading in progress), reading pause, or reading end. After the phone starts to read the text content on the application interface, the reading state is reading.

[0120] In addition, the reading control panel can further include an icon 604 of the text reading function and / or information such as the name of the text reading function (e.g., Xiaoyi Reading).

[0121] In some embodiments, the sentence currently being read by the phone can be displayed at a preset fixed position on the reading control panel 601, for example, at a position close to the top of the reading control panel 601. For example, the sentence currently being read can be displayed from the position of the second row of the reading control panel, so as to facilitate the user to quickly and intuitively locate the text content currently being read, and improve the user experience. The reading control panel 601 can further display the sentence after the sentence being read. The following embodiments are described by taking the second row as an example of the preset fixed position.

[0122] Referring to (a)-(b) in FIG. 6, Figure 8A As the reading proceeds, when the phone reads the next sentence, the text content displayed on the reading control panel scrolls up, and the sentence currently being read is automatically moved to the position of the second row. In some embodiments, the first row of the reading control panel can display the previous sentence of the currently read sentence, and the total reading time is no longer displayed; in other embodiments, the first row of the reading control panel no longer displays the total reading time but displays the remaining reading time and / or the read time.

[0123] In the reading state, if the audio focus held by the voice reading service is occupied by other applications (such as incoming calls, music, video playback, or audio / video calls), the phone automatically pauses the text reading, the reading control panel stops displaying, and a reading control card is generated in the menu bar (or notification bar). In other words, the reading control panel is switched to the reading control card in the menu bar. After the audio focus is released by other applications and the voice reading service re-holds the audio focus, the phone resumes the text reading.

[0124] In the reading state, if the phone exits the current interface display and displays the desktop, displays other application interfaces, or displays the lock screen, in one embodiment, the phone continues to read the text, and the reading control panel is switched to the reading control card in the menu bar. In another embodiment, the phone pauses the reading of the text, and the reading control panel is switched to the reading control card in the menu bar. In another embodiment, the phone stops reading the text, the reading control panel disappears, and the text reading function is exited.

[0125] In some embodiments, the phone displays the identification information of the text corresponding to the sentence currently being read on the reading control panel, so as to intuitively, dynamically, and in real time highlight the text content and reading position currently being read to the user, and improve the interactivity and user experience.

[0126] For example, the identification information can be the text information of the sentence being read, and the text information of the sentence being read is displayed on the reading control panel, and the text information of other sentences is not displayed.

[0127] For another example, the identification information can be the text information of the sentence being read, and the text information of the sentence being read and the text information of several adjacent sentences are displayed on the reading control panel, and the text information of the sentence being read is displayed in a manner different from the display manner of other sentences. For example, as shown in (a) of FIG. 8A, the identification information is the text information of the sentence being read highlighted on the reading control panel, or as shown in (b) of FIG. 8B, the identification information is the text information of the sentence being read underlined, or the identification information can be the text information of the sentence being read displayed in bold, and the like. Figure 6 Figure 6

[0128] In other embodiments, the identification information of the text corresponding to the sentence being read displayed on the reading control panel of the mobile phone is the text information of the line being read in the sentence being read. The display manner of the line text information can be different from the display manner of other lines.

[0129] In the embodiments of the present application, the reading control panel can further include some playback controls for controlling the progress of reading or the speed of reading, and the like. For example, the playback controls can be used to control the reading of the previous sentence, the reading of the next sentence, the pause / resume, the fast forward, the fast backward, the speed, the refresh, the minimization, or the exit, and the like.

[0130] For example, as shown in (b) of FIG. 8B, the reading control panel includes a next sentence control 801 (which can also be referred to as a down control or other names), and the mobile phone detects the operation of the user clicking the next sentence control 801, reads the next sentence text, and as shown in (c) of FIG. 8C, automatically displays the next sentence text in the second line position (i.e., the above-mentioned preset fixed position) of the reading control panel and highlights the next sentence text. Figure 8A Figure 8A For example, as shown in (b) of FIG. 8B, the reading control panel includes a next sentence control 801 (which can also be referred to as a down control or other names), and the mobile phone detects the operation of the user clicking the next sentence control 801, reads the next sentence text, and as shown in (c) of FIG. 8C, automatically displays the next sentence text in the second line position (i.e., the above-mentioned preset fixed position) of the reading control panel and highlights the next sentence text. Figure 8A

[0131] ​​​​After the last sentence of the text on the application interface is read aloud, the phone can announce "Reading complete" to notify the user, and the reading status changes from "Reading in progress" to "Reading finished." In some embodiments, if the phone does not detect any user interaction with the text reading within a preset time after the reading ends, it automatically exits the text reading function. In other embodiments, the phone does not automatically exit the text reading function after the reading ends; it only exits the text reading function after detecting a user instruction to exit.

[0132] For example, such as Figure 8B As shown in (a), the text-to-speech control panel includes a previous sentence control 802 (also known as a move-up control or other names). After the mobile phone detects the user's click on the previous sentence control 802, it reads the previous sentence text aloud, and as shown in (a). Figure 8B As shown in (b), the previous sentence text is highlighted and automatically moved to the second line of the text-to-speech control panel. However, if the phone is currently reading the first sentence of the application interface, the previous sentence control is disabled.

[0133] For example, such as Figure 8C As shown in (a), the text-to-speech control panel includes a pause / resume control 803. When the text-to-speech status is active, the phone detects the user's tap on control 803 and stops reading the text. At this time, as shown in (a), the text-to-speech control panel includes a pause / resume control 803. Figure 8C As shown in (b), the reading state changes from reading aloud to reading paused. Then, after the phone detects the user's click on control 803 again, it resumes reading the text. At this point, as shown... Figure 8C The reading status shown in (a) has changed from reading paused to reading in progress.

[0134] In some embodiments, the text-to-speech control panel may also include a fast-forward control, and the next sentence control may also be used for fast-forward control. After the mobile phone detects that the user has long-pressed or repeatedly clicked the next sentence control, or after the mobile phone detects that the user has clicked the fast-forward control, it determines the starting position of the sentence to be read aloud based on the proportional relationship between the text progress after the user's operation and the entire text of the currently displayed application interface, and begins reading aloud from that starting position.

[0135] In some embodiments, the text-to-speech control panel may also include a rewind control, and the previous sentence control may also be used for rewind control. After the mobile phone detects that the user has long-pressed or tapped the previous sentence control multiple times, or after the mobile phone detects that the user has tapped the rewind control, it determines the starting position of the sentence to be read aloud based on the ratio of the text progress after the user's operation to the entire text, and begins reading aloud from that starting position.

[0136] In addition, the reading control panel can further include a progress bar. After the mobile phone detects that the user drags the progress bar, the mobile phone determines the starting position of the to-be-read sentence corresponding to the dragging point on the progress bar according to the proportional relationship between the text corresponding to the dragging operation and the entire text, and starts reading from the starting position.

[0137] It can be understood that when the user indicates fast forward or fast backward, the position of the dragging point on the progress bar also moves forward / backward accordingly. The mobile phone can determine the starting position of the to-be-read sentence corresponding to the dragging point on the progress bar, and starts reading from the starting position.

[0138] In the embodiments of the present application, the mobile phone can index the identified one or more sentences. The progress bar matches the index of the sentences, and the mobile phone can control based on the granularity of the identified sentences. For example, each sentence can correspond to a node, the progress bar matches the node of the sentence, and after the user drags the progress bar, the dragging point can correspond to the node of a certain sentence according to the proportional relationship, and the mobile phone reads the sentence corresponding to the node. For example, after the user drags the progress bar 800 shown in (a) of Figure 8D , if the dragging point 80 shown in (b) of Figure 8D corresponds to the third sentence, the mobile phone starts reading the third sentence.

[0139] In one implementation, as shown in (a) of Figure 8E , when the user drags the progress bar, the dragging point can only be dragged to the discrete position corresponding to the node of the sentence. After stopping dragging, the mobile phone starts reading from the position of the sentence corresponding to the node where the dragging point is located.

[0140] In another implementation, the node position of the sentence corresponds to the starting position of the sentence. As shown in (b) of Figure 8E , when the user drags the progress bar, the dragging point can be dragged to any position. When the dragging point is dragged to a position between two node positions and the user releases the dragging point, the dragging point automatically corresponds to the position of the latter node, and the mobile phone starts reading from the position of the sentence corresponding to the latter node.

[0141] In another implementation, the node position of the sentence corresponds to the starting position of the sentence. As shown in (c) of Figure 8E , when the user drags the progress bar, the dragging point can be dragged to any position. When the dragging point is dragged to a position between two node positions and the user releases the dragging point, the dragging point automatically corresponds to the position of the former node, and the mobile phone starts reading from the position of the sentence corresponding to the former node.

[0142] In addition, as shown in (d) of Figure 8AAs shown in (b), the text-to-speech control panel includes a speed control 804. After the mobile phone detects that the user clicks the speed control 804 to select a target speed, it reads the text aloud at the speed corresponding to the target speed. For example, the target speed can include 0.5x, 1x, or 2x speed, etc.

[0143] The text-to-speech control panel can also include a refresh control. After the phone detects the user clicking the refresh control, it re-acquires the application's interface content (i.e., re-acquires the second content of the first user interface) and retrieves the text from that content for text-to-speech. For example, due to network issues, the text on the application interface might not be fully loaded, preventing the phone from reading all the text. In this case, the user can use the refresh control to re-acquire the text so that the phone can load and read all the text on the interface. For example, such as... Figure 9A As shown in (a), if the user clicks the refresh control 805 after the phone has been reading aloud for a period of time, the text on the application interface will be retrieved again, and then... Figure 9A As shown in (b) above, the reading resumes, and the progress bar starts moving from the beginning. Understandably, the text or images displayed on the refreshed application interface may have changed compared to before the refresh.

[0144] like Figure 9B As shown in (a), the text-to-speech control panel also includes an exit control 806. After the phone detects the user clicking the exit control 806, it exits the text-to-speech function, stops reading the text, and as shown in (a). Figure 9B The reading control panel shown in (b) disappears.

[0145] If the phone detects that the user has clicked the pause / resume control after the text reading has ended, it will read the text content on the screen from the beginning. If the phone detects that the user has clicked refresh, it will re-acquire the text and read the text content on the screen from the beginning. Alternatively, if the phone determines that the user has not interacted with the text reading function within a preset time (e.g., 5 seconds or 10 seconds) after the text reading has ended, it will exit the text reading function, and the text reading control panel will disappear.

[0146] In addition, see Figure 9CIn (a)-(d), after the phone enables the text-to-speech function, if there is a network anomaly (e.g., no network, weak network, or network connection timeout during text capture and audio synthesis), the text on the application interface cannot be obtained, or the text cannot be read aloud after being obtained, resulting in abnormal results. In this case, the phone can notify the user of the anomaly or remind the user to try again later. In these situations, the refresh control is effective, while other controls are disabled (e.g., grayed out) and have no response after being clicked by the user. In some embodiments, in these situations, if reading aloud does not start within a preset time and no user interaction is detected, the phone exits the text-to-speech function.

[0147] In other embodiments of this application, the mobile phone can also respond to user voice commands to control the reading progress, speed, or exit, such as reading the previous sentence, reading the next sentence, pausing / resume, fast forward, rewind, speed adjustment, refreshing, minimizing, or exiting. For example, after detecting the user's voice command "read the next sentence," the mobile phone stops reading the current sentence and reads the next sentence. As another example, after detecting the user's voice command "Hey Celia, pause reading," the mobile phone stops reading the text content. Yet another example, after detecting the user's voice command "start reading from the sentence 'From Harbin to Beijing,'" the mobile phone switches to reading from that sentence onwards.

[0148] In some embodiments, the form, window size, or position of the reading control panel can also be adjusted. For example, when the reading control panel is displayed in half-screen mode, the phone can respond to the user's dragging of the reading control panel's border, moving the display position of the reading control panel on the screen.

[0149] For example, see Figure 10 In (a) of the text-to-speech control panel, a minimize control 807 is included. After the phone detects a user tapping the minimize control 807, it minimizes the text-to-speech control panel while maintaining uninterrupted text reading. Alternatively, the phone minimizes the text-to-speech control panel while maintaining uninterrupted text reading after the phone detects a user tapping the back button, home button, or menu button on the screen. Alternatively, the phone can minimize the text-to-speech control panel in response to a user's pull-down operation on the control panel, or in response to a user's pull-down operation on control 1003 or its surrounding area.

[0150] Alternatively, when the text-to-speech control panel is displayed in half-screen mode, see [link to relevant documentation]. Figure 10 In (a) of the above, when the phone detects that the user is outside the reading control panel area, the reading control panel is minimized. In some embodiments, an overlay is set on the application interface, and when the phone detects that the user clicks on the overlay area outside the reading control panel area, the reading control panel is minimized.

[0151] In some embodiments of this application, the reading panel is minimized by shrinking the reading control panel to a smaller size, i.e., reducing it to the bottom of the screen, such as... Figure 10 The narrow, tabular control 1001 shown in (b) includes playback controls, icons or names for text-to-speech functions, etc. It is understood that the tabular control 1001 shown in the figures is merely illustrative; the tabular control may have other forms, and more or fewer controls may be displayed on the tabular control 1001, without limitation. Alternatively, in some other embodiments of this application, the text-to-speech panel is minimized to disappear from the application interface, and when the user brings up the menu bar (e.g., by pulling down from the top of the screen), such as... Figure 10 As shown in (c), the mobile phone can display the reading control card 1002 in the menu bar.

[0152] Among them, when the text-to-speech control panel is first or every time it is shrunk, such as Figure 10 As shown in (b), the phone can also prompt the user: You can pull up the panel to bring up the text-to-speech control panel again. Then, after the phone detects the user's pull-up operation starting from control 1001, as shown... Figure 10 As shown in (a), the text-to-speech control panel is restored; or, the text-to-speech control panel is restored after the phone detects a pull-up operation from the bottom of the screen. When the text-to-speech control panel is first / every time it switches to the text-to-speech control card, as shown... Figure 11 As shown in (a), the phone can also prompt the user: You can click on the Xiaoyi Reading Card to bring up the reading control panel again.

[0153] The text-to-speech control card can include the name or icon of the text-to-speech function to indicate that the card is used to control this function. When the user taps the text-to-speech control card, the phone can restore the text-to-speech control panel. For example, in... Figure 11 In the case shown in (b), if the mobile phone detects that the user clicked on the reading control card, it can restore the display as shown in the image. Figure 11 The text-to-speech control panel is shown in (c). It is understood that the reading progress updates as the reading continues. The reading control card may also include playback controls for controlling the reading progress or speed, such as pause / resume, previous sentence, next sentence, fast forward, rewind, speed adjustment, refresh, progress bar, or exit. The reading card may also include one or more of the following information: reading status, elapsed reading time, remaining reading time, or the currently being read sentence. In some embodiments, the reading control card may also display relevant information about the target application, allowing the user to know the source of the text being read by the phone.

[0154] In the embodiments of the present application, the operation of the user indicating the minimization of the reading control panel can be referred to as a fourth operation. For example, the fourth operation can be the operation of the user clicking the minimization control 807.

[0155] For another example, after the phone detects the operation of the user minimizing the reading control panel, the reading control panel disappears from the current interface, and the phone can display the reading control floating ball 1201 as shown in (a) of FIG. 12B, and continue reading the text without interruption. Then, even if the phone displays other interfaces, the text reading process still continues. When the user clicks the reading control floating ball 1201, the phone can display the reading control panel again as shown in (b) of FIG. 12B. In some embodiments, the reading control floating ball 1201 can display relevant information (such as the application name or icon, etc.) of the target application to indicate to the user the source of the text being read. Figure 12 Figure 12

[0156] In some embodiments, the phone can have multiple text reading tasks, and when one of the text reading tasks is executed, the other text reading tasks are suspended. In one technical solution, the menu bar can include reading control cards corresponding to the multiple text reading tasks respectively, and the user can click the corresponding card according to the needs to instruct the phone to read the corresponding text. In another technical solution, each text reading task has a corresponding reading control floating ball, and the user can click the corresponding reading control floating ball according to the needs to instruct the phone to read the corresponding text. In another technical solution, the phone displays a reading control floating ball, and when there are multiple reading tasks, the user can click the reading control floating ball to switch between different reading tasks in sequence. The reading control floating ball can display the name or icon of the application corresponding to the current reading task, and other relevant prompt information to indicate to the user the source of the text being read. After the reading control floating ball is dragged to the edge of the screen by the user, a part of the reading control floating ball can be hidden to reduce the obstruction to the underlying interface. After the reading control floating ball is dragged into the screen by the user, the complete reading control floating ball is displayed.

[0157] For another example, the phone can adjust the size of the reading control panel window in response to the third operation of the user. For example, referring to (a)-(b) of FIG. 10, when the reading control panel is displayed in half of the screen, the phone can display the reading control panel in full screen in response to the upward operation of the user on the reading control panel, or in response to the upward operation of the user on the control 1003 or the nearby area. In addition, adjusting the size of the reading control panel window can also include minimizing the reading control panel, and switching the minimized reading control panel to the reading control panel displayed in half of the screen, etc. Figure 13A

[0158] ​​​In some embodiments, the text displayed on the text-to-speech control panel can scroll up / down in response to a fifth user action (e.g., a swipe). For example, see [link to relevant documentation]. Figure 13A In (b), the text-to-speech control panel is displayed in full screen. After the phone detects the user's upward swipe gesture on the text-to-speech control panel, as shown... Figure 13A As shown in (c) above, scroll up to display the text. Then, as... Figure 13A As shown in (d)-(e), if the phone detects a user's action on a sentence in the text-to-speech control panel (which can be called a second action, such as a tap, double tap, or long press), it will start reading from the user-specified sentence position and automatically move to the second line. Alternatively, as shown in (d)-(e), if the phone detects a user's action on a sentence in the text-to-speech control panel (such as a tap, double tap, or long press), it will start reading from the user-specified sentence position and move to the second line. Figure 13A As shown in (d)-(e), if the mobile phone detects the user's operation on a certain position on the reading control panel (also known as the second operation, such as a click, double-click or long press, etc., a touch operation on a certain position), it will start reading from the sentence corresponding to the position specified by the user and automatically move to the second line.

[0159] Alternatively, after the text scrolls on the text-to-speech control panel, if the phone does not detect the user's text-to-speech control operation within a preset time, then... Figure 13A As shown in (f), the currently being read statement is automatically moved to the top of the second line. Understandably, as time passes, the statement the phone is currently reading may no longer be the same statement that was being read before the text scrolling began.

[0160] It should be noted that when the text-to-speech control panel is displayed in half-screen mode, the text on the control panel can also be scrolled, and users can trigger reading from any position down the text. This will not be elaborated upon here. For an example, see [link to example]. Figure 13B In (a), the text-to-speech control panel is displayed in half-screen mode. If a user action (also known as a secondary action, such as a click, double-click, or long press) on a sentence in the text-to-speech control panel is detected, then... Figure 13B As shown in (b), the mobile phone reads the text aloud from the position specified by the user, highlighting the text and automatically moving it to the second line.

[0161] In some other embodiments, the mobile phone may also respond to a user's action on a statement by reading only that statement, but not continuing to read statements that follow that statement.

[0162] It should be noted that the text displayed on the reading control panel is divided in units of sentences, each sentence corresponds to a display control on the reading control panel, and each display control has its own identifier. After the mobile phone detects a touch operation of a user on a certain sentence that the user expects to read, the mobile phone triggers the display control corresponding to the sentence, and in response to the display control event, the mobile phone acquires the sentence corresponding to the display control and starts reading from the position of the sentence, thereby enabling the user to flexibly, finely, and accurately control reading in units of sentences.

[0163] In addition, in some other embodiments, after the text reading function is turned on, the total reading time display on the reading control panel stops displaying after a preset time length (for example, 10 seconds), and the sentence being read is automatically moved to the first line. In the subsequent reading process, the sentence being read is also automatically moved to the first line.

[0164] In the embodiments of the present application, the mobile phone for executing the text reading method can use various operating systems such as Android or Harmony, and can specifically use the form of software architecture such as layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. The software system and architecture form used by the mobile phone in the embodiments of the present application are not limited. The text reading method provided in the embodiments of the present application is described from the perspective of software implementation.

[0165] The application embodiment takes the Android system with layered architecture as an example to exemplarily illustrate the software structure of the mobile phone. The Android system can be divided into four layers, from top to bottom, the application program layer, the application program framework layer, the Android runtime and the system library, and the kernel layer. The application program layer can include a series of application program packages. The application program framework layer provides the application program interface (API) and the programming framework for the application program of the application program layer. The application program framework layer includes some predefined functions. The Android runtime includes the core library and the virtual machine. The Android runtime is responsible for the scheduling and management of the Android system. The core library contains two parts: one part is the function function called by the java language, and the other part is the core library of the Android. The application program layer and the application program framework layer run in the virtual machine. The virtual machine executes the java file of the application program layer and the application program framework layer into a binary file. The virtual machine is used to execute the object lifecycle management, the stack management, the thread management, the security and exception management, and the garbage collection and other functions. The system library can include a plurality of function modules. For example: the surface manager, the media library, the three-dimensional graphics processing library (for example: OpenGL ES), the 2D graphics engine (for example: SGL) and the like. The kernel layer is the layer between the hardware and the software. The kernel layer at least contains the display driver, the camera driver, the audio driver and the sensor driver.

[0166] In the application embodiment, as shown in Figure 14 The application program layer of the mobile phone can include the target application and the voice reading application, and the application program framework layer can include the window management service and the barrier-free management service and the like. The interface content of the target application is presented through the control, for example, the control can be a button, a progress bar, a single selection box or a text box and the like. The application interface can correspond to a view, each control can correspond to a view node, and the application interface can include a view root node and a plurality of view nodes. The voice reading application includes the text acquisition, the text processing, the reading user interface (UI) and the reading player and the like. The window management service is used for display control and management of the window program, can acquire the display screen size, judge whether there is a status bar, lock the screen, intercept the screen and the like.

[0167] In the embodiments of this application, the target application can register with an accessibility management service. The mobile phone can obtain node information (i.e., the first content of the first user interface) of all view nodes currently displayed on the application interface of the target application through the accessibility management service. This node information includes various information such as the width, height, attributes, or type of text, images, and controls. The voice reading application can obtain the text on the application interface from the node information and then read it aloud. Since the accessibility management service is a system service, it obtains the node information of all view nodes displayed on the mobile phone in real time; therefore, in this way, the mobile phone can obtain the text displayed on any application interface.

[0168] It's important to note that the information a phone obtains from node information includes the text visible on the application interface, as well as information such as node type and height. To ensure that the text-to-speech function can read aloud the relevant text after enabling it, the phone can preprocess this information to remove non-text and obtain the readable text.

[0169] Then, the phone divides the text into sentences. On one hand, it displays the pre-processed text to the user through the text-to-speech control panel. On the other hand, it uploads the pre-processed text to the TTS cloud server in the cloud through the text-to-speech player for text-to-speech conversion. The converted audio file (or audio segment, audio data, audio stream, speech stream, or voice packet, etc.) is returned to the text-to-speech player for playback, thus realizing the text-to-speech function.

[0170] based on Figure 14 The software modules shown, specifically the text-to-speech process provided in this application embodiment, may include:

[0171] The target application's interface content is displayed through a window management service, which then transmits the node information of the interface content to an accessibility management service. The text acquisition module of the voice-to-speech application obtains the node information from the accessibility management service. Subsequently, the text processing module within the voice-to-speech application preprocesses the node information obtained by the text acquisition module to obtain readable text. The readable text in the voice-to-speech application is not limited to Chinese; it also includes English and other languages. This embodiment of the application does not limit the language type of the readable text.

[0172] It is understood that mobile phones can also obtain the text on the application interface through other means besides accessibility management services, such as optical character recognition (OCR) or reader interfaces. This application does not limit the specific method of obtaining the text on the application interface.

[0173] In the embodiments of the present application, the preprocessing process of the text processing module includes filtering, cleaning, typesetting, and dividing sentences, etc.

[0174] The text in the Text-related control on the application interface is usually the main text of the interface content and can be read aloud. The text processing module can filter the obtained text data by retaining the text in the Text-related control, retaining the readable text part, and eliminating the text part that is not needed to be read aloud according to the type or height of the view node. The text information in the Text-related control is usually valid text information that can be read aloud. For example, the Text-related control includes but is not limited to:

[0175] Android control

[0176] android.widget.TextView

[0177] android.widget.EditText

[0178] android.widget.AutoCompleteTextView

[0179] android.widget.MultiAutoCompleteTextView

[0180] android.webkit.WebView

[0181] Huawei control

[0182] com.huawei.uikit.hwtextview.widget.HwTextView

[0183] com.huawei.uikit.hwedittext.widget.HwEditText

[0184] Text cleaning is used to filter the text data and eliminate the part that the user does not expect to read aloud, such as label, picture address, table, or website address, etc.

[0185] For example, the text processing module can eliminate the label type character that the user does not expect to read aloud by using the following regular expression:

[0186] private static final Pattern PATTERN_LABEL = Pattern.compile("(?:<([a-zA-Z]+?)(?:(:(?:\\s[^<>\\s]+?= [^<>\\s]+?)*?))?(?:= [^<>\\s]+?)?\\s?>.*?(?:< / \\1>)|\"+(?:< [a-zA-Z]+? (?:(:(?:= [^<>\\s]+?\\s? / >))|(:(?:\\s[^<>\\s]+?= [^<>\\s]+?)*?\\s? / >)))") ;

[0187] private static final Pattern PATTERN_IMG = Pattern.compile("[a-zA-Z0-9_]+?\\? [a-zA-Z0-9_]+?= [^&\\s]+(?:& [a-zA-Z0-9_]+?= [^&\\s]+)*") ;

[0188] private static final Pattern PATTERN_JAVA_SCRIPT = Pattern.compile("^javascript:.*? (?:(:(?:\\(.*?\\))|;)") ;

[0189] private static final Pattern PATTERN_INVALID_CHARACTERS = Pattern.compile("^[^\\u4e00-\\u9fa5a-zA-Z0-9]*$") ;

[0190] For another example, the text processing module can eliminate the characters that do not contain Chinese, English, numbers, etc. that the user does not expect to be read by the following way:

[0191] private static final Pattern PATTERN_INVALID_CHARACTERS = Pattern.compile("^[^\\u4e00-\\u9fa5a-zA-Z0-9]*$") ;

[0192] For another example, the text processing module can eliminate the meaningless content that the user does not expect to be read by the following way, for example, at least one uppercase, lowercase, number, and only contains uppercase, lowercase, number and some abnormal characters! #$%^&*_-=?.:+ / @, and the length is greater than 20:

[0193] private static final Pattern PATTERN_SPECIAL_CHARACTERS = Pattern.compile("^(?=.*[A-Z])(?=.*[a-z])(?=.*[0-9])[a-zA-Z0-9!#$%^&*_\\-=?.:+ / @]{20,}$").

[0194] The text processing module can perform sentence processing on the filtered and cleaned text, thereby identifying one or more sentences. For example, the text processing module can perform sentence processing by standard punctuation marks. For example, the following first type of punctuation marks can be used for sentence processing:

[0195] Period. A kind of end-of-sentence mark, mainly indicating the declarative mood of the sentence.

[0196] Question mark. A kind of end-of-sentence mark, mainly indicating the interrogative mood of the sentence.

[0197] Exclamation mark. A kind of end-of-sentence mark, mainly indicating the exclamatory mood of the sentence.

[0198] Ellipsis. A kind of mark, indicating the omission of some content in the segment and the discontinuity of meaning.

[0199] Period. A kind of end-of-sentence mark, mainly indicating the declarative mood of the sentence.

[0200] For another example, the following second type of punctuation marks cannot be used for sentence processing:

[0201] Quotation marks. A kind of mark, indicating directly quoted content or components that need to be specifically pointed out in the segment. Quotation marks also include “ ”, ‘ ’, etc.

[0202] Parentheses. A kind of mark, indicating the content of the segment, supplementary explanation, or other specific meaning statements. Parentheses also include [], 〔 〕, 【 】 and {}.

[0203] Title marks. A kind of mark, indicating the names of various works appearing in the segment.

[0204] In this way, when the second type of punctuation marks includes the first type of punctuation marks, the first type of punctuation marks cannot be used for sentence processing; when the first type of punctuation marks is not included in the second type of punctuation marks, the first type of punctuation marks can be used for sentence processing.

[0205] For another example, the following situations cannot use the period. as a sentence:

[0206] English abbreviation Mr. Wed.

[0207] Inter-digit (IP address, date, decimal number, segment number after the dot) (192.168.0.1 2009.3.13.1 16526 1.)

[0208] Email 89745859@qq.com

[0209] Website https: / / login.tom.com / login /

[0210] The text after the sentence processing can be further processed for layout. The text processing module can first sort the text according to the view node, and then adjust according to the height information of each view node, adjust the adjacent rows with uniform height to the same row. For example, each row on the application interface includes 4 application icons, if no layout is performed, each application icon name obtained by the speech reading application occupies a row; after layout, the names of the four application icons are arranged in the same row. Each sentence after layout is independent.

[0211] Then, the text processing module can send the entire text after layout to the reading UI for display. The reading UI performs window management and text display management, so that each sentence in the entire text is independently displayed as a paragraph on the reading control panel. Each sentence has a corresponding number. In some embodiments, each sentence can correspond to a display control respectively, and the mobile phone can determine that the user selects the sentence corresponding to the display control after detecting the operation of the user on the display control.

[0212] As described previously, the text is divided by sentence, and each sentence has a corresponding number, which can be used to index the sentence. The text processing module transmits the preprocessed text to the reading player. In some embodiments, as shown in Figure 14 The reading player estimates the total reading time according to the character number of the preprocessed text and the reading speed experience value and other parameters. In other embodiments, the reading player can send the preprocessed text to the TTS cloud server, and the TTS cloud server estimates the total reading time according to the character number of the text, the conversion and synthesis speed of the text to audio and other parameters, and notifies the reading player of the total reading time.

[0213] Then, the reading player is initialized, the estimated total reading time is displayed, and each sentence in the preprocessed text is sent to the TTS cloud server in the cloud according to the order. Referring to Figure 15, the TTS cloud server processes the text into speech sentence by sentence. For each sentence, the TTS cloud server can generate multiple audio segments. The TTS cloud server outputs and sends each synthesized audio segment (e.g., in the mp3 format) to the reading player on the mobile phone as soon as the audio segment is synthesized. For example, the TTS cloud server can use the interface of the hypertext transfer protocol (HTTP) protocol to download each audio segment to the mobile phone. Each audio segment has a short playback duration, such as 20 ms to 50 ms. Each audio segment corresponds to a small number of characters, such as one or two characters. The reading player caches the received audio segments in a cache 1 (or cache queue 1). The reading player starts playing as soon as the first audio segment is cached. In this way, the reading player starts playing as soon as the first audio segment is synthesized, that is, the user can quickly hear the reading voice after waiting for the first audio segment to be synthesized. The user's waiting time for starting reading is short, and the starting time can be reduced to the order of milliseconds. The reading is real-time, and the user experience is good.

[0214] As shown in Figure 16 , each sentence can be divided into multiple audio segments for speech conversion and broadcasting. Each sentence can be understood as a text batch. The TTS cloud server can return the data of several audio segments to the reading player, which is stored in the cache 1 after being received, and then taken out for broadcasting to achieve streaming and real-time playback.

[0215] In addition, after the audio file of each sentence is cached, the TTS cloud server sends an end identifier, and the reading player triggers the sending of the next sentence to the TTS cloud server. As shown in Figure 15 , the reading player adds the cached audio file of the entire sentence to a cache 2 (or cache queue 2). As shown in Figure 17 , the entire text (i.e., the text corresponding to the application interface, which can also be referred to as an article) can include multiple sentences (i.e., multiple batches), and the cache 2 can cache the complete audio file of each batch. Moreover, because the TTS synthesis speed is much faster than the playback speed, all audio can be synthesized and cached quickly during playback. In this way, the sentences cached in the cache 2 generally precede the sentences being broadcast. When the user drags the progress bar (which can be dragged backward or forward), if the progress after the dragging falls into the cache 2, the corresponding audio file can be obtained from the cache 2 to achieve immediate playback without waiting, and the real-time responsiveness is strong, and the user experience is good. If the progress after the user drags the progress bar falls into a sentence that does not exist in the cache 2, the synthesis and caching of the text from the sentence to the end are triggered immediately.

[0216] That is, as shown inFigure 17 As shown, the cache 2 is used to cache audio information corresponding to complete single sentence texts. When audio data of a certain sentence is attempted to be obtained from the cache 2, if there is no such data, it indicates that there is no synthesized audio file for the sentence, and at this time, a synthesis and decoding process of the sentence is initiated in real time. During the synthesis and decoding process, audio data is returned several times, and after being received, the audio data is cached in the queue, and the audio data is taken out to start the broadcast, so as to realize streaming real-time broadcast.

[0217] It should be noted that the cache 1 and the cache 2 can be different cache queues, or can be two parts in the same cache queue, and the embodiments of the present application are not limited thereto. The cache design of the cache 1 and the cache 2 described above can achieve the effect that the play control and the concurrent processing do not affect each other.

[0218] In addition, the reading player can also perform play control on the reading process. For example, for the operation of dragging the progress bar by the user, the reading player finds the corresponding sentence according to the proportion of the drag point to the progress bar, and starts the broadcast from the sentence, or sends the sentence and the subsequent text to the TTS cloud server for voice synthesis. The fast forward and fast backward operations can also be converted into the proportion to perform corresponding processing.

[0219] For the operation of adjusting the broadcast speed by the user, the reading player adjusts the broadcast speed by setting a corresponding control parameter such as AudioTrack.

[0220] For the operation of pausing / resuming by the user, the reading player pauses / resumes the request of audio data to the TTS cloud server, and pauses / resumes the play of AudioTrack.

[0221] For the operation of clicking the refresh button by the user, the voice reading application reacquires text data from the application interface, and reperforms the play process.

[0222] For the operation of clicking the minimize “—” control, the Back button, the Menu button and the like by the user, the reading control panel is minimized, the background continues to play, and the reading control panel is shrunk to the menu bar.

[0223] For the operation of clicking the exit “X” control by the user, the voice reading application stops reading the text, releases the resources occupied by the voice reading application, and destroys the reading UI.

[0224] For example, after a user opens an article on a news app, they instruct the app to activate the text-to-speech function. The app retrieves the article from the top of the search results and preprocesses it. The text is then sent to a TTS (Text-to-Speech) cloud server. The audio data output from the TTS cloud server is continuously transmitted to the app, which caches the data and begins playback. Users can then pause, resume, fast forward, rewind, refresh, or exit the app.

[0225] also, Figure 18 This diagram illustrates a logical block diagram of a text-to-speech method provided in an embodiment of this application. For example... Figure 18 As shown, the text player transmits text to the TTS cloud server sentence by sentence. Upon receiving a sentence, the TTS cloud server converts the text into an audio segment corresponding to the speech stream and transmits the compressed audio file (e.g., compressed into an MP3 format) to the text player. The text player decodes the received audio file of the sentence and stores the corresponding audio file in cache 1, segment by segment. For example, cache 1 includes audio segment data packets (package) 1, package 2, ..., package n. Each package represents one packet of speech data, corresponding to one audio segment. After a sentence is cached, it is sent to cache 2. Furthermore, after the audio file of a sentence is cached in cache 1, the text player can send the next sentence to the TTS cloud server and repeat the above process. For example, cache 2 includes sentence data packets (sentence) 1, sentence 2, ..., sentence n. Each sentence corresponds to a complete audio file of a sentence.

[0226] like Figure 18 As shown, when the text being read is the same sentence as the audio file stored in cache 1, AudioTrack retrieves the audio file from cache 1 and plays it. When the text being read is not the same sentence as the audio file stored in cache 1, AudioTrack reads the audio file for that sentence from cache 2 and plays it. Furthermore, after a sentence of text has finished playing, the text reader controls the UI to highlight the next sentence and scroll the text.

[0227] like Figure 19As shown, if the to-be-read sentence changes due to user instructions such as fast forward, fast backward, dragging the progress bar, continuous clicking of the previous sentence or the next sentence, etc., the reading player can calculate the proportion of the text progress corresponding to the user operation, find the corresponding starting sentence, that is, find the to-be-read sentence. If the reading player determines that the to-be-read sentence exists in the cache 2, the AudioTrack reads the audio file of the sentence from the cache 2 and immediately plays. If the reading player determines that the to-be-read sentence does not exist in the cache 2, the TTS cloud server is triggered to send the sentence. In this case, the cache 1 is caching and reading the same sentence, and the AudioTrack reads the audio file from the cache 1 and plays in real time.

[0228] The embodiment of the application further provides a module interaction timing diagram. As shown in Figure 20 The mobile phone end includes a northbound interface, a state machine, a reading player and a decoding module, and the TTS cloud server includes a synthesis module. In the embodiment of the application, the northbound interface can include a play interface, a pause interface or a stop interface, etc. The play control interface can be called by the upper voice reading application.

[0229] After the start play interface of the northbound interface is called by the voice reading application, the state machine of the reading player selects to ignore the play or to execute the play according to the current play state of the reading player. If the state machine determines to execute the play according to the current state, the play control logic is started to be executed in units of sentences. The play control logic can specifically include that the reading player (which can also be called a play control center) attempts to obtain the audio data of the to-be-played sentence from the cache 2. If the cache 2 does not have the audio data of the sentence, the synthesis module in the TTS cloud server is sent for audio synthesis, and the synthesized audio file is decoded by the decoding module of the mobile phone end. In the synthesis and decoding process, the intermediate audio file is returned to the reading player. Then, the reading player can store the audio file in the cache 1. The reading player can call the audio.write interface to write data, and constantly read data from the cache 1 and write. The reading player can perform a single sentence play start callback. Moreover, after the synthesis and decoding of the whole sentence are completed, the audio file is stored in the cache 2. After the text play is completed, the reading player can perform a single sentence play end callback. The reading player automatically updates the play sequence and starts the play of the next sentence, and then repeats the play control logic.

[0230] In summary, this application provides a text-to-speech method. A mobile phone can read text aloud from any application interface to the user and provides a text-to-speech control panel for user interaction. The phone can intuitively, explicitly, and dynamically inform the user of the current reading position through highlighted text, preventing the user from getting lost. Users can easily perform reading control operations such as pause / resume, previous / next sentence, fast forward, rewind, drag the progress bar, adjust speed, and refresh the interface text. They can also freely click on any paragraph in the text-to-speech control panel to trigger reading from that point, allowing users to easily re-listen to parts they didn't understand or skip uninteresting sections to important parts of interest. After the user drags the progress bar, the highlighted text dynamically follows the changing sentences, providing real-time information about the current reading position and improving the interactive experience. Furthermore, the caching design enables quick, real-time start of reading and rapid resumption of reading after changing the reading position (such as fast forward, rewind, or dragging the progress bar), resulting in a better user experience.

[0231] In addition, in conjunction with the above embodiments and corresponding drawings, another embodiment of this application provides a method for reading aloud displayed content. This method can be used in a [missing information - likely a specific application or system]. Figure 1 The structure shown is implemented on an electronic device. The relevant content of the above embodiments and accompanying drawings applies to this embodiment and will not be repeated here. See also... Figure 21 The method includes:

[0232] 2101. The electronic device displays the first user interface.

[0233] As mentioned above, the first user interface can be any interface displayed on the mobile phone, the interface of a first specific application, or the interface of an application other than the first specific application, etc.

[0234] 2102. The electronic device receives the user's first operation.

[0235] The first operation can be used to instruct the electronic device to activate the text-to-speech function and obtain the text content to be read aloud. For example, the first operation can be the user triggering the voice command "Hey Celia, read it aloud", or the user clicking the read-aloud control, etc.

[0236] 2103. The electronic device responds to the first operation and obtains the first content of the first user interface.

[0237] The first content includes text information on the first user interface. For example, the first content may be node information of each view node on the first user interface.

[0238] 2104、The electronic device displays a second user interface, and content displayed by the second user interface includes text in the first content; wherein the second user interface covers part of the display area of the first user interface.

[0239] For example, the second user interface can be the interface corresponding to the reading control panel described above.

[0240] 2105、The electronic device reads a first sentence in the content of the second user interface, and displays identification information of the first sentence being read in the text corresponding to the second user interface.

[0241] The identification information is used to prompt the user about the content of the sentence being read.

[0242] The electronic device can read the text on the first user interface and display the identification information through the second user interface superimposed on part of the display area of the first user interface, thereby dynamically and intuitively prompting the user about the content being read, and the interaction with the user is good, and the user experience is good. For example, the first user interface can be the interface of a target application, and the second user interface can be the interface corresponding to the reading control panel.

[0243] It can be understood that the above is described by taking the electronic device as a mobile phone as an example, and when the electronic device is a tablet computer or other device, the method described in the above embodiments can still be used for text reading, which will not be described here.

[0244] In another text reading scheme, as shown in (a) of Figure 22 , the voice assistant of the mobile phone can support text reading for a specific manufacturer's browser and part of the applications of the manufacturer. This scheme is limited to the internal application of a specific manufacturer and is not applicable to other applications outside the specific manufacturer. In this scheme, the reading progress can be advanced or reversed in time, and the highlighted text can be changed in real time when the reading progress bar is dragged. However, it does not support the user to click on any paragraph of the article to trigger reading from this point.

[0245] In another text reading scheme, as shown in (b) of Figure 22 , the mobile phone can enter the accessibility mode in response to the user's operation on the floating control, and then read the text in this mode. Since the accessibility mode is for people with disabilities, it is not convenient for normal users. In addition, this scheme does not display the content being read, the interaction with the user is limited, and it does not support dragging the progress bar, highlighting the text, clicking to start reading, and other interactions. Moreover, the floating control displayed on the interface also disturbs the user.

[0246] It can be understood that, in order to realize the above functions, the electronic device comprises hardware and / or software modules corresponding to the respective functions. The algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is realized in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered beyond the scope of the present application.

[0247] The embodiments of the present application can divide the functional modules of the electronic device according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The integrated module can be realized in the form of hardware. It should be noted that the division of the modules in the embodiments is illustrative and is only a logical functional division. Actual implementation can have another division manner. For example, in some embodiments, the electronic device can include display unit, processing unit, and reading unit modules.

[0248] The embodiments of the present application also provide an electronic device, which includes one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are configured to store computer program codes, the computer program codes including computer instructions, which, when executed by the one or more processors, cause the electronic device to perform the steps of the above-mentioned related method and realize the text reading method in the above-mentioned embodiments.

[0249] The embodiments of the present application also provide an electronic device, as shown in Figure 23 The electronic device includes a display screen (or screen) 2301, one or more processors 2302, a memory 2303, and one or more computer programs 2304. The above-mentioned devices can be connected through one or more communication buses 2305. The one or more computer programs 2304 are stored in the memory 2303 and are configured to be executed by the one or more processors 2302. The one or more computer programs 2304 include instructions that can be used to execute each step in the above-mentioned embodiments. The above-mentioned method embodiments involve all related content of each step, which can be cited to the function description of the corresponding entity device, and will not be described here.

[0250] For example, the processor 2302 can be specifically the processor 110 shown in Figure 1 The memory 2303 can be specifically the internal memory 121 shown in Figure 1 The display screen 2301 can be specifically the display screen 120 shown in Figure 1The display screen 194 is shown.

[0251] Embodiments of the present application further provide a computer readable storage medium, which stores computer instructions, when the computer instructions are run on an electronic device, the electronic device is caused to perform the above related method steps to implement the text reading method in the above embodiments.

[0252] Embodiments of the present application further provide a computer program product, when the computer program product is run on a computer, the computer is caused to perform the above related steps to implement the text reading method performed by the electronic device in the above embodiments.

[0253] In addition, embodiments of the present application further provide a device, which can be a chip, a component or a module, and the device can include a processor and a memory connected to each other; wherein the memory is configured to store computer execution instructions, and when the device is running, the processor can execute the computer execution instructions stored in the memory to enable the chip to perform the text reading method performed by the electronic device in the above method embodiments.

[0254] The electronic device, the computer readable storage medium, the computer program product or the chip provided by the embodiments can be used to perform the corresponding method provided above, and thus the beneficial effects achieved thereby can refer to the beneficial effects of the corresponding method provided above, which will not be described herein again.

[0255] Through the above description of the embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0256] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0257] The units described as separate components may or may not be physically separate, and the components displayed as units may be a physical unit or multiple physical units, that is, may be located in one place, or also can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0258] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0259] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical scheme of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions to make a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0260] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for reading out content of a display, applied to an electronic device, the method comprising: receiving a request for reading out content of a display; and reading out the content of the display in response to the request. The method comprises: displaying a first user interface; the first user interface is an arbitrary application interface; receiving a first operation of a user; in response to the first operation, obtaining first content of the first user interface by a system-level accessibility management service; displaying a second user interface, the content displayed by the second user interface comprising text in the first content; wherein the second user interface is displayed on the upper layer of the original first user interface and shields part of the display area of the first user interface; reading a first sentence in the content of the second user interface, and displaying identification information of the corresponding text of the first sentence being read in the second user interface; detecting a second operation on the second user interface; in response to the second operation on the second user interface, controlling the electronic device to read a second sentence corresponding to the second operation based on the text position or text content corresponding to the second operation, or read from the position of the second sentence.

2. The method of claim 1, wherein, Before displaying the second user interface, the method further comprises: identifying one or more sentences from the obtained first content of the first user interface.

3. The method of claim 1, wherein, The method further comprises: displaying identification information of the corresponding text of the second sentence being read in the second user interface.

4. The method of claim 1, wherein, The method further comprises: The second operation is a touch operation of the user on the text position corresponding to the second sentence displayed by the second user interface; based on the detected touch operation, confirming the obtaining of the voice corresponding to the second sentence and reading.

5. The method of claim 4, wherein, The method further comprises: displaying the second user interface based on the identified one or more sentences, each of the one or more sentences displayed by the second user interface corresponding to a control; The method further comprises: based on the second operation corresponding to the touch operation, triggering the control corresponding to the second sentence; in response to the control event, obtaining the voice of the second sentence corresponding to the control and reading.

6. The method according to any one of claims 2-5, characterized in that, The method further comprises: indexing the identified one or more sentences; The method further comprises: the content displayed by the second user interface further comprises a play control, which can be used to control the progress of reading or the speed of reading; the play control comprises a progress bar, the progress bar matches the index of the one or more sentences, and the progress bar is controlled based on the granularity of the identified sentences; the play control comprises a down control for controlling the current reading progress to switch to the next sentence; and / or, the play control comprises an up control for controlling the current reading progress to switch to the previous sentence.

7. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: detecting a third operation of a user; in response to the third operation, adjusting the size of the window corresponding to the second user interface.

8. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: detecting a fourth operation of a user; in response to the fourth operation, minimizing the window corresponding to the second user interface, and continuing to read the current text content without interruption. After minimizing the window corresponding to the second user interface, the method further includes: restoring the second user interface by the card corresponding to the reading function, displaying the current reading text content and progress.

9. The method of claim 6, wherein, The playback control includes a refresh control, and the method further includes: In response to the refresh control, stopping reading the current text and reacquiring the second content of the first user interface for reading.

10. The method of claim 2, wherein, The method further includes: sending the recognized one or more sentences to a server for text-to-speech processing; receiving and caching the voice obtained after the server processing; reading the voice obtained after the server processing in real time, or sequentially caching the voice obtained after the server processing according to the recognized one or more sentences.

11. The method of claim 10, wherein, The method further includes: The electronic device caches the voice obtained after the server processing through two queues; The first queue is used to cache one or more voice packets of the sentence currently being processed by the server; The second queue is used to cache the voice corresponding to the one or more sentences processed by the server; The reading of the first sentence in the content of the second user interface includes: If the current voice received from the server is the voice of the first sentence, the voice corresponding to the first sentence is obtained from the first queue for reading; or If the current voice received from the server is not the voice of the first sentence, the voice corresponding to the first sentence is obtained from the second queue for reading.

12. The method of claim 2, wherein, Before displaying the second user interface, the method further includes: processing the obtained first content of the first user interface to remove non-text information to obtain the text in the first content; displaying the processed text on the second user interface.

13. The method of claim 12, wherein, Before displaying the second user interface, the method further includes: punctuating the processed text according to special punctuation marks to obtain the recognized one or more sentences.

14. The method according to any one of claims 1-5, characterized in that, The method further includes: In response to the user's voice instruction, controlling the reading progress or the reading speed.

15. The method according to any one of claims 1-5, characterized in that, The method further includes: detecting a fifth operation of the user; In response to the fifth operation of the user, scrolling the text content on the second user interface and continuing to read the current text content without interruption.

16. An electronic device, comprising: It includes: a screen for displaying a user interface; one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to perform the method of reading the displayed content according to any one of claims 1-15.

17. A computer-readable storage medium, characterized in that, It includes computer instructions that, when executed on a computer, cause the computer to perform the method of reading the displayed content according to any one of claims 1-15.

18. A computer program product, characterised in that, When the computer program product is executed on a computer, it causes the computer to perform the method of reading the displayed content according to any one of claims 1-15.

Citation Information

Patent Citations

  • Text-to-speech interface featuring visual content supplemental to audio playback of text documents

    WO2020023070A1