Screen text information reading aloud method and apparatus, storage medium, and electronic apparatus

By judging the currently opened application type in the terminal screen and calling the corresponding text acquisition method, the problem that the terminal screen reading function cannot cover all interfaces of the terminal is solved, and automatic acquisition and reading of invisible parts of text content in browser applications and local applications is realized.

WO2025091942A1PCT designated stage expired Publication Date: 2025-05-08ZTE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/100219
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-06-19
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

In the prior art, the screen reading function of the terminal cannot cover all interfaces of the terminal, especially in browser applications and local applications, and text content of invisible parts of the screen cannot be automatically obtained.

Method used

By judging the currently opened application type in the terminal screen, calling the corresponding text acquisition method, obtaining text information in the current page, including the visible and invisible areas of the screen. For browser applications, use the server to perform URL parsing to obtain text information.

Benefits of technology

It realizes voice broadcasting of all terminal interfaces, covering part of the invisible text content in browser applications and local applications, and meets the interface reading needs of multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024100219_08052025_PF_FP_ABST
    Figure CN2024100219_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a screen text information reading aloud method and apparatus, a storage medium, and an electronic apparatus. The method comprises: determining the application type of a currently opened application in a screen of a terminal, wherein the term "application" comprises a browser application; calling a corresponding text acquisition mode on the basis of different application types, and acquiring text information of a current page of the application, wherein the current page comprises a screen visible area and a screen invisible area that can be displayed only by means of a page turning or pull-down operation; and then reading aloud the acquired text information.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for broadcasting screen text information, storage medium and electronic device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on Chinese patent application CN 202311461294.3, filed on November 03, 2023, entitled “Method and device for broadcasting screen text information, storage medium and electronic device”, and claims the priority of the patent application, and all the disclosed contents are incorporated into this application by reference. Technical Field

[0003] The embodiments of the present disclosure relate to the field of communications, and in particular, to a method and device for broadcasting screen text information, a storage medium, and an electronic device. Background Art

[0004] In the related art, screen reading for the elderly is not yet fully adapted. Some terminals cannot broadcast the content of browser applications or support reading and broadcasting of local applications. Some terminals also use the talkback function for blind users, but it can only read aloud the text of the currently clicked content.

[0005] Summary of the Invention

[0006] The embodiments of the present disclosure provide a method and device for broadcasting screen text information, a storage medium, and an electronic device, so as to at least solve the problem in the related art that the screen reading function of a terminal cannot cover all interfaces of the terminal.

[0007] According to one embodiment of the present disclosure, a method for broadcasting screen text information is provided, comprising: determining the application type of an application currently opened on a terminal screen, wherein the application includes a browser application; calling a corresponding text acquisition method according to different application types to acquire text information on a current page of the application, wherein the current page includes a visible area of ​​the screen and an invisible area of ​​the screen that can only be displayed by turning the page or pulling down the page; and broadcasting the acquired text information.

[0008] According to another embodiment of the present disclosure, a device for broadcasting screen text information is provided, including: a scene recognition module, a text recognition module, and a voice broadcast module, wherein the scene recognition module is configured to determine the application type of an application currently opened on the screen of a terminal; the text recognition module is configured to call a corresponding text acquisition method according to different application types to obtain text information in a current page of the application, wherein the current page includes a visible area of ​​the screen and an invisible area of ​​the screen that needs to be displayed according to page turning and pull-down operations; the voice broadcast module is configured to broadcast the acquired text information.

[0009] According to another embodiment of the present disclosure, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when running.

[0010] According to another embodiment of the present disclosure, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] FIG1 is a hardware structure block diagram of a mobile terminal for a method for broadcasting screen text information according to an embodiment of the present disclosure;

[0012] FIG2 is a schematic diagram of a network structure of an operation method for broadcasting screen text information according to an embodiment of the present disclosure;

[0013] FIG3 is a flow chart of a method for broadcasting screen text information according to an embodiment of the present disclosure;

[0014] FIG4 is a structural block diagram of a device for broadcasting screen text information according to an embodiment of the present disclosure;

[0015] FIG5 is a structural block diagram of a device for broadcasting screen text information according to an embodiment of the present disclosure;

[0016] FIG6 is a structural block diagram of a device for broadcasting screen text information according to an embodiment of the present disclosure;

[0017] FIG7 is a schematic diagram of the switch logic of the screen reading initiation module according to an embodiment of the present disclosure;

[0018] FIG8 is a logic diagram of a scene recognition module according to an embodiment of the present disclosure;

[0019] FIG9 is a logical diagram of a text recognition module according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the accompanying drawings and in combination with the embodiments.

[0021] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0022] Currently, most screen reading solutions on the market use screenshot recognition to obtain text. This method cannot obtain text content in the invisible part of the browser scene.

[0023] See Patent 1: CN115248650.

[0024] Patent 1 receives a first input from a user to the electronic device, determines at least one area to be processed on the display interface of the electronic device in response to the first input, receives a second input from the user to the electronic device, determines a first target area from the area to be processed in response to the second input, and reads aloud the text content in the first target area.

[0025] There are two obvious differences between the embodiment of the present disclosure and Patent 1:

[0026] 1. Patent 1 requires the user to select an area on the screen to be read aloud, and only the selected part is read aloud.

[0027] 2. Patent 1 cannot achieve automatic reading of the invisible part of the browser text. There are two cases of invisibility:

[0028] (1) You need to scroll down or turn the page to see the rest of the page. (2) You can correctly identify the end of the article and the pop-up advertisement section.

[0029] See Patent 2: 2022CN-0681075.

[0030] The sources and methods of adding text snippets to the reading resource collection of Patent 2 may include at least one of the following: taking a screenshot of the content displayed on the screen of the electronic device, identifying the screenshot, and storing the text snippets obtained by identifying the screenshot in the reading resource collection; capturing an image through the camera of the electronic device, identifying the captured image, and storing the text snippets obtained by identifying the image in the reading resource collection, etc.

[0031] There are two obvious differences between the embodiment of the present disclosure and Patent 2:

[0032] 1. Patent 2 uses the method of screenshot or video recording to first obtain the screenshot, and then obtain the text through image recognition.

[0033] 2. This recognition method also cannot automatically obtain text in the invisible part of the screen. There are two cases of invisible text:

[0034] (1) You need to scroll down or turn the page to see the rest of the page. (2) You can correctly identify the end of the article and the pop-up advertisement section.

[0035] In an embodiment of the present application, the browser application includes a webview control, and some third-party browser applications have also extended their own webview controls. The accessibility of the Android native system is to find the text in the textview in the view control by traversing the view control of the interface. Based on this, the Android native system method is unable to parse the webview control, resulting in the inability to obtain text. In one embodiment of the present application, application classification is performed by determining whether the scenario includes the use of the webview control. Through the embodiment of the present application, the reading of all interfaces of the terminal can be covered to meet the reading needs of applications in a variety of different scenarios. Different scenarios include local applications, social applications, or browser applications, but are not limited to these scenarios.

[0036] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking operation on a mobile terminal as an example, FIG1 is a hardware structure block diagram of a mobile terminal of a method for broadcasting screen text information in an embodiment of the present disclosure. As shown in FIG1 , the mobile terminal may include one or more (only one is shown in FIG1 ) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that the structure shown in FIG1 is only for illustration and does not limit the structure of the mobile terminal. For example, the mobile terminal may also include more or fewer components than those shown in FIG1 , or have a configuration different from that shown in FIG1 .

[0037] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for broadcasting screen text information in the embodiment of the present disclosure. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the mobile terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0038] The transmission device 106 is used to receive or send data via a network. Examples of such networks may include wireless networks provided by the mobile terminal's communications provider. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0039] The embodiment of the present application can run on the network architecture shown in Figure 2. Figure 2 is a schematic diagram of the operating network structure of the method for broadcasting screen text information according to the embodiment of the present disclosure. As shown in Figure 2, the network architecture includes: a terminal and a server, wherein the terminal and the server can communicate in a network interactive manner, the terminal is used to provide a page to be read aloud, wherein the page includes a visible area and an invisible area, and the server is used to extract text information from the scene page of the webview of the browser application, and then transmit the text information to the terminal for reading aloud.

[0040] In this embodiment, a method for broadcasting screen text information running on the above-mentioned mobile terminal or network architecture is provided. FIG3 is a flow chart of the method for broadcasting screen text information according to an embodiment of the present disclosure. As shown in FIG3 , the process includes the following steps:

[0041] Step S302, determining the application type of the application currently opened on the screen of the terminal, where the application includes a browser application;

[0042] In the actual implementation process, applications also include local applications and social applications. As for application scenarios, local applications basically do not use webview scenarios, social applications partially use webview scenarios, and browser applications all use webview scenarios.

[0043] Step S304: Calling a corresponding text acquisition method according to different application types to acquire text information on the current page of the application, wherein the current page includes a visible area of ​​the screen and an invisible area of ​​the screen that can only be displayed by turning the page or pulling down the screen;

[0044] In an exemplary embodiment, corresponding text acquisition methods are called according to different application types to obtain text information in the current page of the application, including: when the application type is a browser application, the terminal reports the Uniform Resource Locator (URL) information of the browser application to the server; the terminal receives the text information obtained by the server by parsing the URL information.

[0045] In the browser application scenario of the embodiment of the present application, by receiving the user's operating instructions to the electronic device, in response to the operating instructions, the page to be broadcast of the browser application on the electronic device is determined, and the text content in the page to be broadcast is read aloud. Among them, the electronic device includes at least a mobile phone terminal, and the operating instructions include at least instructions such as turning on the reading switch and confirming the reading. The page to be broadcast includes a visible area of ​​the screen, and an invisible area of ​​the screen that can only be displayed according to page turning and pull-down operations. The present application realizes the acquisition of the text content of the invisible part in the browser scenario, and the automatic reading of the invisible part of the browser text. Among them, the acquisition of the invisible part includes not needing to pull down or turn the page to see other parts of the page, or identifying the end of the article and popping out of the advertising part, etc.

[0046] In an exemplary embodiment, the server uses an anti-crawler mechanism to parse URL information.

[0047] In the actual implementation process, the use of anti-crawler mechanism to parse URL information can avoid the introduction of useless text information such as text obfuscation and image disguise.

[0048] In an exemplary embodiment, the server configures the authentication environment for the authenticated website through a backend browser.

[0049] In the actual implementation process, the website content cannot be obtained for the authenticated website, and an identical environment needs to be configured in the back-end browser.

[0050] In an exemplary embodiment, the server uses a multi-threaded concurrent mechanism to parse URL information.

[0051] In the actual implementation process, there may be hundreds, thousands, or even tens of thousands of users requesting to obtain text information at the same time, so concurrent configuration optimization is required.

[0052] In an exemplary embodiment, the corresponding text acquisition method is called according to different application types to obtain text information in the current page of the application, and also includes: when the application type is a local application or a social application, the terminal uses the accessibility method to obtain text information.

[0053] Step S306: broadcast the acquired text information.

[0054] During actual implementation, a main switch for the screen reading function may be set on the terminal. Screen reading can only be performed when the main switch for the screen reading function is turned on.

[0055] Through the above steps, by determining the application type of the currently open application on the terminal screen, where the application includes a browser application; calling the corresponding text acquisition method according to the different application types to obtain the text information on the current page of the application, where the current page includes the visible area of ​​the screen and the invisible area of ​​the screen that can only be displayed by turning the page or pulling down; and then announcing the acquired text information. This solves the problem in the related art that the screen reading function of the terminal cannot cover all the terminal interfaces, and achieves the effect of meeting the interface reading needs of different terminal application scenarios.

[0056] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the embodiment of the present disclosure is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD-ROM), including a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method described in the embodiment of the present disclosure.

[0057] In this embodiment, a device for broadcasting screen text information is also provided. The device is used to implement the above-mentioned embodiments and preferred embodiments. Details already described are not repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0058] Figure 4 is a structural block diagram of a device for broadcasting screen text information according to an embodiment of the present disclosure. As shown in Figure 4, the broadcasting device 40 includes: a scene recognition module 410, a text recognition module 420, and a voice broadcasting module 430, wherein the scene recognition module 410 is configured to determine the application type of the application currently opened on the screen of the terminal, and the application includes a browser application; the text recognition module 420 is configured to call the corresponding text acquisition method according to different application types to obtain text information in the current page of the application, wherein the current page includes a visible area of ​​the screen and an invisible area of ​​the screen that needs to be displayed according to page turning and pull-down operations; the voice broadcasting module 430 is configured to broadcast the acquired text information.

[0059] Figure 5 is a structural block diagram of a device for broadcasting screen text information according to an embodiment of the present disclosure. As shown in Figure 5, in addition to all the modules shown in Figure 4, the broadcasting device 50 also includes: a server 510 set independently of the terminal, wherein the server 510 is configured to receive the uniform resource locator URL information of the browser application from the terminal when the application type is a browser application, and parse the URL information to obtain text information.

[0060] In the actual implementation process, the server is set up independently of the terminal in the embodiment of the present disclosure for illustration only. The function of the server to obtain text information from the application interface of the browser application can also be integrated on the terminal without using an independent server. The implementation method can adopt the method of adding a new chip or text recognition application on the terminal, which will not be repeated here.

[0061] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0062] An embodiment of the present disclosure further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when running.

[0063] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0064] An embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0065] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0066] The examples in this embodiment can refer to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.

[0067] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present disclosure can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented using program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be made into individual integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Thus, the embodiments of the present disclosure are not limited to any specific combination of hardware and software.

[0068] In order to enable those skilled in the art to better understand the technical solutions of the embodiments of the present disclosure, the following is explained in conjunction with the embodiments.

[0069] The embodiments of the present disclosure provide a method and device for broadcasting screen text information, which can cover voice broadcasting of all interfaces of the terminal, mainly including three scenarios: broadcasting of local applications, broadcasting of social applications, or broadcasting of browser applications.

[0070] Figure 6 is a structural block diagram of a device for broadcasting screen text information according to an embodiment of the present disclosure. As shown in Figure 6, the broadcasting device 60 includes: a screen reading initiation module 610, a scene recognition module 620, a text recognition module 630, and a voice broadcast module 640. Among them, the screen reading initiation module 610 is used to turn on and off the screen reading function. The scene recognition module 620, the text recognition module 630, and the voice broadcast module 640 have the same functions as the scene recognition module, text recognition module, and voice broadcast module in the above embodiments, among which the scene recognition module 620 is used to determine whether the application currently opened to be read aloud is a browser application. The text recognition module 630 is used to call different text acquisition methods for different applications to obtain text information. The voice broadcast module 640 is used to broadcast the acquired text information.

[0071] The broadcasting device 60 shown in FIG6 can be provided on the terminal in the form of various functional modules.

[0072] In the embodiment of the present disclosure, the acquisition of text information of the browser application is performed through the server, including the following contents: (1) the user opens the browser; (2) the broadcast function is called through a voice command; (3) it is confirmed that voice broadcast is required, and the URL of the current traffic device is obtained in the framework; (4) the URL is sent to the server side; (5) the URL is processed on the server side; (6) the server crawls the web page text content and pushes it to the designated client (mobile phone side); (7) the mobile phone side processes the text and reads it aloud.

[0073] Server-side URL processing includes: a) website parsing with anti-crawler mechanisms (such as text obfuscation and image camouflage). b) authentication environment configuration. Authenticated websites cannot access their content, requiring a completely identical environment to be configured in the backend browser. c) concurrent processing. Hundreds, thousands, or even tens of thousands of user requests can be processed simultaneously, requiring configuration optimization when concurrency is required.

[0074] In the above (2) invoking the broadcast function through voice commands, invoking the broadcast function through voice commands is only one way to invoke the broadcast function through voice commands. Correspondingly, in the embodiment of the present disclosure, the main switch of the screen reading function is set through the above screen reading initiation module. In actual implementation, the main switch of the screen reading function can be set in the voice assistant, and a floating window can be directly displayed on the screen as an entrance, etc.

[0075] FIG7 is a schematic diagram of the switch logic of the screen reading initiation module according to an embodiment of the present disclosure, as shown in FIG7 , including:

[0076] Step S702: Add a "screen reading" switch via the voice assistant;

[0077] In one embodiment, the main switch of the screen reading function can be set in the setting module of the mobile terminal;

[0078] Step S704, turn on the "Screen Reading" switch;

[0079] In one embodiment, screen reading can only be performed when the screen reading function is turned on. Therefore, it is necessary to turn on the "screen reading" switch through the mobile phone settings;

[0080] Step S706: Turn on the reading function through the voice command "screen reading".

[0081] In one embodiment, after the user turns on the screen reading main switch, he needs to call the voice assistant and turn on the reading function through the voice assistant. He can turn on the reading function by saying "screen reading" as a voice command to read the screen.

[0082] FIG8 is a logic diagram of a scene recognition module according to an embodiment of the present disclosure. As shown in FIG8 , the following steps are included:

[0083] Step S802: Obtain information of the currently opened top application.

[0084] Step S804: Determine the application type of the current top application.

[0085] Step S806: calling different text recognition methods for different application types.

[0086] As shown in FIG8 , in an embodiment of the present disclosure, when performing scene recognition, it is necessary to first obtain information about the top application opened on the mobile terminal, and then determine the application type of the current top application, and instruct the text recognition module to call different text recognition methods according to different application types.

[0087] FIG9 is a logic diagram of a text recognition module according to an embodiment of the present disclosure. As shown in FIG9 , the following steps are included:

[0088] Step S902: determine whether it is a browser application.

[0089] Step S904: If it is a browser application, obtain the URL information of the browser application.

[0090] Step S906: If it is not a browser application, accessibility is used to obtain text information.

[0091] Step S908: In the case of a browser application, the obtained URL information is uploaded to the server.

[0092] Step S910: Parse the URL information in the server to obtain text information.

[0093] Step S912: Send the text information obtained by the server back to the mobile phone.

[0094] Step S914: determine whether the text information is effective and optimize the text information.

[0095] In actual implementation, to avoid errors or grammatically incorrect content in text messages, the text messages can be checked and optimized based on normal word order, etc., wherein the checking and optimization can be performed using conventional text checking and optimization techniques. If the checking result is invalid, the text message can be directly terminated without being read aloud.

[0096] Step S916: calling the announcement module to announce the text.

[0097] In summary, the method and device for broadcasting screen text information provided by the embodiments of the present disclosure can effectively solve the problem of missing browser application broadcasts currently encountered by various manufacturers. It can provide users with a complete broadcast in any occasion, achieve true full-factor broadcast, and is suitable for the application of aging-friendly products, greatly improving the user experience.

[0098] The disclosed embodiments are intended to improve the convenience for elderly people to obtain screen information when using mobile phones, especially when reading news and other articles, they do not need to keep looking at the phone, and can obtain information about the article through the screen reading function. The disclosed embodiments use different solutions for text recognition in three scenarios. Local applications and social applications obtain text through accessibility. Browser applications need to send it to the server side through the network and parse it on the server side. The former does not require network conditions, but the latter requires a network environment. The disclosed embodiments implement a full-factor screen reading function, which is mainly used to cooperate with the terminal aging-friendly strategy. The technical means of the disclosed embodiments can well solve the problem of incomplete broadcasting for browser scenarios in the current screen reading, thereby increasing the competitiveness of the terminal's aging-friendly products.

[0099] The above description is merely a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will appreciate that various modifications and variations of the present disclosure are possible. Any modifications, equivalent substitutions, or improvements made within the principles of the present disclosure should be included within the scope of protection of the present disclosure.

Claims

1. A method for broadcasting screen text information, comprising: Determine the application type of the application currently opened on the screen of the terminal; Calling corresponding text acquisition methods according to different application types to acquire text information in a current page of the application, wherein the current page includes a visible area of ​​the screen and an invisible area of ​​the screen that can only be displayed according to page turning and pull-down operations; The acquired text information is broadcasted.

2. The method according to claim 1, wherein: The calling of corresponding text acquisition methods according to different application types to acquire text information in the current page of the application includes: In the case where the application type is a browser application, the terminal reports the uniform resource locator URL information of the browser application to the server; The terminal receives text information obtained by the server by parsing the URL information.

3. The method according to claim 1, wherein: The method of calling corresponding text acquisition methods according to different application types to acquire text information in the current page of the application also includes: In the case where the application type is a local application or a social application, the terminal obtains the text information in an accessibility manner.

4. The method according to claim 2, wherein: The server uses an anti-crawler mechanism to parse the URL information.

5. The method according to claim 2, wherein: The server configures the authentication environment for the authenticated website through the back-end browser.

6. The method according to claim 2, wherein: The server uses a multi-threaded concurrent mechanism to parse the URL information.

7. A device for broadcasting screen text information, comprising: Scene recognition module, text recognition module, voice broadcast module, among which, The scene recognition module is configured to determine the application type of the application currently opened on the screen of the terminal; The text recognition module is configured to call corresponding text acquisition methods according to different application types to obtain text information in the current page of the application, wherein the current page includes a visible area of ​​the screen and an invisible area of ​​the screen that can only be displayed according to page turning and pull-down operations; The voice broadcast module is configured to broadcast the acquired text information.

8. The device according to claim 7, wherein: Also includes: A server independently provided from the terminal, wherein: The server is configured to receive the The browser application obtains the uniform resource locator URL information and parses the URL information to obtain the text information.

9. A computer-readable storage medium having a computer program stored therein, wherein: When the computer program is executed by a processor, the method described in any one of claims 1 to 6 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Text reading method and device, storage medium and electronic equipment

    CN116088976A

  • Screen text information broadcasting method and device, storage medium and electronic device

    CN117519554A

  • Browser system, voice proxy server, link item reading-aloud method, and storage medium storing link item reading-aloud program

    JP1999110186A

  • A method to access web page text information that is difficult to read.

    WO2002069322A1