Voice broadcasting method and device

By monitoring the display events of the user interaction page, using the browser interface and voice broadcast module to generate voice broadcast content, the problem of users not paying attention to page promptly is solved, and real-time voice broadcast and resource optimization are achieved.

CN120491923APending Publication Date: 2025-08-15BEIJING BAIJU YIXING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510531906.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the prompt information of the user interaction page is mainly displayed in text, icons or animations, and users may not be able to pay attention to these prompts in time.

Method used

Provide a voice broadcast method, by monitoring the display events of the user interaction page, using the browser interface and the configured voice broadcast module, to generate and broadcast voice content, ensuring that the user can obtain information in a timely manner even if the user does not actively watch the page.

Benefits of technology

Real-time voice broadcasting is realized, avoiding user missed prompts, improving the timeliness and efficiency of information acquisition, reducing auditory fatigue, and optimizing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491923A_ABST
    Figure CN120491923A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses a voice broadcasting method and device. The method comprises the steps that when a display event of a user interaction page is triggered, text content is determined according to the display event; the text content is sent to a browser through an interface of the browser, so that the browser carries out voice synthesis processing on the text content, and voice broadcast content is generated; and broadcasting the voice broadcast content by using a voice broadcast module configured by the browser according to the preset voice parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a voice broadcasting method and device. Background Art

[0002] Currently, prompt information in user interaction pages is mainly presented in the form of text, icons or animations to serve as a reminder to users.

[0003] Although these methods can serve as prompts to users in most cases, the display in the form of text, icons or animations requires users to actively watch, and in some scenarios users may not notice these prompts in time. Summary of the Invention

[0004] In view of this, the present invention provides a voice broadcasting method and device.

[0005] In a first aspect, the present invention provides a voice broadcast method, which includes: when a display event of a user interaction page is triggered, determining the text content according to the display event; sending the text content to the browser through the browser interface, so that the browser performs voice synthesis processing on the text content to generate voice broadcast content; and using the voice broadcast module configured in the browser to broadcast the voice broadcast content according to preset voice parameters.

[0006] The voice announcement method provided in this embodiment, when a display event is triggered on a user-interactive page, retrieves the text content, sends it to the browser via a browser interface for speech synthesis processing, and generates a voice announcement. Finally, the browser's configured voice announcement module announces the information according to preset voice parameters. This means that even if the user isn't actively viewing the text, icons, or animated prompts on the page, they can still obtain relevant information in a timely manner through voice announcements, avoiding the possibility of missing the prompt due to inactivity.

[0007] In one possible implementation, the user interaction page is a front-end page; and when a display event of the user interaction page is triggered, the text content is determined based on the display event, including: when a new node is added to the DOM tree of the front-end page and a page life cycle event is triggered, element information that meets preset conditions is determined from the node; and text content corresponding to the element information is determined from the element information.

[0008] The voice broadcast method provided in this embodiment can capture nodes in real time by monitoring new nodes in the DOM tree, without relying on fixed event binding or polling mechanisms, thereby improving detection efficiency. In addition, by using preset conditions (such as element tags, class names, attributes, etc.), target elements can be accurately screened to avoid processing irrelevant content.

[0009] In one possible implementation, the user interaction page is a mini-program; and when a display event of the user interaction page is triggered, the text content is determined based on the display event, including: listening to the display event of the pop-up component through the life cycle hook of the mini-program; when the display event of the pop-up component is triggered, the text content is determined based on the display event.

[0010] The voice broadcast method provided in this embodiment monitors display events through lifecycle hooks, which can ensure that text is extracted immediately after the pop-up content rendering is completed, avoiding data loss due to asynchronous loading.

[0011] In one possible implementation, a voice broadcast module configured by a browser is used to broadcast the voice broadcast content according to preset voice parameters, including: detecting whether there is historical broadcast content identical to the voice broadcast content in a hash table; if there is no historical broadcast content identical to the voice broadcast content in the hash table, executing the step of using the voice broadcast module configured by the browser to broadcast the voice broadcast content according to the preset voice parameters; if there is historical broadcast content identical to the voice broadcast content in the hash table, detecting the playback time of the historical broadcast content; when the difference between the playback time and the current time is greater than a difference threshold, executing the step of using the voice broadcast module configured by the browser to broadcast the voice broadcast content according to the preset voice parameters.

[0012] The voice broadcast method provided in this embodiment uses a hash table to record historical broadcast content. If the hash table contains historical broadcast content that is identical to the voice broadcast content, the playback time of the historical broadcast content is detected. When the difference between the playback time and the current time is greater than a difference threshold, the voice broadcast module configured in the browser is executed to broadcast the voice broadcast content according to preset voice parameters, thereby avoiding repeated broadcasting of the same information in a short period of time and reducing the user's auditory fatigue.

[0013] In one possible implementation, the voice broadcast module configured by the browser is used to broadcast the voice broadcast content according to preset voice parameters, including: detecting the priority of the voice broadcast content; when the priority of the voice broadcast content is higher than the priority of all the content to be voice broadcast in the voice queue and there is target voice broadcast content, stopping broadcasting the target voice broadcast content, and executing the step of broadcasting the voice broadcast content according to the preset voice parameters using the voice broadcast module configured by the browser.

[0014] The voice broadcast method provided in this embodiment immediately stops the current broadcast and switches to the new content when the new voice content has a higher priority than all other content in the queue, ensuring that urgent information is not delayed. Furthermore, by determining the priority, the voice channel is prevented from being occupied for a long time by low-priority content, thereby improving resource utilization.

[0015] In a second aspect, the present invention provides a voice broadcast device, which includes: a determination module, which is used to determine the text content according to the display event when the display event of the user interaction page is triggered; a generation module, which is used to send the text content to the browser through the browser interface, so that the browser performs speech synthesis processing on the text content and generates voice broadcast content; and a broadcast module, which is used to use the voice broadcast module configured by the browser to broadcast the voice broadcast content according to preset voice parameters.

[0016] In one possible implementation, the user interaction page is a front-end page; and the determination module includes: a first determination unit, which is used to determine element information that meets preset conditions from the node when a new node is added to the DOM tree of the front-end page and a page life cycle event is triggered; and a second determination unit, which is used to determine the text content corresponding to the element information from the element information.

[0017] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the voice broadcast method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0018] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the voice broadcast method of the first aspect or any corresponding embodiment thereof.

[0019] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, which are used to enable a computer to execute the voice broadcast method of the above-mentioned first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 is a flow chart of a voice broadcasting method according to an embodiment of the present invention;

[0022] Figure 2 is a structural block diagram of a voice broadcasting device according to an embodiment of the present invention;

[0023] Figure 3 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0025] According to an embodiment of the present invention, an embodiment of a voice broadcast method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0026] In this embodiment, a voice broadcast method is provided, which can be used in computer devices such as computers, servers, etc. Figure 1 : is a flow chart of a voice broadcast method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0027] Step S101: When a display event of a user interaction page is triggered, text content is determined according to the display event.

[0028] The user interaction page can be the interface through which users interact with the system. The user interaction page can be a webpage, web application, or mobile interface, etc., without specific limitation here. A display event can indicate a state change event of a page or element from hidden to visible (such as DOMContentLoaded, IntersectionObserver triggering, etc.).

[0029] The text content may indicate the text information that needs to be converted into voice broadcast (such as title, description, prompt information, etc.).

[0030] In specific implementations, you can capture display events on user-interactive pages (such as page load completion, element entry into the viewport, user action triggering, etc.). Determine the source of the text content to be broadcast (such as HTML elements, dynamic data, user input, etc.). When a display event on a user-interactive page is triggered, you can extract the corresponding text content from the page. For example, if a user opens a news page, the news headline will be automatically broadcast after the page loads.

[0031] As an example, you can use JavaScript to query specific elements in a page and extract text content.

[0032] Step S102: sending the text content to the browser through the browser interface, so that the browser performs speech synthesis processing on the text content to generate speech broadcast content.

[0033] The browser interface may refer to a standard API provided by the browser (e.g., the Web Speech API). The speech synthesis process may refer to the process of converting text content into speech data. The voice broadcast content may refer to the speech data generated by the browser (which can be directly played). In a specific implementation, the Web Speech API (e.g., the SpeechSynthesis interface) may be used to send text content to the browser. The browser then converts the text content into speech data.

[0034] In one possible implementation, voice resources can be preloaded: call speechSynthesis.getVoices() during application initialization to preload available voice libraries. Generate offline audio files for high-frequency prompts (such as "Submission successful") to reduce real-time synthesis overhead.

[0035] Step S103: Utilize the voice broadcast module configured in the browser to broadcast the voice broadcast content according to preset voice parameters.

[0036] The voice broadcast module can instruct the browser's built-in voice playback function (such as SpeechSynthesis). The preset voice parameters can be pre-configured voice playback parameters (such as speech rate and volume). For example, playing news at a slower speed, or playing songs at a faster speed, etc., are not specifically limited here.

[0037] During specific implementation, the parameters of the voice broadcast (such as speech speed, volume, pitch, etc.) can be pre-set, and the voice broadcast module of the browser is called to play the generated voice content.

[0038] The voice announcement method provided in this embodiment, when a display event is triggered on a user-interactive page, retrieves the text content, sends it to the browser via a browser interface for speech synthesis processing, and generates a voice announcement. Finally, the browser's configured voice announcement module announces the information according to preset voice parameters. This means that even if the user isn't actively viewing the text, icons, or animated prompts on the page, they can still obtain relevant information in a timely manner through voice announcements, avoiding the possibility of missing the prompt due to inactivity.

[0039] In one possible implementation, the user interaction page is a front-end page; and the above step S101 includes:

[0040] Step S1011: When a new node is added to the DOM tree of the front-end page and a page lifecycle event is triggered, element information that meets preset conditions is determined from the node.

[0041] The DMO tree represents the tree structure of the Document Object Model (DOM), which represents the hierarchical structure of an HTML document. Each HTML tag corresponds to a node, and nodes are connected through parent-child relationships. Page lifecycle events are triggered when the browser loads, parses, and renders the page, such as DOMContentLoaded (DOM loading complete) and load (all resources loaded).

[0042] The front-end page can be a web page that users directly access. It can be composed of HTML, CSS and JavaScript, and is responsible for displaying content and interactive logic.

[0043] The preset conditions can satisfy the conditions of the content to be played, wherein the preset conditions can include the label name, class name, attribute value, etc. of the node to be played.

[0044] In practice, you can use MutationObserver to monitor changes in the DOM tree and extract the text content when a toast or popup element appears. You can also monitor page lifecycle events (such as onLoad and onShow) and obtain the text properties of a popup component when it is created or displayed.

[0045] Step S1012: determining text content corresponding to the element information from the element information.

[0046] Create a SpeechSynthesisUtterance instance and set its text property to the captured content. When listening for new node events in the DOM tree, traverse the newly added nodes and filter the target elements based on preset conditions (such as tag name, class name, and attributes). Extract the target element's attributes or structural information (such as ID, class name, and child nodes). Then call the SpeechSynthesisUtterance instance's speak method to start the speech broadcast.

[0047] In one possible implementation, the presence of window.speechSynthesis is detected. If not, the server-side speech synthesis interface (such as Baidu TTS API) is downgraded. The voice library is automatically matched based on navigator.language, supporting dynamic switching between Chinese and English.

[0048] The voice broadcast method provided in this embodiment can capture nodes in real time by monitoring new nodes in the DOM tree, without relying on fixed event binding or polling mechanisms, thereby improving detection efficiency. In addition, by using preset conditions (such as element tags, class names, attributes, etc.), target elements can be accurately screened to avoid processing irrelevant content.

[0049] In one possible implementation, the user interaction page is a mini-program; and the above step S101 includes:

[0050] Step S1013: Monitor the display event of the pop-up component through the life cycle hook of the mini program.

[0051] Monitor the display events of pop-up components through the mini program lifecycle hooks (such as onShow and onReady).

[0052] In one possible implementation, embedding the Web Speech API in the Mini Program WebView component requires verifying the support of the basic library version (such as WeChat kernel version ≥ 8.0.16).

[0053] Step S1014: when a display event of the pop-up component is triggered, the text content is determined according to the display event.

[0054] When the display event of the pop-up component is triggered, you can call wx.createSelectorQuery() to obtain the pop-up content text.

[0055] In one possible implementation, wx.createInnerAudioContext is used to play pre-generated audio files, and the server needs to cooperate in generating audio (MP3 / WAV format).

[0056] In one possible implementation, the applet pop-up window text is passed to the Web Speech API in the WebView for processing via postMessage communication.

[0057] The voice broadcast method provided in this embodiment monitors display events through lifecycle hooks, which can ensure that text is extracted immediately after the pop-up content rendering is completed, avoiding data loss due to asynchronous loading.

[0058] In one possible implementation, step S103 includes:

[0059] Step S1031 , detecting whether there is a historical broadcast content identical to the voice broadcast content in the hash table.

[0060] The hash table can be a table that stores historical broadcast data. Each broadcast data played can be stored in the hash table. When the browser performs speech synthesis processing on the text content to generate speech broadcast content, it can further check whether there is historical broadcast content corresponding to the broadcast content in the hash table.

[0061] Step S1032: If the hash table does not contain the same historical broadcast content as the voice broadcast content, the step of using the voice broadcast module configured by the browser to broadcast the voice broadcast content according to the preset voice parameters is executed.

[0062] If the hash table does not contain any historical broadcast content that is identical to the voice announcement, the voice announcement is new. The browser's voice announcement module can then be used to announce the announcement using pre-set voice parameters. For example, if the announcement is "Turn on the lights," and the hash table does not contain any historical broadcast content for "Turn on the lights," the browser's voice announcement module can be used to announce the announcement using pre-set voice parameters.

[0063] In one possible implementation, the hash table can be compared to determine whether there is historical content identical to the voice announcement. Specifically, a neural network model can be used to determine similarity based on the voice announcement. If the similarity reaches 95%, the historical content and the voice announcement are identical.

[0064] Step S1033: If the hash table contains the same historical broadcast content as the voice broadcast content, the playback time of the historical broadcast content is detected.

[0065] Step S1034: When the difference between the play time and the current time is greater than the difference threshold, the step of using the voice broadcast module configured by the browser to broadcast the voice broadcast content according to the preset voice parameters is executed.

[0066] The difference threshold can indicate the critical value for broadcasting the voice broadcast content. When the difference exceeds the difference threshold, the voice broadcast content is broadcast. If the hash table contains historical broadcast content identical to the voice broadcast content, the broadcast time of the historical broadcast content needs to be determined. When the difference between the playback time and the current time exceeds the difference threshold, the voice broadcast content is broadcast using the browser-configured voice broadcast module according to preset voice parameters.

[0067] In one possible implementation, when the difference between the playback time and the current time is not greater than a difference threshold, the voice announcement module configured by the browser is not executed to announce the voice announcement content according to preset voice parameters. Instead, the voice announcement content can be placed in a queue and configured with a timestamp. When the next voice announcement content is the same as the voice announcement content with the configured timestamp, the voice announcement content in the queue can be directly detected.

[0068] The voice broadcast method provided in this embodiment uses a hash table to record historical broadcast content. If the hash table contains historical broadcast content that is identical to the voice broadcast content, the playback time of the historical broadcast content is detected. When the difference between the playback time and the current time is greater than a difference threshold, the voice broadcast module configured in the browser is executed to broadcast the voice broadcast content according to preset voice parameters, thereby avoiding repeated broadcasting of the same information in a short period of time and reducing the user's auditory fatigue.

[0069] In one possible implementation, step S103 includes:

[0070] Step S1035: Detect the priority of the voice broadcast content.

[0071] The priority of the voice broadcast content may be a pre-set priority. When the browser performs speech synthesis processing on the text content to generate the voice broadcast content, the priority of the voice broadcast content may be detected.

[0072] Step S1036: When the priority of the voice broadcast content is higher than the priority of all the contents to be voice broadcast in the voice queue and there is target voice broadcast content, stop broadcasting the target voice broadcast content and execute the step of broadcasting the voice broadcast content according to the preset voice parameters using the voice broadcast module configured by the browser.

[0073] If the priority of the voice broadcast content is higher than the priority of all the content to be voice broadcast in the voice queue and there is a target voice broadcast content, then the voice broadcast content needs to be broadcast first, then stop broadcasting the target voice broadcast content, and execute the step of broadcasting the voice broadcast content according to the preset voice parameters using the voice broadcast module configured by the browser.

[0074] As an example, an online customer service system can support voice announcements. User-submitted inquiries are announced according to their priority level, which is dynamically set by the system based on the message type (e.g., urgent messages have a higher priority, ordinary messages have a lower priority). The system is currently announcing a normal message: "Hello, your order has been accepted." At this moment, the system receives a high-priority message: "Urgent Notice: System Maintenance is Imminent. Please Save Your Data Immediately!" The system detects that the priority of the new message, M, is higher than the priority of all messages in the current queue. The system immediately interrupts the current announcement and begins announcing the high-priority message.

[0075] The voice broadcast method provided in this embodiment immediately stops the current broadcast and switches to the new content when the new voice content has a higher priority than all other content in the queue, ensuring that urgent information is not delayed. Furthermore, by determining the priority, the voice channel is prevented from being occupied for a long time by low-priority content, thereby improving resource utilization.

[0076] In this embodiment, a voice broadcast device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments. The details already described will not be repeated here. As used below, the term "module" can mean a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0077] This embodiment provides a voice broadcast device, such as Figure 2 As shown, it includes: a determination module 201, which is used to determine the text content according to the display event when the display event of the user interaction page is triggered; a generation module 202, which is used to send the text content to the browser through the browser interface, so that the browser performs speech synthesis processing on the text content and generates voice broadcast content; a broadcast module 203, which is used to use the voice broadcast module configured by the browser to broadcast the voice broadcast content according to preset voice parameters.

[0078] In one possible implementation, the determination module 201 includes: a first determination unit, used to determine element information that meets preset conditions from the node when a new node is added to the DOM tree of the front-end page and a page life cycle event is triggered; and a second determination unit, used to determine the text content corresponding to the element information from the element information.

[0079] In one possible implementation, the determination module 201 includes: a monitoring unit for monitoring the display event of the pop-up component through the life cycle hook of the mini program; and a third determination unit for determining the text content according to the display event when the display event of the pop-up component is triggered.

[0080] In one possible implementation, the broadcast module 203 includes: a first detection unit, used to detect whether there is historical broadcast content identical to the voice broadcast content in the hash table; a first execution unit, used to execute the step of broadcasting the voice broadcast content according to preset voice parameters using the voice broadcast module configured by the browser if there is no historical broadcast content identical to the voice broadcast content in the hash table; a second detection unit, used to detect the playback time of the historical broadcast content if there is historical broadcast content identical to the voice broadcast content in the hash table; a second execution unit, used to execute the step of broadcasting the voice broadcast content according to preset voice parameters using the voice broadcast module configured by the browser when the difference between the playback time and the current time is greater than a difference threshold.

[0081] In one possible implementation, the broadcast module 203 includes: a third detection unit, used to detect the priority of the voice broadcast content, and a third execution unit, used to stop broadcasting the target voice broadcast content when the priority of the voice broadcast content is higher than the priority of all the content to be voice broadcast in the voice queue and there is target voice broadcast content, and execute the step of broadcasting the voice broadcast content according to preset voice parameters using the voice broadcast module configured by the browser.

[0082] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0083] The voice broadcast device in this embodiment is presented in the form of a functional unit, where the functional unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0084] The embodiment of the present invention also provides a computer device having the above Figure 2 The voice broadcast device shown.

[0085] See also Figure 3 , Figure 3 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 3 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 3 A processor 10 is taken as an example.

[0086] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0087] The memory 20 stores instructions that can be executed by at least one processor 10, so as to enable at least one processor 10 to execute the method shown in the above embodiment.

[0088] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0089] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0090] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 3 The bus connection is taken as an example.

[0091] The input device 30 can receive input digital or character information and generate key signal input related to user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0092] The computer device further includes a communication interface for the computer device to communicate with other devices or a communication network.

[0093] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0094] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0095] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A voice broadcasting method, characterized in that: The method comprises: When a display event of a user interaction page is triggered, determining text content according to the display event; Sending the text content to the browser through the browser interface, so that the browser performs speech synthesis processing on the text content to generate speech broadcast content; The voice broadcast module configured by the browser is used to broadcast the voice broadcast content according to preset voice parameters.

2. The voice broadcast method according to claim 1, characterized in that: The user interaction page is a front-end page; and when a display event of the user interaction page is triggered, determining text content according to the display event, including: When a new node is added to the DOM tree of the front-end page and a page lifecycle event is triggered, element information that meets the preset conditions is determined from the node; Determine text content corresponding to the element information from the element information.

3. The voice broadcast method according to claim 1, wherein: The user interaction page is a mini-program; and when a display event of the user interaction page is triggered, determining text content according to the display event, including: Monitor the display events of the pop-up component through the lifecycle hook of the mini program; When a display event of the pop-up component is triggered, text content is determined according to the display event.

4. The voice broadcast method according to claim 1, wherein: The voice broadcast module configured by the browser broadcasts the voice broadcast content according to preset voice parameters, including: Checking whether there is a historical broadcast content identical to the voice broadcast content in the hash table; If the hash table does not contain the same historical broadcast content as the voice broadcast content, executing the step of using the voice broadcast module configured by the browser to broadcast the voice broadcast content according to the preset voice parameters; If the hash table contains the same historical broadcast content as the voice broadcast content, detecting the playback time of the historical broadcast content; When the difference between the play time and the current time is greater than a difference threshold, the step of using a voice broadcast module configured by the browser to broadcast the voice broadcast content according to preset voice parameters is executed.

5. The voice broadcasting method according to claim 1, characterized in that: The voice broadcast module configured by the browser broadcasts the voice broadcast content according to preset voice parameters, including: Detecting the priority of the voice broadcast content; When the priority of the voice broadcast content is higher than the priority of all the to-be-voice-broadcast contents in the voice queue and there is target voice broadcast content, stop broadcasting the target voice broadcast content and execute the step of broadcasting the voice broadcast content according to preset voice parameters using the voice broadcast module configured by the browser.

6. A voice broadcasting device, characterized in that: The device comprises: A determination module, configured to determine text content according to a display event of a user interaction page when the display event is triggered; A generating module, configured to send the text content to the browser via the browser interface, so that the browser performs speech synthesis processing on the text content to generate speech broadcast content; The broadcast module is used to broadcast the voice broadcast content according to preset voice parameters using the voice broadcast module configured by the browser.

7. The voice broadcasting device according to claim 6, characterized in that: The user interaction page is a front-end page; and the determination module includes: A first determining unit is configured to determine element information that meets preset conditions from a new node in the DOM tree of the front-end page when a page lifecycle event is triggered; The second determining unit is configured to determine text content corresponding to the element information from the element information.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the voice broadcast method according to any one of claims 1 to 5 by executing the computer instructions.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the voice broadcast method according to any one of claims 1 to 5.

10. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to enable a computer to execute the voice broadcasting method according to any one of claims 1 to 5.