Voice broadcast method, device, terminal, storage medium and program product

CN115641833BActive Publication Date: 2026-09-25GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211285185.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-20
Publication Date
2026-09-25
Estimated Expiration
2042-10-20

AI Technical Summary

Technical Problem

[0004]然而,直接将界面文本进行语义提取后进行语音播报,无法准确表达出该结构化图形用户界面的关键信息,进而通过语音播报向用户提供的图形用户界面的界面信息与视觉交互的差距较大,智能程度较低

Benefits of technology

[0019]在本申请实施例中,终端可以在用户不方便进行视觉交互的情况下,获取结构化图形用户界面的文本语义特征以及图像语义特征,并根据文本语义特征以及图像语义特征生成交互文本后对用户进行语音播报。其中,结合图像语义特征能够使得最终生成的交互文本更加贴合视觉交互中用户获取到的信息。另一方面根据文本语义特征和图像语义特征生成交互文本,能够将结构化信息中离散、割裂的文字内容进行整合,使得交互文本的内容更加流畅、自然。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115641833B_ABST
    Figure CN115641833B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a voice broadcast method and device, a terminal, a storage medium and a program product, and belong to the field of human-computer interaction. The method comprises: acquiring an interface image of a structured graphical user interface, the structured graphical user interface being used for structured display of content; acquiring text semantic features corresponding to each text block in the interface image and image semantic features, the text block being obtained by division based on the distribution of text in the interface image; generating interactive text corresponding to the structured graphical user interface based on the text semantic features and the image semantic features, the interactive text being used for describing the content contained in the structured graphical user interface; and performing voice broadcast on the interactive text. The scheme provided by the embodiments of the present application can perform voice broadcast on discrete information in the structured graphical user interface, so that the user can obtain the same information as visual interaction through voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction technology, and in particular to a voice broadcasting method, device, terminal, storage medium, and program product. Background Technology

[0002] With the development of software technology, structured graphical user interfaces have gradually become the mainstream end-user interaction interface. At the same time, voice, as a human-computer interaction method, has also been widely used, and users' demand for obtaining terminal user interface information through voice is constantly increasing.

[0003] In related technologies, text in a structured graphical user interface is acquired, its semantic features are extracted, and then TTS (Text To Speech) technology is used to convert the semantic features into speech, which is then used to convey graphical user interface information to the user through voice broadcast.

[0004] However, directly extracting the semantics of the interface text and then using voice playback cannot accurately express the key information of the structured graphical user interface. Consequently, the interface information provided to the user through voice playback differs significantly from the visual interaction, resulting in a low level of intelligence. Summary of the Invention

[0005] This application provides a voice broadcasting method, apparatus, terminal, storage medium, and program product. The technical solution is as follows:

[0006] On one hand, embodiments of this application provide a voice broadcasting method, the method comprising:

[0007] Obtain the interface image of the structured graphical user interface, which is used to display content in a structured manner;

[0008] The text semantic features and image semantic features corresponding to each text block in the interface image are obtained, and the text blocks are divided based on the distribution of text in the interface image;

[0009] Based on the text semantic features and the image semantic features, interactive text corresponding to the structured graphical user interface is generated, and the interactive text is used to describe the content contained in the structured graphical user interface;

[0010] The interactive text is read aloud via voice.

[0011] On the other hand, embodiments of this application provide a voice broadcasting device, the device comprising:

[0012] The interface image acquisition module is used to acquire the interface image of the structured graphical user interface, which is used to display content in a structured manner.

[0013] The semantic feature acquisition module is used to acquire the text semantic features and image semantic features corresponding to each text block in the interface image, wherein the text blocks are divided based on the distribution of text in the interface image;

[0014] An interactive text generation module is used to generate interactive text corresponding to the structured graphical user interface based on the text semantic features and the image semantic features. The interactive text is used to describe the content contained in the structured graphical user interface.

[0015] The voice broadcast module broadcasts the interactive text via voice.

[0016] On the other hand, embodiments of this application provide a terminal, the terminal including a processor and a memory; the memory stores at least one instruction, the at least one instruction being executed by the processor to implement the voice broadcasting method as described above.

[0017] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the voice broadcasting method as described above.

[0018] On the other hand, embodiments of this application provide a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice broadcasting method provided in various optional implementations of the above aspects.

[0019] In this embodiment, the terminal can acquire textual and image semantic features of a structured graphical user interface when visual interaction is inconvenient for the user. Based on these features, it generates interactive text and then delivers it via voice. Combining image semantic features makes the final generated interactive text more closely aligned with the information obtained by the user during visual interaction. Furthermore, generating interactive text based on both textual and image semantic features integrates discrete and fragmented textual content within structured information, resulting in a smoother and more natural interactive text. Attached Figure Description

[0020] Figure 1 A flowchart of a voice broadcasting method provided in an exemplary embodiment of this application is shown;

[0021] Figure 2 A schematic diagram of a structured graphical user interface provided in an exemplary embodiment of this application is shown;

[0022] Figure 3 A flowchart illustrating the semantic feature extraction process provided in an exemplary embodiment of this application is shown;

[0023] Figure 4 A schematic diagram of a structured graphical user interface provided in another exemplary embodiment of this application is shown;

[0024] Figure 5 A flowchart illustrating the interactive text generation process provided in an exemplary embodiment of this application is shown;

[0025] Figure 6 A schematic diagram of a structured graphical user interface provided in another exemplary embodiment of this application is shown;

[0026] Figure 7 A schematic diagram of a voice interaction process provided in an exemplary embodiment of this application is shown;

[0027] Figure 8 This paper shows a structural block diagram of a voice broadcasting device according to an embodiment of the present application;

[0028] Figure 9 A structural block diagram of a terminal provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0030] It should be noted that all embodiments of this application are executed on the terminal after the voice broadcast program is triggered, and this application does not limit the method of triggering the voice broadcast program. Furthermore, all structured graphical user interface data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with relevant national and regional laws, regulations, and standards.

[0031] Figure 1 A flowchart illustrating a voice broadcasting method provided in an illustrative embodiment of this application is shown. The method includes:

[0032] Step 101: Obtain the interface image of the structured graphical user interface, which is used to display content in a structured manner.

[0033] A structured graphical user interface (GUI) is a way to display application content or system messages in a structured manner. It can be provided by the application itself or automatically generated by the operating system. Common structured GUIs include desktop cards and application interfaces.

[0034] The terminal can capture the interface image of the structured graphical user interface by taking a screenshot, and then send the captured image to the terminal's built-in processor to process the image and adjust its pixel size for subsequent steps.

[0035] Step 102: Obtain the text semantic features and image semantic features corresponding to each text block in the interface image. The text blocks are divided based on the distribution of text in the interface image.

[0036] The features of a text block include textual semantic features and image semantic features. Textual semantic features are mainly related to the text content and position of the text block, while image semantic features refer to the text form features such as the size, color, and font of the text in the text block.

[0037] After the terminal obtains the interface image of the structured graphical user interface, it needs to acquire the text semantic features and image semantic features of the interface image.

[0038] In a structured graphical user interface, text is the most direct content that can reflect key information. Therefore, before obtaining the text semantic features and image semantic features of the interface image, it is necessary to divide the text into multiple text blocks according to the distribution of text in the interface image, and then obtain the features of each text block.

[0039] In structured graphical user interfaces (GUIs), text conveying key information is typically given special fonts, colors, or sizes to distinguish it from other text providing less important information. This allows users to obtain key information more directly and clearly during visual interactions. Similarly, in voice interactions, interactive text generated by combining image semantic features can contain more key information, enabling users to obtain the same information through voice prompts as they would through visual interactions. Therefore, it is necessary to acquire the image semantic features of the interface images.

[0040] In one possible implementation, the terminal extracts features from the text block image using an image feature extraction network. Since the size of the divided text blocks varies, and image feature extraction networks typically have certain requirements regarding the size of the input image, the size of the text block image needs to be adjusted before the terminal acquires the semantic features of the image.

[0041] Then, the input text block image is processed by an image feature extraction network. Different image feature extraction networks have various algorithms, including scale-invariant feature transformation, histogram of oriented gradients, and neural network feature extraction. Different image feature extraction networks with different performance can be selected according to different terminals and different human-computer interaction scenarios. This application does not limit the type of image feature extraction network algorithm.

[0042] Optionally, a bottleneck network is used to obtain the semantic features of the text block image. The resized text block image is input into the bottleneck network, and then passed through a 1*1 convolutional module, a DW (Deep-Wise) module, and an SE (Squeeze and Excite) convolutional module in sequence. After that, it passes through a 1*1 convolutional module and is then added to the original input tensor, so that the output tensor still carries the original features of the input tensor.

[0043] Step 103: Generate interactive text corresponding to the structured graphical user interface based on text semantic features and image semantic features.

[0044] Interactive text is used to describe key information contained in the structured graphical user interface.

[0045] After the terminal acquires the semantic features of the text and the image, it fuses the two to obtain the semantic expression that represents each text block. After further fusion, the semantic features of the graphical user interface are generated. Subsequently, the text representation of this structured graphical user interface is generated through text generation and text summarization, which can then be used as the interactive text.

[0046] One approach is to use text generation techniques that convert numerical values ​​into text to process the semantic features of the structured graphical interface, then simplify them using text summarization algorithms, and finally generate interactive text.

[0047] Step 104: Read the interactive text aloud via voice.

[0048] After the terminal generates interactive text, it will read the content of the interactive text aloud. In one possible implementation, the terminal uses Text-to-Speech (TTS) technology to convert the interactive text into a speech signal. TTS mainly consists of two parts: speech analysis and an acoustic system. Speech analysis mainly analyzes the text information of the input interactive text, determines the text structure and language, standardizes the text, converts the text into phonemes, and finally predicts sentence and prosody. The acoustic system part generates corresponding audio based on the results of speech analysis, thus realizing the function of converting interactive text into a speech signal.

[0049] like Figure 2As shown, for the first graphical user interface 201, after generating interactive text, the terminal broadcasts the following message: "Your China Eastern Airlines flight MU5332 is scheduled to depart from Shenzhen Bao'an T3 at 07:15 on July 17th and is scheduled to arrive at Shanghai Pudong T1 at 09:35. Check-in counters are E01-E18, and the boarding gate is F28." After the user's flight lands, the first graphical user interface 201 will switch to the second graphical user interface 202. For the second graphical user interface 202, the following message will be broadcast: "Your China Eastern Airlines flight MU5332 arrived at Shanghai Pudong T1 at 09:35. The flight time was 2 hours and 20 minutes, the flight distance was 1343 km, and the baggage carousel number is 26."

[0050] In summary, in this embodiment, the terminal can acquire the text semantic features and image semantic features of the structured graphical user interface when the user is unable to perform visual interaction. Based on these features, it generates interactive text and then provides voice prompts to the user. Combining image semantic features makes the final generated interactive text more closely match the information obtained by the user during visual interaction. Furthermore, generating interactive text based on both text and image semantic features integrates the discrete and fragmented text content within the structured information, resulting in a smoother and more natural interactive text.

[0051] On the one hand, the text in a graphical user interface (GUI) block can represent the content presented to the user to a certain extent. On the other hand, the image of the text block can also display the highlighted text in the GUI; its different colors, fonts, and sizes can indicate the importance of the text to a certain extent. Therefore, semantic feature extraction of the interface image of a structured GUI is achieved by extracting both the textual semantic features and the image semantic features of the text blocks as the semantic features of the structured GUI. This ensures that the generated interactive text content is as similar as possible to the information that the visual interaction user can obtain, achieving the same effect as voice broadcasting and visual interaction.

[0052] The following illustrative embodiment will illustrate the process by which a terminal obtains the text semantic features and image semantic features corresponding to each text block in an interface image.

[0053] Figure 3 A flowchart illustrating a semantic feature extraction process provided in an exemplary embodiment of this application is shown. The process includes:

[0054] Step 301: Perform text detection on the interface image using a text detection model to obtain the text location.

[0055] The terminal needs to divide text blocks according to the text distribution in the interface image. Therefore, text detection is required to locate the position of the text contained in the graphical user interface.

[0056] Text detection can be achieved through text detection models, including regression-based and segmentation-based methods. Specifically, methods such as Textbox, EAST, and region segmentation with boundary calibration can be selected. Different text detection models have different structures, resulting in varying detection speeds and different requirements for terminal processing performance. This application does not limit the method used for text detection.

[0057] In one possible implementation, the terminal employs natural image text detection based on a connected text proposal network. First, a small region of fixed pixels is moved and monitored in the interface image. The detector can then output the text-to-non-text ratio score and the predicted y-axis coordinate of K anchor points at each window position. Then, an RNN (Recurrent Neural Network) is used to connect these small regions, and finally, the edges are refined.

[0058] Step 302: Based on the text location, the interface image is divided into text blocks.

[0059] After obtaining the text location, the terminal divides the interface image into multiple text blocks based on the different text locations. Since the content of the text at different locations in a structured graphical user interface differs, the size of the text blocks also varies.

[0060] like Figure 4 As shown, in the structured graphical user interface 401, seven different text blocks are divided according to their positions: the first text block 402, the second text block 403, the third text block 404, the fourth text block 405, the fifth text block 406, the sixth text block 407, the seventh text block 408, and the eighth text block 409. Among these, the second text block 403 and the fourth text block 405 have more text content and larger font sizes, therefore, their areas are also larger. For text that is close together in the structured graphical user interface, the terminal may also divide it into a single text block.

[0061] After the terminal receives the text block, it needs to extract the text information from the text block to obtain the semantic features of the text block. The semantic features of the text block include text semantic features and image semantic features.

[0062] Step 303: Perform optical character recognition on the text block to obtain the text content information of the text block.

[0063] When the terminal extracts semantic features of text, it uses optical character recognition to search, extract and recognize the characters in the text block. It determines the shape by detecting the light and dark patterns, and then uses character recognition methods to translate the shape into computer text.

[0064] In one possible implementation, the terminal can perform character recognition on text blocks using a CRNN (Convolutional Recurrent Neural Network). In the CRNN model, convolutional layers and max-pooling layers from a CNN (Convolutional Neural Networks) model (with fully connected layers removed) are used to construct new convolutional layers to extract semantic features from the input text block. These features are then passed through an RNN recurrent layer to predict the label distribution for each frame of the image. Finally, the label distribution predicted by the RNN is converted into a label sequence. In a dictionary-based mode, the label sequence with the highest probability is selected for prediction. Ultimately, the text content information of the text block is obtained.

[0065] Step 304: Determine the text position information of the text block. The text position information is used to characterize the position of the text block in the interface image.

[0066] After obtaining the text location, the terminal can determine the text location information of the text block based on the absolute position of the text block in the interface image and the relative position of the text block with other text blocks.

[0067] Step 305: Encode the text content information and text position information to obtain the text encoding vector of the text block.

[0068] For text content information, the terminal can encode it using methods such as one-hot encoding, matrix-based distributed representation, clustering-based distributed representation, and neural network-based distributed representation. Optionally, the terminal encodes the text content by querying the position of characters in a dictionary within a text block. This dictionary can be created based on a pre-stored corpus or based on the text in the graphical user interface.

[0069] For text location information, normalization can be used for encoding, ultimately resulting in a numerical encoding vector.

[0070] After the terminal encodes the text content information and the text location information respectively, the encoding results are concatenated to obtain the text encoding vector corresponding to the text block.

[0071] Step 306: Based on the text encoding vector, feature extraction is performed using a text semantic feature extraction model to obtain the text semantic features of the text block.

[0072] After obtaining the text encoding vector from the terminal, features can be extracted using methods such as the Transformer model, constructing an evaluation function, and the skip-gram model to obtain the text semantic features corresponding to the text blocks.

[0073] In one possible implementation, the terminal uses a Transformer model for feature extraction. The Transformer model consists of an encoder framework and a decoder framework. Each encoder and decoder framework is composed of multiple independent feature extractors stacked together. The initial input text encoding vector, after passing through the encoder framework, yields a matrix or vector, which is analogous to an encoding of the text encoding vector. The decoder structure is flexible and can decode the matrix or vector according to different application scenarios. The output result is the text semantic features of the text block.

[0074] Step 307: Extract features from the image information of the text block to obtain the image semantic features of the text block.

[0075] While extracting features from the text content and position information of the text block, the terminal also needs to extract features from the image information of the text block so that the final generated interactive text contains more key information from the structured graphical user interface.

[0076] The terminal can extract features from image information by using methods such as LBP (Local Binary Pattern) and color moments, and finally obtain the semantic features of the image.

[0077] In this embodiment, the terminal identifies the text position in the interface image and divides the text into blocks based on the text position. This allows it to accurately identify the text position in the entire interface image. Then, it identifies the text in the text block using optical character recognition. After encoding the text position information and text content together and extracting features, it obtains text semantic features. This makes the generated interactive text have structured information of the graphical user interface, making it smoother, more coherent, and easier for users to understand. Furthermore, combined with the image semantic features of the text block, the generated interactive text can better reflect the key information presented by the current graphical user interface.

[0078] After acquiring text and image semantic features, the terminal generates semantic text corresponding to the structured graphical user interface (GUI) based on these features. Then, it determines the interactive text for the GUI, which is text with interactive attributes. The terminal uses text and image semantic features to determine whether the text has interactive attributes. Interactive attributes indicate whether the interface element corresponding to the text block can be triggered by the user. For example, in a card containing a text block with the text "View Orders," the color and font of this text are different from other texts, and the color and font size of interactive elements are usually consistent with this text. Therefore, the terminal will likely determine this text to be interactive and have interactive attributes.

[0079] Once the interactive text is determined, the terminal will generate interactive text corresponding to the structured graphical user interface based on the semantic text and the interactive text.

[0080] Among them, semantic text is generated by the terminal based on text semantic features and image semantic features. After extracting and fusing the semantic fusion features to obtain the interface semantic features of the structured graphical user interface, the interface semantic features are then input into the text generation model.

[0081] Specifically, the terminal first extracts and fuses the semantic fusion features corresponding to each text block to obtain the first semantic features of the structured graphical user interface. Then, the first interface semantic features are input into the first text generation model to obtain the semantic text corresponding to the structured graphical user interface. The first text generation model has text generation and text summarization functions.

[0082] The first text generation model can generate interactive text based on the semantic features of the input interface.

[0083] In one possible implementation, the framework of the first text generation model includes a signal analysis module, a data interpretation module, a document planning module, and a microplanning and implementation module. The signal analysis module takes the semantic features of the first interface as input, uses various data analysis methods to detect basic patterns in the semantic features of the first interface, and outputs discrete data patterns. The data interpretation module takes the basic patterns and events in the semantic features of the first interface as input, analyzes the basic patterns and events to infer more complex and abstract messages, and infers the relationships between messages, finally outputting high-level messages and the relationships between messages, such as causal relationships, temporal relationships, etc. The document planning module takes the messages and the relationships between messages as input, analyzes and determines the messages to be mentioned in the interactive text, determines the structure of the interactive text, and finally outputs the messages to be mentioned and the document structure. The microplanning and implementation module takes the simplified messages to be mentioned and the document structure as input, and outputs the final interactive text using natural language generation technology. Furthermore, the first text generation model also includes a text summarization module, used to simplify the messages to be mentioned in the output, using a text summarization algorithm to obtain the simplified messages to be mentioned and the document structure.

[0084] The first text generation model enables the final generated interactive text to have correct syntax and voice, as well as information from a structured graphical user interface.

[0085] The process of generating interactive text on a terminal will be described below through an illustrative embodiment.

[0086] Figure 5 A flowchart illustrating a process for generating interactive text according to an exemplary embodiment of this application is shown, the process including:

[0087] Step 501: The text semantic features and image semantic features corresponding to each text block are fused to generate semantic fusion features corresponding to each text block.

[0088] After obtaining the text semantic features and image semantic features, the terminal will fuse the text semantic features and image semantic features of each text block to obtain the semantic expression of each text block. The fusion can be direct fusion or weighted fusion. The fusion methods that can be used include maximum fusion, MFB (Multimodal Factorized Bilinear Pooling) and artificial neural network methods.

[0089] Step 502: Based on the semantic fusion features corresponding to the text block, classify the text block using the first classification model to obtain the first classification result. The first classification result is used to characterize whether the text block contains key text.

[0090] Since structured graphical user interfaces may contain a large number of text blocks, if all the text content in all the text blocks is read aloud, there may be too much interactive text content. Then, if the text is read aloud, it will waste a lot of time. In addition, in a structured graphical user interface, the text content in some text blocks is not important content of the structured graphical user interface and does not need to convey its information to the user. Concise interactive text is more conducive to the user understanding of the main information of the graphical user interface.

[0091] Therefore, after the terminal obtains the semantic fusion features corresponding to the text blocks, the first classification result divides all text blocks into text blocks containing key text and text blocks not containing key text. This process is implemented using the first classification model.

[0092] In one possible implementation, a first classification model is pre-trained to calculate the probability of key semantic features of the input. The terminal inputs the semantic fusion features corresponding to the text blocks into the first classification model to obtain an output result, which is the first classification result representing the probability of keyness. Text blocks with probabilities higher than a threshold are identified as text blocks containing key text. This first classification model can be trained based on key labels corresponding to semantic features of a large number of samples.

[0093] Step 503: Based on the first classification result, select the target text block from the text block. The target text block contains key text.

[0094] The first classification result is divided into two categories: text blocks containing key text and text blocks not containing key text. The terminal selects the text blocks containing key text as the target text blocks. Each target text block contains key text.

[0095] Step 504: Extract and fuse the semantic fusion features corresponding to the target text block to obtain the second interface semantic features of the structured graphical user interface.

[0096] After the terminal determines the target text block, it extracts and fuses the semantic fusion features corresponding to the target text block. This step can be achieved through a neural network model.

[0097] In one possible implementation, a Transformer model employing a self-attention mechanism extracts and fuses features from the semantic fusion vector corresponding to the target text block. The input to the Transformer model is the semantic fusion features corresponding to each text block. Each encoder network block in the Transformer model contains a multi-head self-attention layer and a feedforward neural network layer. The structure of each decoder is similar to that of the encoder, with each decoder network block containing an additional multi-head self-attention layer. For the purpose of optimizing the deep network, the entire network uses residual connections and performs normalization processing on the network layers. The Transformer model with a self-attention mechanism can obtain the second interface semantic features of a structured graphical user interface.

[0098] Step 505: Input the semantic features of the second interface into the second text generation model to obtain the semantic text corresponding to the structured graphical user interface.

[0099] The second text generation model has text generation and text summarization functions. It can generate interactive text based on the semantic features of the second interface. The architecture of the second text generation model is the same as that of the first text generation model, and will not be described in detail here.

[0100] Step 506: Based on the text semantic features and image semantic features corresponding to the text block, the text block is classified by the second classification model to obtain the second classification result. The second classification result is used to characterize whether the text block contains interactive text.

[0101] The second classification model is used to classify the interactivity of text blocks. If a text block contains interactive text, it is an interactive text block; if a text block does not contain interactive text, it is a non-interactive text block.

[0102] In one possible implementation, a second classification model is pre-trained to calculate the interactivity probability of input semantic features. The terminal inputs the textual semantic features and image semantic features of text blocks into the second classification model to obtain an output result, which represents the interactivity probability. Text blocks with probabilities higher than a threshold are identified as interactive text blocks. This second classification model can be trained based on the interactivity labels corresponding to the semantic features of a large number of samples.

[0103] Step 507: Based on the second classification result, determine the interactive text in the structured graphical user interface.

[0104] The second classification result includes interactive text blocks and non-interactive text blocks, where interactive text blocks contain interactive text.

[0105] Step 508: Generate interactive text corresponding to the structured graphical user interface based on semantic text and interactive text.

[0106] The interactive text is generated by the terminal using a simple template to combine interactive text and semantic text.

[0107] For example, the template could be "...For further action, you can say...", which guides the user to issue voice commands for interaction. It can be illustrative, such as... Figure 2 The interactive text generated by the terminal in the first graphical user interface 201 can be: "Your China Eastern Airlines flight MU5332 is scheduled to depart from Shenzhen Bao'an T3 at 07:15 on July 17th and arrive at Shanghai Pudong T1 at 09:35. Check-in counters are E01-E18, and the boarding gate is F28. For further operations, you can say 'Flight Assistant,' 'View Orders,' or 'Book a Hotel.'" The interactive text in the second graphical user interface can be: "Your China Eastern Airlines flight MU5332 arrived at Shanghai Pudong T1 at 09:35. The flight time is 2 hours and 20 minutes, the flight distance is 1343 km, and the baggage carousel number is 26. For further operations, you can say 'Flight Assistant,' 'View Orders,' or 'Book a Hotel.'"

[0108] In this embodiment, the terminal uses a first classification model to determine whether there is key text in the text block, ensuring that the generated interactive text contains key information in the structured graphical user interface, making the interactive text content more concise and accurate. Furthermore, the terminal uses a second classification model to determine the interactivity of the text, enabling the terminal to provide structured graphical user interface information to the user via voice broadcast while also guiding the user to operate the graphical user interface based on voice, thus achieving a fully voice-based structured graphical user interface interaction.

[0109] In practical applications, a structured graphical user interface may contain multiple card components, such as... Figure 6 As shown, the first graphical user interface 601 contains multiple card components.

[0110] In this scenario, the terminal first uses a card detection model to divide the structured graphical user interface image into multiple images corresponding to different cards, such as... Figure 6 The second graphical user interface 602 is divided by dotted lines to obtain interface images corresponding to the first card 603, the second card 604, the third card 605, and the fourth card 606, respectively. Then, the steps in the embodiments of this application are executed for each interface image. If the data of a certain card changes, the voice broadcast process of the interface image corresponding to that card can be triggered to regenerate the interactive text.

[0111] After the terminal broadcasts key information in the structured graphical user interface, the user can also interact with the terminal by voice, issuing voice commands. After receiving the user's voice commands, the terminal executes the corresponding operations, completing the human-computer interaction process. The following is an illustrative example illustrating the process of the terminal receiving user voice commands and performing voice interaction.

[0112] Figure 7 The illustration shows a schematic diagram of a voice interaction process provided in an exemplary embodiment of this application, the process including:

[0113] Step 701: Upon receiving a voice command, generate the command text corresponding to the voice command.

[0114] After receiving a voice command, the terminal sequentially executes a beamforming algorithm, front-end signal processing, and an ASR (Automatic Speech Recognition) algorithm to generate the corresponding command text.

[0115] Optionally, the front-end signal processing uses the ANC (Active Noise Cancellation) algorithm to eliminate ambient noise; the AEC (Acoutic Echo Cancellation) algorithm to eliminate voice echo in the terminal's voice broadcast; and the AGC (Automatic Gain Control) algorithm to adjust the amplitude range of the voice signal so that the amplitude of the processed output signal is stable.

[0116] Step 702: Determine the matching text in the structured graphical user interface that matches the instruction text.

[0117] After the terminal generates the instruction text, it matches the text content information obtained from optical character recognition (OCR) of the structured graphical user interface (GUI) to determine the matching text within the GUI that matches the instruction text. The matching operation can be implemented using a text matching algorithm, which will not be elaborated upon in this embodiment.

[0118] Once the terminal identifies the matching text, it can execute the operation indicated by the voice command based on the position of the text block to which the matching text belongs. However, in practical applications, there may be situations where the matching text is not interactive. Therefore, to avoid program execution errors, it is necessary to determine whether the matching text has an interactive attribute before executing the operation indicated by the voice command.

[0119] Step 703: If the matched text has an interactive attribute, execute the operation indicated by the voice command based on the position of the text block to which the matched text belongs.

[0120] The interactive attribute refers to whether the element corresponding to the matched text can be triggered.

[0121] If the matched text determined by the terminal has interactive attributes, the element at the corresponding position in the structured graphical user interface is determined according to the position of the matched text in the interface image, and the element is triggered to perform the operation indicated by the voice command.

[0122] Step 704: If the matched text does not have interactive attributes, provide interactive prompts.

[0123] If the matching text is determined to lack interactive attributes, the terminal will provide a non-interactive prompt to the user based on the content of the instruction text, and will also broadcast to the user the text information in the graphical user interface that has interactive attributes.

[0124] For example, targeting Figure 2 In the first graphical user interface 201, when the user issues a voice command such as "Click on the check-in counter", the terminal determines that "check-in counter" in the first graphical user interface 201 is non-interactive text. At this time, the terminal's voice broadcast content can be: "The check-in counter you indicated does not support interaction. If you need further operation, you can say flight assistant, view order, or book hotel".

[0125] In this embodiment, after the terminal acquires the user's voice signal, it determines the text that matches the instruction text, and then judges the interactive attributes of the matched text. For interactive matched text, the terminal performs an operation at the corresponding position. For non-interactive matched text, the terminal prompts the user and provides interactive interface content, so that the terminal can achieve complete interaction between the user and the terminal solely through voice, and can guide the user so that the user can more clearly understand the interactive interface content and can more smoothly carry out voice interaction.

[0126] In one possible implementation, if the terminal does not obtain the view tree corresponding to the structured graphical user interface, it obtains the interface image of the structured graphical user interface, and obtains the text semantic features and then the image semantic features, and then generates interactive text for voice broadcasting.

[0127] In another possible implementation, if the terminal obtains the view tree corresponding to the structured graphical user interface, it can generate interactive text corresponding to the structured graphical user interface based on the view tree, and then perform voice broadcast.

[0128] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0129] Figure 8A structural block diagram of a voice broadcasting device according to an embodiment of this application is shown. The device may include:

[0130] The interface image acquisition module 801 is used to acquire the interface image of the structured graphical user interface, which is used to display content in a structured manner.

[0131] The semantic feature acquisition module 802 is used to acquire the text semantic features and image semantic features corresponding to each text block in the interface image, wherein the text blocks are divided based on the distribution of text in the interface image;

[0132] The interactive text generation module 803 is used to generate interactive text corresponding to the structured graphical user interface based on the text semantic features and the image semantic features. The interactive text is used to describe the content contained in the structured graphical user interface.

[0133] The voice broadcast module 804 broadcasts the interactive text via voice.

[0134] Optionally, the semantic feature acquisition module 802 is used for:

[0135] The text location is obtained by performing text detection on the interface image using a text detection model.

[0136] Based on the text location, the interface image is divided to obtain the text block;

[0137] Feature extraction is performed on the text information of the text block to obtain the text semantic features of the text block;

[0138] Feature extraction is performed on the image information of the text block to obtain the image semantic features of the text block.

[0139] Optionally, the semantic feature acquisition module 802 is used for:

[0140] Optical character recognition is performed on the text block to obtain the text content information of the text block;

[0141] Determine the text position information of the text block, the text position information being used to characterize the position of the text block in the interface image;

[0142] The text content information and the text position information are encoded to obtain the text encoding vector of the text block;

[0143] Based on the text encoding vector, feature extraction is performed using a text semantic feature extraction model to obtain the text semantic features of the text block.

[0144] Optionally, the interactive text generation module 803 is used for:

[0145] Based on the text semantic features and the image semantic features, generate semantic text corresponding to the structured graphical user interface;

[0146] Based on the text semantic features and the image semantic features, the interactive text in the structured graphical user interface is determined, and the interactive text is text with interactive attributes.

[0147] The interactive text corresponding to the structured graphical user interface is generated based on the semantic text and the interactive text.

[0148] Optionally, the interactive text generation module 803 includes:

[0149] The fusion feature generation unit is used to fuse the text semantic features and the image semantic features corresponding to each text block to generate semantic fusion features corresponding to each text block;

[0150] A semantic feature generation unit is used to extract and fuse the semantic fusion features to obtain the interface semantic features of the structured graphical user interface.

[0151] The semantic text generation unit is used to input the interface semantic features into the text generation model to obtain the semantic text corresponding to the structured graphical user interface.

[0152] Optionally, the semantic feature generation unit is used for:

[0153] The semantic fusion features corresponding to each text block are extracted and fused to obtain the first interface semantic features of the structured graphical user interface;

[0154] The semantic text generation unit is used for:

[0155] The semantic features of the first interface are input into the first text generation model to obtain the semantic text corresponding to the structured graphical user interface. The first text generation model has text generation and text summarization functions.

[0156] Optionally, the semantic feature generation unit is used for:

[0157] Based on the semantic fusion features corresponding to the text block, the text block is classified by a first classification model to obtain a first classification result. The first classification result is used to characterize whether the text block contains key text.

[0158] Based on the first classification result, a target text block is selected from the text block, and the target text block contains the key text;

[0159] The semantic fusion features corresponding to the target text block are extracted and fused to obtain the second interface semantic features of the structured graphical user interface;

[0160] The semantic text generation unit is used for:

[0161] The semantic features of the second interface are input into the second text generation model to obtain the semantic text corresponding to the structured graphical user interface. The second text generation model has text generation and text summarization functions.

[0162] Optionally, the interactive text generation module 803 is used for:

[0163] Based on the text semantic features and image semantic features corresponding to the text block, the text block is classified by a second classification model to obtain a second classification result. The second classification result is used to characterize whether the text block contains the interactive text.

[0164] Based on the second classification result, the interactive text in the structured graphical user interface is determined.

[0165] Optionally, the device further includes:

[0166] The instruction text generation module is used to generate the instruction text corresponding to the received voice instruction.

[0167] The matching text determination module is used to determine the matching text in the structured graphical user interface that matches the instruction text;

[0168] The instruction execution module is used to execute the operation indicated by the voice instruction based on the position of the text block to which the matched text belongs.

[0169] Optionally, the instruction execution module is used for:

[0170] If the matched text has interactive attributes, the operation indicated by the voice command is executed based on the position of the text block to which the matched text belongs;

[0171] The device further includes:

[0172] An interactive prompt module is used to provide interactive prompts when the matched text does not have the interactive attribute.

[0173] Optionally, the interface image acquisition module 801 is used for:

[0174] If the view tree corresponding to the structured graphical user interface is not obtained, the interface image of the structured graphical user interface is obtained;

[0175] The device further includes:

[0176] The second interactive text generation module is used to generate the interactive text corresponding to the structured graphical user interface based on the view tree obtained from the structured graphical user interface.

[0177] Please refer to Figure 9 This diagram illustrates a structural block diagram of a terminal provided in an exemplary embodiment of this application. The terminal 900 can be implemented as the terminal in the various embodiments described above. The terminal 900 may include one or more components such as a processor 910 and a memory 920.

[0178] The processor 910 may include one or more processing cores. The processor 910 connects to various parts within the terminal 900 using various interfaces and lines, and performs various functions and processes data of the terminal 900 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 920, and by calling data stored in the memory 920. Optionally, the processor 910 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 910 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), Neural-network Processing Unit (NPU), and modem. Specifically, the CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on the touch screen; the NPU is used to implement Artificial Intelligence (AI) functions; and the modem is used to handle wireless communication. Understandably, the aforementioned modem may also be implemented separately as a single chip, rather than being integrated into the processor 910.

[0179] The memory 920 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 920 may include a non-transitory computer-readable storage medium. The memory 920 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 920 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described below, etc.; the data storage area may store data created according to the use of the terminal 900 (such as audio data, phone book, etc.).

[0180] In addition, those skilled in the art will understand that the structure of the terminal 900 shown in the above figures does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the terminal 900 may also include a display screen, camera assembly, microphone, speaker, radio frequency circuit, input unit, sensors (such as accelerometer, angular velocity sensor, light sensor, etc.), audio circuit, WiFi module, power supply, Bluetooth module, etc., which will not be described in detail here.

[0181] This application also provides a computer-readable storage medium storing at least one piece of program code, which is loaded and executed by a processor to implement the voice broadcasting method described in the above embodiments.

[0182] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice broadcasting method provided in various optional implementations of the above aspects.

[0183] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0184] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A voice broadcasting method, characterized in that, The method includes: Obtain the interface image of the structured graphical user interface, which is used to display content in a structured manner; The text semantic features and image semantic features corresponding to each text block in the interface image are obtained, and the text blocks are divided based on the distribution of text in the interface image; Based on the text semantic features and the image semantic features, generate semantic text corresponding to the structured graphical user interface; Based on the text semantic features and the image semantic features, interactive text in the structured graphical user interface is determined. The interactive text is text with interactive attributes, which means that the interface element corresponding to the location of the text block can be triggered. Based on the semantic text and the interactive text, interactive text corresponding to the structured graphical user interface is generated. The interactive text is used to describe the content contained in the structured graphical user interface. The interactive text is obtained by concatenating the semantic text and the interactive text using a template. The interactive text is read aloud via voice.

2. The method according to claim 1, characterized in that, The step of obtaining the text semantic features and image semantic features corresponding to each text block in the interface image includes: The text location is obtained by performing text detection on the interface image using a text detection model. Based on the text location, the interface image is divided to obtain the text block; Feature extraction is performed on the text information of the text block to obtain the text semantic features of the text block; Feature extraction is performed on the image information of the text block to obtain the image semantic features of the text block.

3. The method according to claim 2, characterized in that, The step of extracting features from the text information of the text block to obtain the text semantic features of the text block includes: Optical character recognition is performed on the text block to obtain the text content information of the text block; Determine the text position information of the text block, the text position information being used to characterize the position of the text block in the interface image; The text content information and the text position information are encoded to obtain the text encoding vector of the text block; Based on the text encoding vector, feature extraction is performed using a text semantic feature extraction model to obtain the text semantic features of the text block.

4. The method according to claim 1, characterized in that, The step of generating semantic text corresponding to the structured graphical user interface based on the text semantic features and the image semantic features includes: The semantic features of the text and the semantic features of the image corresponding to each text block are fused to generate semantic fusion features corresponding to each text block; The semantic fusion features are extracted and fused to obtain the interface semantic features of the structured graphical user interface; The semantic features of the interface are input into the text generation model to obtain the semantic text corresponding to the structured graphical user interface.

5. The method according to claim 4, characterized in that, The step of extracting and fusing the semantic fusion features to obtain the interface semantic features of the structured graphical user interface includes: The semantic fusion features corresponding to each text block are extracted and fused to obtain the first interface semantic features of the structured graphical user interface; The step of inputting the interface semantic features into a text generation model to obtain the semantic text corresponding to the structured graphical user interface includes: The semantic features of the first interface are input into the first text generation model to obtain the semantic text corresponding to the structured graphical user interface. The first text generation model has text generation and text summarization functions.

6. The method according to claim 4, characterized in that, The step of extracting and fusing the semantic fusion features to obtain the interface semantic features of the structured graphical user interface includes: Based on the semantic fusion features corresponding to the text block, the text block is classified by a first classification model to obtain a first classification result. The first classification result is used to characterize whether the text block contains key text. Based on the first classification result, a target text block is selected from the text block, and the target text block contains the key text; The semantic fusion features corresponding to the target text block are extracted and fused to obtain the second interface semantic features of the structured graphical user interface; The step of inputting the interface semantic features into a text generation model to obtain the semantic text corresponding to the structured graphical user interface includes: The semantic features of the second interface are input into the second text generation model to obtain the semantic text corresponding to the structured graphical user interface. The second text generation model has text generation and text summarization functions.

7. The method according to claim 1, characterized in that, The step of determining the interactive text in the structured graphical user interface based on the text semantic features and the image semantic features includes: Based on the text semantic features and image semantic features corresponding to the text block, the text block is classified by a second classification model to obtain a second classification result. The second classification result is used to characterize whether the text block contains the interactive text. Based on the second classification result, the interactive text in the structured graphical user interface is determined.

8. The method according to claim 1, characterized in that, The method further includes: Upon receiving a voice command, generate the command text corresponding to the voice command; Determine the matching text in the structured graphical user interface that matches the instruction text; Based on the position of the text block to which the matched text belongs, the operation indicated by the voice command is executed.

9. The method according to claim 8, characterized in that, The step of executing the operation indicated by the voice command based on the position of the text block to which the matched text belongs includes: If the matched text has interactive attributes, the operation indicated by the voice command is executed based on the position of the text block to which the matched text belongs; The method further includes: If the matched text does not have the interactive attribute, provide an interactive prompt.

10. The method according to claim 1, characterized in that, The process of obtaining the interface image of the structured graphical user interface includes: If the view tree corresponding to the structured graphical user interface is not obtained, the interface image of the structured graphical user interface is obtained; The method further includes: If the view tree corresponding to the structured graphical user interface is obtained, the interactive text corresponding to the structured graphical user interface is generated based on the view tree.

11. A voice broadcasting device, characterized in that, The device includes: The interface image acquisition module is used to acquire the interface image of the structured graphical user interface, which is used to display content in a structured manner. The semantic feature acquisition module is used to acquire the text semantic features and image semantic features corresponding to each text block in the interface image, wherein the text blocks are divided based on the distribution of text in the interface image; An interactive text generation module is used to generate semantic text corresponding to the structured graphical user interface based on the text semantic features and the image semantic features; Based on the text semantic features and the image semantic features, interactive text in the structured graphical user interface is determined. The interactive text is text with interactive attributes, which means that the interface element corresponding to the location of the text block can be triggered. Based on the semantic text and the interactive text, interactive text corresponding to the structured graphical user interface is generated. The interactive text is used to describe the content contained in the structured graphical user interface. The interactive text is obtained by concatenating the semantic text and the interactive text using a template. The voice broadcast module broadcasts the interactive text via voice.

12. A terminal, characterized in that, The terminal includes a processor and a memory; the memory stores at least one instruction, which is executed by the processor to implement the voice broadcasting method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the voice broadcasting method as described in any one of claims 1 to 10.

14. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice broadcasting method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Voice broadcast method and device, electronic device and computer readable storage medium

    CN109462689A

  • Visual question and answer method and device, electronic equipment and storage medium

    CN114445826A