Display method, device and electronic equipment

CN115291826BActive Publication Date: 2026-08-18VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210927614.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-03
Publication Date
2026-08-18
Estimated Expiration
2042-08-03

AI Technical Summary

Technical Problem

[0004]本申请实施例的目的是提供一种显示方法、装置和电子设备,能够解决电子设备与智能语音系统交互的效率低的问题

Benefits of technology

[0016]第六方面,本申请实施例提供一种计算机程序产品,该程序产品被存储在存储介质中,该程序产品被至少一个处理器执行以实现如第一方面所述的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115291826B_ABST
    Figure CN115291826B_ABST
Patent Text Reader

Abstract

The application discloses a display method, device and electronic equipment, and belongs to the technical field of electronics. The method comprises the following steps: in the case that audio transmitted by an intelligent voice system is received, the audio is converted into text to obtain target text; N text contents are extracted from the target text, wherein different text contents are used for indicating different voice services, and N is a positive integer; and the N text contents are displayed in correspondence with N controls, wherein different controls correspond to different text contents, and the controls are used for allowing a user to select a voice service indicated by the text content corresponding to the control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic technology, specifically relating to a display method, device, and electronic device. Background Technology

[0002] Intelligent voice systems are based on Natural Language Processing (NLP), Automatic Speech Recognition (ASR), and Text-to-Speech (TTS) technologies to enable outbound and inbound voice calls. They can communicate with customers using natural and realistic dialogue, helping businesses improve outbound call efficiency.

[0003] Currently, users typically interact with intelligent voice systems through electronic devices such as smartphones and tablets, allowing them to select desired services based on the audio. However, because intelligent voice systems can speak too quickly, users need to repeatedly listen to the audio when interacting with them via electronic devices, resulting in low efficiency. Summary of the Invention

[0004] The purpose of this application is to provide a display method, apparatus, and electronic device that can solve the problem of low efficiency in the interaction between electronic devices and intelligent voice systems.

[0005] In a first aspect, embodiments of this application provide a display method, including:

[0006] Upon receiving audio transmitted by the intelligent voice system, the audio is converted into text to obtain the target text;

[0007] Extract N text contents from the target text, where different text contents are used to indicate different voice services, and N is a positive integer;

[0008] The N text contents are displayed in N corresponding controls, with different controls corresponding to different text contents, and the controls are used for users to select the voice service indicated by the corresponding text content.

[0009] Secondly, embodiments of this application provide a display device, including:

[0010] An audio conversion module is used to convert audio received from an intelligent voice system into text to obtain target text.

[0011] The text content extraction module is used to extract N text contents from the target text, where different text contents are used to indicate different voice services, and N is a positive integer.

[0012] The display module is used to display the N text contents corresponding to N controls, with different controls corresponding to different text contents, and the controls are used for users to select the voice service indicated by the corresponding text content.

[0013] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0014] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0015] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0016] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0017] In this embodiment, upon receiving audio transmitted by the intelligent voice system, the audio is first converted into text to obtain the target text. Then, N text contents indicating different voice services are extracted from the target text. Finally, the N text contents are displayed corresponding to N controls. Thus, during a voice call between the user and the intelligent voice system via an electronic device, the user can operate the controls to select the desired voice service using the text contents displayed on the electronic device. This reduces the occurrence of the user repeatedly listening to the audio output by the voice system, thereby improving the efficiency of the user's interaction with the intelligent voice system via the electronic device. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating an embodiment of the display method provided in this application;

[0019] Figure 2 This is a schematic diagram of the display interface in an embodiment of the display method provided in this application;

[0020] Figure 3This is another schematic diagram of the display interface in an embodiment of the display method provided in this application;

[0021] Figure 4 This is another schematic diagram of the display interface in an embodiment of the display method provided in this application;

[0022] Figure 5 This is another schematic diagram of the display interface in an embodiment of the display method provided in this application;

[0023] Figure 6 This is a schematic diagram of the structure of an embodiment of the display device provided in this application;

[0024] Figure 7 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application;

[0025] Figure 8 This is a schematic diagram of another embodiment of the electronic device provided in this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0027] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0028] The display method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0029] Please see Figure 1 This is a flowchart illustrating the display method provided in an embodiment of this application. This display method is applied to electronic devices, such as... Figure 1 As shown, the method includes the following steps:

[0030] Step 101: Upon receiving audio transmitted by the intelligent voice system, convert the audio into text to obtain the target text;

[0031] Step 102: Extract N text contents from the target text. Different text contents are used to indicate different voice services, and N is a positive integer.

[0032] Step 103: Display N text contents corresponding to N controls. Different controls correspond to different text contents, and the controls are used for users to select the voice service indicated by the corresponding text content.

[0033] In this embodiment, upon receiving audio transmitted by the intelligent voice system, the audio is first converted into text to obtain the target text. Then, N text contents indicating different voice services are extracted from the target text. Finally, the N text contents are displayed corresponding to N controls. Thus, during a voice call between the user and the intelligent voice system via an electronic device, the user can operate the controls to select the desired voice service using the text contents displayed on the electronic device. This reduces the occurrence of the user repeatedly listening to the audio output by the voice system, thereby improving the efficiency of the user's interaction with the intelligent voice system via the electronic device.

[0034] In step 101 above, upon receiving audio transmitted by the intelligent voice system, the audio is converted into text to obtain the target text.

[0035] The above-mentioned conversion of audio into text to obtain the target text can be achieved by having an audio-to-text tool pre-installed in the electronic device. This tool can convert the received audio in real time to obtain the target text.

[0036] For example, during the interaction between the aforementioned electronic device and the intelligent voice system, the electronic device can use its audio-to-text tool to convert the received audio into the following text 1 (i.e., the target text) in real time:

[0037] To access service A1, press 1; to access service A2, press 2; to access service A3, press A3; to access service A5, press 5; to access service A6, press 6; to access service A7, press 7; to access service A8, press 8; to access service A9, press 9, and so on.

[0038] In step 102 above, after the electronic device obtains the target text, the electronic device can extract N text contents from the target text.

[0039] The above N text contents are used to indicate different voice services. For example, the following 9 text contents can be extracted from the above text 1: If you need service A1, If ​​you need service A2, If you need service A3, If you need service A4, If you need service A5, If you need service A6, If you need service A7, If you need service A8 and If you need service A9.

[0040] The above-mentioned extraction of N text contents from the target text can be achieved by the electronic device having multiple preset fields, and the target text being divided into multiple text contents through these preset fields; or by using all the contents of the target text as the above-mentioned text contents if the target text does not have the aforementioned preset fields.

[0041] For example, the electronic device has a preset field "Please press X". When the target text is the text 1, the electronic device can use the text before each "Please press X" as the text content, that is, it can extract the above 9 text contents: If you need to provide services A1, If ​​you need to provide services A2, If you need to provide services A3, If you need to provide services A4, If you need to provide services A5, If you need to provide services A6, If you need to provide services A7, If you need to provide services A8 and If you need to provide services A9.

[0042] In step 103 above, after the electronic device extracts the above N text contents, the electronic device can display the corresponding N controls so that the user can operate each control according to the text content to select the voice service indicated by the text content corresponding to the control.

[0043] The above-mentioned display of N text contents corresponding to N controls can be achieved by an electronic device displaying N controls on its display interface, with each of the N controls displaying its corresponding text content.

[0044] For example, when the electronic device displays a desktop interface, after extracting the aforementioned nine text contents, the electronic device can replace the displayed content of the nine application icons on the desktop interface with the aforementioned nine text contents. When the user clicks on any of the nine application icons based on their text contents, the electronic device can send a command to the intelligent voice system to request the intelligent voice system to provide the voice service indicated by the text content displayed on the clicked control. For example, if the user clicks on control 3 which displays "Service A3 required," the electronic device can request the intelligent voice system to provide service A3. In this case, the aforementioned nine application icons become nine controls.

[0045] In some implementations, the above-described conversion of audio into text to obtain the target text includes:

[0046] The audio is converted into text, and identifiers are added to the converted text to obtain target text containing multiple identifiers. The identifiers are used to identify the start, end and interruption of the audio content.

[0047] The above extraction of N text contents from the target text includes:

[0048] Based on multiple identifiers, the target text is divided into multiple subtexts, each of which includes the text between two identifiers;

[0049] Among multiple subtexts, at least one subtext is identified as having N text contents.

[0050] In this embodiment, by adding an identifier to the target text during the generation of the target text, dividing the target text into multiple sub-texts based on the identifier, and determining N text contents in the multiple sub-texts, the efficiency and accuracy of extracting the N text contents from the target text can be further improved.

[0051] The aforementioned addition of identifiers to the converted text can be achieved by an electronic device adding identifiers to the converted text based on the audio content. These identifiers can include start and end characters and separators. Start and end characters indicate the beginning and end of the audio content, and the audio content represented by these characters can be all or part of the audio; while separators indicate breaks in the audio content.

[0052] For example, if the pause in a received audio segment is longer than or equal to a preset duration, add the symbol "|" (i.e., separator) after the text converted before the pause; if the audio content of a received segment is "Please press [number, *, #]", add the symbol "@|" (i.e., start and end character) after the text converted from that segment; and add the symbol "@|" at the beginning of the audio segment, etc.

[0053] The above method of dividing the target text into multiple subtexts based on multiple identifiers can be achieved by taking the text between any two adjacent identifiers as subtexts.

[0054] Alternatively, if the identifiers mentioned above include delimiters and start / end characters, the above method of dividing the target text into multiple subtexts based on multiple identifiers may include: taking the text between two adjacent start / end characters as subtexts.

[0055] For example, suppose the target text with added identifiers is: @|a service|c service|d service|Press 1@|b service|Press 2@|…..@|X service|Press 9@|. Then, the electronic device can segment the target text according to the symbol "@|" to obtain the subtext: "a service|c service|d service|Press 1", "|b service|Press 2", ... and "X service|Press 9".

[0056] The above method of determining at least one subtext as N text contents among multiple subtexts may include: obtaining a keyword list corresponding to the intelligent voice system; matching each subtext with a keyword in the keyword list; and determining at least one subtext that matches a keyword as N text contents.

[0057] The aforementioned acquisition of the keyword list corresponding to the intelligent voice system can be achieved when the electronic device and the intelligent voice system are connected via voice, and the intelligent voice system sends its stored keyword list to the electronic device.

[0058] For example, assuming the above keyword list includes keywords: a, b, c, d, ... and X, and the electronic device obtains the subtext: "a service|c service|d service|Press 1", "|b service|Press 2", ... and "X service|Press 9", then the electronic device can obtain the above N text contents as "a service, c service, d service", "b service", ... and "X service".

[0059] In some implementations, determining at least one subtext as N text contents among multiple subtexts includes:

[0060] Among multiple subtexts, at least one subtext is identified that matches the service instruction information of each service option in the intelligent voice interaction interface, wherein the intelligent voice interaction interface includes at least one service option and each service option displays service instruction information.

[0061] Based on at least one subtext matched by each service option, obtain the text content that matches the service instruction information of the service option, and obtain N text contents that match the service instruction information of N service options.

[0062] In this embodiment, by extracting N text contents matching the service indication information of N service options from the target text, the efficiency and accuracy of extracting the aforementioned N text contents from the target text can be further improved.

[0063] During a voice call between an electronic device and an intelligent voice system, the electronic device can display the aforementioned intelligent voice interaction interface. This intelligent voice interaction interface includes at least one service option, and each service option displays service instruction information. This service instruction information can be text corresponding to a portion of the audio in the aforementioned audio, and it can correspond to a voice service that the intelligent voice system can provide.

[0064] For example, during a voice call between the aforementioned electronic device and the intelligent voice system, the electronic device can display, for instance, the following: Figure 2 The intelligent voice interaction interface shown includes 10 numeric virtual buttons (i.e., service options), and each numeric virtual button displays a corresponding number. For example, numeric virtual button 21 displays the number "7", and so on.

[0065] The above-mentioned determination of at least one subtext that matches the service indication information of each service option among multiple subtexts may be to match each subtext with the service indication information of each service option, and determine at least one subtext that matches the service indication information of the service option among at least one subtext containing the service indication information of the service option.

[0066] For example, in electronic devices displaying such Figure 2 In the case of the intelligent voice interaction interface shown, if the above subtext includes: "a service|c service|d service|Please press 1", "|b service|Please press 2", ... and "X service|Please press 9", then the electronic device can determine "a service|c service|d service|Please press 1" as the subtext that matches the virtual key displaying the number "1", determine "|b service|Please press 2" as the subtext that matches the virtual key displaying the number "2", ..., and determine "|X service|Please press 9" as the subtext that matches the virtual key displaying the number "9".

[0067] The above-mentioned method of obtaining text content that matches the service instruction information of the service options based on at least one subtext matched by each service option, and obtaining N text contents that match the service instruction information of N service options, can be achieved by constructing a mapping table based on at least one subtext matched by each of the N service options. The mapping table includes N mapping relationships, each mapping relationship being the relationship between a service option and at least one service content text, and each service content text being extracted from the subtext matched by the service option. The text content is then obtained based on at least one service content in each mapping relationship, resulting in N text contents.

[0068] For example, if it is determined that "a service|c service|d service|press 1" matches a virtual key displaying the number "1", "|b service|press 2" matches a virtual key displaying the number "2", ... and "|X service|press 9" matches a virtual key displaying the number "9", the electronic device can segment each sub-text according to the separator "|" and add the segmented operation text and the numbers of the virtual keys to a mapping table. In this mapping table, "a service", "c service", "d service" are mapped to the virtual key with the number "1", "b service" is mapped to the virtual key with the number "2", ..., and "X service" is mapped to the virtual key with the number "9".

[0069] The above method of obtaining text content based on at least one service content in each mapping relationship to obtain N text content can be achieved by using all service content in each mapping relationship as text content. For example, for the mapping relationship between "service a", "service c", "service d" and the virtual key with the number "1", "service a c service d service" can be used as text content.

[0070] In some implementations, obtaining the text content matching the service instruction information of each service option based on at least one subtext matched by each service option includes:

[0071] Extract the content of at least one subtext that matches each service option to obtain the text content that matches the service instruction information of the service option.

[0072] In this embodiment, by refining the content of at least one subtext that matches each service option, text content that matches the service instruction information of the service option is obtained, thereby making the displayed text content more concise and improving the display effect.

[0073] The above-mentioned extraction of the content of at least one subtext matching each service option can be achieved by electronic devices using Named Entity Recognition (NER) technology to extract the content of at least one subtext matching each service option and obtain the text content.

[0074] For example, regarding the mapping relationship between "service a", "service c", and "service d" and the virtual key with the number "1", we can extract "service a", "service c", and "service d" to obtain the text content "acd" that matches the virtual key with the number "1"; regarding the mapping relationship between "service b" and the virtual key with the number "2", we can extract "service b" to obtain the text content "b" that matches the virtual key with the number "1", and so on.

[0075] In some implementations, the above-described conversion of audio into text to obtain the target text includes:

[0076] The first part of the audio is converted into text to obtain the first sub-text.

[0077] If the first subtext matches the service indication information of the preset target service option and the value of the preset flag variable is the first preset value, the second part of the audio is converted into text to obtain the second subtext, and the value of the flag variable is updated to the second preset threshold. The target service option is at least one service option in the intelligent voice interaction interface, the second part of the audio is the part of the audio received after the first part of the audio, and the target text includes the first subtext and the second subtext.

[0078] If the first subtext matches the service instruction information of the preset target service option and the value of the flag variable is the second preset value, stop converting the second part of the audio into text and find the target subtext including the first subtext on the electronic device.

[0079] In this embodiment, during the process of the electronic device receiving audio, it can determine whether to convert the entire audio based on the first sub-text obtained from the first part of the audio conversion, the service indication information of the target service option, and the variable of the flag variable, or to search for the pre-stored target text in the electronic device, thereby avoiding repeated audio conversion.

[0080] The first part of the audio mentioned above can be a portion of the audio converted to the target text. For example, the first part of the audio could be a portion of the received audio containing the message "Service a, Service c, Service d, please press 1", etc.

[0081] The target service option can be any service option in the intelligent voice interaction interface. Specifically, the target service option is usually the service option whose service instruction information is played first. For example, the target service option can be a virtual button with the number "1".

[0082] After the electronic device obtains the first subtext, it can match the first subtext with the service indication information of the target service option. For example, if the first subtext is "service a, service c, service d, please press 1", and the target service option is a virtual key with the number "1", the electronic device can determine that the first subtext matches the service indication information of the target service option.

[0083] The aforementioned flag variable can be any variable defined as needed, and the value of the flag variable can include a first preset value and a second preset value. The first preset value is used to indicate that there has been no previous successful matching between the first subtext and the service indication information of the target service option; the second preset value is used to indicate that there has been a previous successful matching between the first subtext and the service indication information of the target service option.

[0084] If the first subtext matches the service instruction information of the preset target service option, and the value of the preset flag variable is the first preset value, the electronic device can convert the second part of the audio into text to obtain the second subtext, and update the value of the flag variable to the second preset threshold.

[0085] For example, suppose the electronic device initializes the start flag variable (i.e., the preset flag variable mentioned above) StartFlag to False (i.e., the first preset value). If the first sub-text is "Service a, Service c, Service d, press 1", and the target service option is a virtual button with the number "1", and StartFlag is False, the electronic device can convert the remaining audio into text, obtaining the second sub-text "Service b, press 2; ...; Service X, press 9", and update StartFlag to True, and save "Service a, Service c, Service d, press 1; Service b, press 2; ...; Service X, press 9". If the first sub-text matches the service indication information of the preset target service option, and the flag variable's value is the second preset value, the conversion of the second part of the audio into text stops, and the electronic device finds the target sub-text including the first sub-text.

[0086] For example, if the first subtext is “service a, service c, service d, please press 1”, the target service option is a virtual button with the number “1”, and StartFlag is True, the electronic device determines that it has previously converted the audio and obtained the target text. Therefore, the electronic device finds the target text of the audio based on “service a, service c, service d, please press 1”.

[0087] In some implementations, the above-described method of displaying N text contents corresponding to N controls includes:

[0088] On the intelligent voice interaction interface, N text entries are displayed corresponding to N service options; or...

[0089] A pop-up window is displayed in the display interface. The pop-up window includes N controls, and N text contents are displayed on the N controls.

[0090] In this embodiment, N text contents can be displayed on the intelligent voice interaction interface corresponding to N service options, or a pop-up window including N controls displaying N text contents can be displayed on the display interface, thereby making the way of displaying N text contents more flexible.

[0091] The display interface of the aforementioned electronic device may be an intelligent voice interaction interface; or it may be other interfaces besides the aforementioned intelligent voice interaction interface, such as an application interface, etc.

[0092] For example, when an electronic device is communicating with an intelligent voice system and the electronic device displays an instant messaging application interface, the electronic device can display a pop-up window in the instant messaging application interface. The pop-up window can display 9 controls, and each of the 9 controls displays the aforementioned 9 text contents.

[0093] In some implementations, the above-mentioned display of N text contents corresponding to N service options on the intelligent voice interaction interface includes:

[0094] Display N text items among N service options in the intelligent voice interaction interface, where each service option includes service instructions and text content; or,

[0095] Update the service instructions for each of the N service options in the intelligent voice interaction interface to N text contents.

[0096] In this embodiment, the electronic device may display the text content and service instruction information side by side on the service options, or it may replace the service instruction information of the service options with the text content, thereby making the display method more flexible.

[0097] The aforementioned electronic device can display N text contents in N service options, and displays N text contents in N service options.

[0098] For example, such as Figure 3 As shown, during a voice call between an electronic device and a smart voice system for travel services, the electronic device can display each text content alongside the corresponding service instructions, such as displaying the number "1" alongside "passenger transport", the number "2" alongside "freight transport", and so on.

[0099] The aforementioned electronic device can also update the service instruction information of each of the N service options into N text contents, thereby enlarging and displaying the N text contents to improve the display effect.

[0100] For example, such as Figure 4As shown, during a voice call between an electronic device and a smart voice system for travel services, the electronic device can update the number "1" to "passenger transport", the number "2" to "freight transport", and so on.

[0101] In some implementations, the above-mentioned display of N text contents corresponding to N service options on the intelligent voice interaction interface further includes:

[0102] If the service instructions for each of the N service options in the intelligent voice interaction interface are updated to N text contents, then the display positions of the N service options are updated.

[0103] In this embodiment, when the service instruction information of each of the N service options is updated to N text content, the electronic device can also update the display position of the N service options, thereby further improving the display effect of the electronic device.

[0104] For example, if the above N service options are distributed in a rectangular array on the intelligent voice interaction interface, the electronic device can update the N service options to be distributed in a circular array, and so on.

[0105] In this embodiment of the application, when the above-mentioned intelligent voice interaction interface displays the above-mentioned N controls, when the electronic device receives touch input (such as click or press input) to any control, the electronic device can send an instruction to the intelligent voice system. The instruction is used to instruct the intelligent voice system to provide the voice service indicated by the text content corresponding to the control targeted by the touch input.

[0106] In some implementations, the aforementioned intelligent voice interaction interface displays N controls and a first control, with the N controls distributed around the first control.

[0107] After displaying N text contents corresponding to N controls as described above, it can also include:

[0108] Upon receiving target input to the first control and the second control among the N controls, in response to the target input, a target instruction is sent to the intelligent voice system. The target instruction is used to indicate the voice service indicated by the text content corresponding to the second control.

[0109] In this embodiment, when the electronic device receives the target input, it can send a target instruction to the intelligent voice system to instruct the intelligent voice system to provide the voice service indicated by the text content corresponding to the second control, thereby reducing user misoperation.

[0110] The target input can be any input made to the first control and the second control. For example, the target input can be a sliding operation between the first control and the second control; or it can be a clicking operation of clicking the first control and the second control in sequence; or it can be a dragging operation of dragging the first control onto the second control, and so on.

[0111] For example, such as Figure 5 As shown, if the electronic device receives an operation from the user who first presses the "press and slide" control (i.e., the first control) and slides to the "freight" control (i.e., the second control), the electronic device can send a target instruction to the intelligent voice system. The target instruction is used to instruct the intelligent voice system to provide freight voice service.

[0112] In addition, the aforementioned electronic devices can also update the background and other aspects of the aforementioned intelligent voice interaction interface, which will not be elaborated here.

[0113] The display method provided in this application can be executed by a display device. This application uses a display device executing the display method as an example to illustrate the display device provided in this application.

[0114] Please see Figure 6 This is a schematic diagram of the structure of the display device provided in the embodiments of this application. Figure 6 The display device 600 shown includes:

[0115] The audio conversion module 601 is used to convert audio into text upon receiving audio transmitted from the intelligent voice system, thereby obtaining the target text.

[0116] The text content extraction module 602 is used to extract N text contents from the target text. Different text contents are used to indicate different voice services, and N is a positive integer.

[0117] Display module 603 is used to display N text contents corresponding to N controls, with different controls corresponding to different text contents, and the controls are used for users to select the voice service indicated by the corresponding text contents.

[0118] In some embodiments, the audio conversion module 601 is specifically used for:

[0119] The audio is converted into text, and identifiers are added to the converted text to obtain target text containing multiple identifiers. These identifiers are used to identify the start, end, and interruptions of the audio content.

[0120] The text content extraction module 602 may include:

[0121] A text segmentation unit is used to segment target text into multiple subtexts based on multiple identifiers, each subtext including the text between two identifiers;

[0122] The text content determination unit is used to determine at least one subtext as N text contents among multiple subtexts.

[0123] In some implementations, the text content determination unit includes:

[0124] A matching sub-unit is used to identify at least one sub-text that matches the service indication information of each service option among multiple sub-texts;

[0125] The text content acquisition subunit is used to acquire the text content that matches the service indication information of each service option based on at least one subtext matched by each service option, and to obtain N text contents that match the service indication information of N service options.

[0126] In some implementations, the text content acquisition unit is specifically used for:

[0127] Extract the content of at least one subtext that matches each service option to obtain the text content that matches the service instruction information of the service option.

[0128] In some embodiments, the audio conversion module 601 includes:

[0129] The first conversion unit is used to convert the first part of the audio into text to obtain the first sub-text.

[0130] The second conversion unit is used to convert the second part of the audio into text to obtain the second subtext when the first subtext matches the service indication information of the preset target service option and the value of the preset flag variable is a first preset value, and to update the value of the flag variable to a second preset threshold. The target service option is at least one service option in the intelligent voice interaction interface, the second part of the audio is the part of the audio received after the first part of the audio, and the target text includes the first subtext and the second subtext.

[0131] The search unit is used to stop converting the second part of the audio into text and to find the target subtext including the first subtext in the electronic device when the first subtext matches the service indication information of the preset target service option and the value of the flag variable is a second preset value.

[0132] In some implementations, the display module 603 is specifically used for:

[0133] On the intelligent voice interaction interface, N text entries are displayed corresponding to N service options; or...

[0134] A pop-up window is displayed in the display interface. The pop-up window includes N controls, and N text contents are displayed on the N controls.

[0135] In some implementations, the display module 603 is specifically used for:

[0136] Display N text items among N service options in the intelligent voice interaction interface, where each service option includes service instructions and text content; or,

[0137] Update the service instructions for each of the N service options in the intelligent voice interaction interface to N text contents.

[0138] In some embodiments, the display module 603 is further configured to:

[0139] The update module is used to update the display position of the N service options when the service indicator information of each service option in the N service options of the intelligent voice interaction interface is updated to N text content.

[0140] In some implementations, the intelligent voice interaction interface displays N controls and a first control, with the N controls distributed around the first control.

[0141] The device 600 may further include:

[0142] The sending module is used to send a target instruction to the intelligent voice system in response to the target input when it receives target input to the first control and the second control among N controls. The target instruction is used to indicate the voice service indicated by the text content corresponding to the second control.

[0143] The display device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0144] The display device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0145] The display device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments can achieve the same technical effect, and will not be described again here to avoid repetition.

[0146] Optionally, such as Figure 7 As shown, this application embodiment also provides an electronic device 700, including a processor 701 and a memory 702. The memory 702 stores a program or instructions that can run on the processor 701. When the program or instructions are executed by the processor 701, they implement the various steps of the above-described display method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0147] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.

[0148] Figure 8 A schematic diagram of the hardware structure of the electronic device used to implement the embodiments of this application.

[0149] The electronic device 800 includes, but is not limited to, components such as: radio frequency unit 801, network module 802, audio output unit 803, input unit 804, sensor 805, display unit 806, user input unit 807, interface unit 808, memory 809, and processor 810.

[0150] Those skilled in the art will understand that the electronic device 800 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 810 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0151] The processor 810 is used for:

[0152] Upon receiving audio transmitted from the intelligent voice system, the audio is converted into text to obtain the target text;

[0153] Extract N text contents from the target text. Different text contents are used to indicate different voice services. N is a positive integer.

[0154] N text contents are displayed as N controls, with different controls corresponding to different text contents. The controls are used to allow users to select the voice service indicated by the corresponding text content.

[0155] In some implementations, the processor 810 is also configured to:

[0156] The audio is converted into text, and identifiers are added to the converted text to obtain target text containing multiple identifiers. The identifiers are used to identify the start, end and interruption of the audio content.

[0157] Based on the plurality of identifiers, the target text is divided into a plurality of subtexts, each of the subtexts comprising the text between two identifiers;

[0158] Among the plurality of subtexts, at least one of the subtexts is identified as one of the N text contents.

[0159] In some implementations, the processor 810 is also configured to:

[0160] Among the plurality of subtexts, at least one subtext is determined to match the service indication information of each service option in the intelligent voice interaction interface, wherein the intelligent voice interaction interface includes at least one of the service options and each of the service options displays service indication information.

[0161] Based on at least one subtext matched by each of the service options, obtain the text content that matches the service indication information of the service option, and obtain N text contents that match the service indication information of N service options.

[0162] In some implementations, the processor 810 is also configured to:

[0163] Extract the content of at least one subtext that matches each service option to obtain the text content that matches the service instruction information of the service option.

[0164] In some implementations, the processor 810 is also configured to:

[0165] The first part of the audio is converted into text to obtain the first sub-text.

[0166] If the first subtext matches the service indication information of the preset target service option and the value of the preset flag variable is the first preset value, the second part of the audio is converted into text to obtain the second subtext, and the value of the flag variable is updated to the second preset threshold. The target service option is at least one service option in the intelligent voice interaction interface, the second part of the audio is the part of the audio received after the first part of the audio, and the target text includes the first subtext and the second subtext.

[0167] If the first subtext matches the service instruction information of the preset target service option and the value of the flag variable is the second preset value, stop converting the second part of the audio into text and find the target subtext including the first subtext on the electronic device.

[0168] In some implementations, the processor 810 is also configured to:

[0169] The N text contents are displayed on the intelligent voice interaction interface corresponding to the N service options; or...

[0170] A pop-up window is displayed in the display interface, and the pop-up window includes N controls, on which the N controls display the corresponding N text contents.

[0171] In some implementations, the processor 810 is also configured to:

[0172] Display N text items among N service options in the intelligent voice interaction interface, where each service option includes service instructions and text content; or,

[0173] Update the service instructions for each of the N service options in the intelligent voice interaction interface to N text contents.

[0174] In some implementations, the processor 810 is also configured to:

[0175] If the service instructions for each of the N service options are updated to N text contents, then the display position of the N service options is updated.

[0176] In some implementations, the aforementioned intelligent voice interaction interface displays the N controls and a first control, with the N controls distributed around the first control.

[0177] The 810 processor is also used for:

[0178] Upon receiving target input to the first control and the second control among the N controls, in response to the target input, a target instruction is sent to the intelligent voice system. The target instruction is used to indicate the voice service indicated by the text content corresponding to the second control.

[0179] The display device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments can achieve the same technical effect, and will not be described again here to avoid repetition.

[0180] It should be understood that, in this embodiment, the input unit 804 may include a graphics processing unit (GPU) 8041 and a microphone 8042. The GPU 8041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 806 may include a display panel 8061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 807 includes at least one of a touch panel 8071 and other input devices 8072. The touch panel 8071 is also called a touch screen. The touch panel 8071 may include a touch detection device and a touch controller. Other input devices 8072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0181] The memory 809 can be used to store software programs and various data. The memory 809 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 809 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 809 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0182] Processor 810 may include one or more processing units; optionally, processor 810 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 810.

[0183] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described display method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0184] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0185] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described display method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0186] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0187] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0188] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0189] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0190] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A display method, characterized in that, include: Upon receiving audio transmitted by an intelligent voice system, the audio is converted into text, and identifiers are added to the converted text to obtain target text containing multiple identifiers. These identifiers are used to identify the start, end, and interruptions of the audio content. Based on the plurality of identifiers, the target text is divided into a plurality of subtexts, each of the subtexts comprising the text between two identifiers; Among the plurality of subtexts, at least one subtext is determined to match the service indication information of each service option in the intelligent voice interaction interface, wherein the intelligent voice interaction interface includes at least one of the service options and each of the service options displays service indication information. Based on at least one subtext matched by each of the service options, obtain the text content that matches the service indication information of the service option, and obtain N text contents that match the service indication information of N service options. Different text contents are used to indicate different voice services, where N is a positive integer. The N text contents are displayed in N corresponding controls, with different controls corresponding to different text contents, and the controls are used for users to select the voice service indicated by the corresponding text content.

2. The method according to claim 1, characterized in that, The method includes: The first part of the audio is converted into text to obtain the first sub-text. If the first sub-text matches the service indication information of the preset target service option and the value of the preset flag variable is the first preset value, the second part of the audio is converted into text to obtain the second sub-text, and the value of the flag variable is updated to the second preset value. The target service option is at least one service option in the intelligent voice interaction interface, the second part of the audio is the part of the audio received after the first part of the audio, and the target text includes the first sub-text and the second sub-text. If the first subtext matches the service indication information of the preset target service option, and the value of the flag variable is the second preset value, then the conversion of the second part of the audio into text is stopped, and the target text including the first subtext is found on the electronic device.

3. The method according to claim 1, characterized in that, The step of displaying the N text contents corresponding to N controls includes: The N text contents are displayed on the intelligent voice interaction interface corresponding to the N service options; or... A pop-up window is displayed in the display interface, and the pop-up window includes N controls, on which the N controls display the corresponding N text contents.

4. The method according to claim 3, characterized in that, The display of the N text contents corresponding to the N service options on the intelligent voice interaction interface includes: The N text contents are displayed among N service options in the intelligent voice interaction interface, and each service option includes the service instruction information and the text content; or, Update the service instruction information of the N service options in the intelligent voice interaction interface to the N text content.

5. The method according to claim 4, characterized in that, The method of displaying the N text contents corresponding to the N service options on the intelligent voice interaction interface also includes: When the service instruction information of the N service options in the intelligent voice interaction interface is updated to the N text content, the display position of the N service options is updated.

6. A display device, characterized in that, include: The audio conversion module is used to convert the audio received from the intelligent voice system into text, and add identifiers to the converted text to obtain target text including multiple identifiers. The identifiers are used to identify the start, end and interruption of the audio content. The text content extraction module is used to segment the target text into multiple sub-texts based on the multiple identifiers, each sub-text including text between two identifiers; among the multiple sub-texts, determine at least one sub-text that matches the service indication information of each service option in the intelligent voice interaction interface, the intelligent voice interaction interface including at least one of the service options, and each service option displaying service indication information; based on the at least one sub-text matched by each service option, obtain the text content that matches the service indication information of the service option, to obtain N text contents that match the service indication information of N service options, different text contents are used to indicate different voice services, where N is a positive integer; The display module is used to display the N text contents corresponding to N controls, with different controls corresponding to different text contents, and the controls are used for users to select the voice service indicated by the corresponding text content.

7. The apparatus according to claim 6, characterized in that, The audio conversion module includes: The first conversion unit is used to convert the first part of the audio into text to obtain the first sub-text. The second conversion unit is configured to convert the second part of the audio into text to obtain the second sub-text when the first sub-text matches the service indication information of the preset target service option and the value of the preset flag variable is a first preset value, and update the value of the flag variable to the second preset value. The target service option is at least one service option in the intelligent voice interaction interface, the second part of the audio is a part of the audio received after the first part of the audio, and the target text includes the first sub-text and the second sub-text. The search unit is configured to stop converting the second part of the audio into text and locate the target text including the first subtext in the electronic device when the first subtext matches the service indication information of the preset target service option and the value of the flag variable is the second preset value.

8. The apparatus according to claim 6, characterized in that, The display module is specifically used for: The N text contents are displayed on the intelligent voice interaction interface corresponding to the N service options; or... A pop-up window is displayed in the display interface, and the pop-up window includes N controls, on which the N controls display the corresponding N text contents.

9. The apparatus according to claim 8, characterized in that, The display module is specifically used for: The N text contents are displayed in the N service options of the intelligent voice interaction interface, and each of the service options includes the service instruction information and the text contents; or, Update the service instruction information of the N service options in the intelligent voice interaction interface to the N text content.

10. The apparatus according to claim 9, characterized in that, The display module is also used for: When the service instruction information of the N service options in the intelligent voice interaction interface is updated to the N text content, the display position of the N service options is updated.

11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the display method as described in any one of claims 1-5.

12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the display method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Service initiation techniques

    CN101978390A

  • Visualization IVR realization method and mobile terminal

    CN105704106A