Interface prompting method and device based on voice interaction, equipment and storage medium
Patent Information
- Application Number
- CN202211742881.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-31
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2042-12-31
AI Technical Summary
[0003]然而,当显示界面上的显示内容包括图标或者特殊字符时,会存在用户不知道该图标或者特殊字符如何用语音进行表达的问题,导致用户与终端之间语音交互困难
[0025] The above solutions may not be suitable for users who are unsure how to describe special target controls, such as graphical controls and controls with special characters, using voice commands. By providing prompts for the text descriptions of some target controls at different times during the display process, users are guided to use text that the terminal can recognize to describe the target controls during voice interaction. This makes it easier for users to describe these difficult-to-express target controls, improving the effectiveness of voice interaction between users and the terminal to some extent.
Smart Images

Figure CN116126442B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, and in particular to a method, apparatus, device, and storage medium for providing interface prompts based on voice interaction. Background Technology
[0002] With the continuous development of intelligent voice technology, the "see it, say it" function is being rapidly adopted on terminals. This means that users can directly speak voice control commands based on the content displayed on the terminal's screen to achieve voice interaction with the terminal.
[0003] However, when the displayed content includes icons or special characters, users may not know how to express those icons or special characters using voice, making voice interaction between users and the terminal difficult. Summary of the Invention
[0004] The main technical problem solved by this invention is to provide a method, apparatus, device and computer-readable storage medium for providing interface prompts based on voice interaction, which can improve the effect of voice interaction between users and terminals.
[0005] To address the aforementioned technical problems, this application provides a technical solution: a voice-interactive interface prompting method, comprising: determining at least one target control in a display interface, wherein the at least one target control includes at least one of a graphical control and a special character control; obtaining text description content corresponding to the at least one target control; and prompting the text description content corresponding to some target controls at different time periods during the display process, wherein the text description content corresponding to the target controls is used to guide the user to describe the target controls using text that the terminal can recognize during voice interaction with the terminal.
[0006] The steps for filtering target controls for each time period include: obtaining the number of each target control; and filtering out the target controls corresponding to each time period from at least one target control based on the number of each target control.
[0007] Specifically, based on the number of each target control, a subset of target controls corresponding to each time period are selected from at least one target control, including: extracting target numbers from the numbers of at least one target control according to the sampling rules matching each time period; and determining the target control corresponding to the target number as a subset of target controls corresponding to the time period.
[0008] The sampling rules include: matching the remainder of the number divided by the total number of time periods with the order of the time periods, or any of the random rules.
[0009] The prompts can be displayed on the screen or given via voice announcement, or at least one of these methods.
[0010] The prompts can be displayed on the display interface, with the text description corresponding to the target control displayed in the associated area of the target control or within the display graphic corresponding to the target control. The associated area is the area in the display interface that is outside the display graphic corresponding to the target control but within a preset distance range of the display graphic corresponding to the target control.
[0011] The determination of at least one target control in the display interface includes: determining at least one target control in the display interface in response to the activation of the voice interaction function, or determining at least one target control in the display interface in response to the activation of the voice interaction function and the activation of the control guidance service.
[0012] The determination of at least one target control in the display interface includes one of the following methods: reading the attribute information of each control in the display interface, identifying controls with special markers in the attribute information as target controls; obtaining guidance requirement reference information for each control in the display interface, predicting the accuracy probability of each control's description based on the guidance requirement reference information, and selecting the target control from the controls based on the accuracy probability, wherein the guidance requirement reference information includes at least one of the following: the display content on the display interface within a preset range of each control, the historical description of each control, and the user characteristics of the current user, wherein the historical description of the control represents the frequency with which the control was correctly described in historical voice interactions; and / or, obtaining the text description content corresponding to at least one target control, including: obtaining the text description content corresponding to each target control from the attribute information of each target control.
[0013] The method further includes, before prompting the text description content corresponding to each target control, the following steps: in response to the number of prompts for the text description content corresponding to the target control exceeding a first preset number, not prompting the text description content corresponding to the target control; and / or, after prompting the text description content corresponding to each target control, the method further includes: in response to one of the following conditions being met, stopping the prompting of the text description content corresponding to the target control: the voice interaction function is turned off; the prompting time for the text description content corresponding to the target control exceeds a preset time; the number of voice hits for the target control exceeds a second preset number.
[0014] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a voice-interactive interface prompting device, the device comprising: a determining module, configured to determine at least one target control in a display interface, the at least one target control including at least one of a graphical control and a special character control; an acquiring module, configured to acquire text description content corresponding to at least one target control; and a prompting module, configured to prompt the text description content corresponding to some target controls at different times during the display process, the text description content corresponding to the target controls being used to guide the user to describe the target controls using text recognizable by the terminal during voice interaction with the terminal.
[0015] The prompting module is used to obtain the number of each target control; based on the number of each target control, it filters out some target controls corresponding to each time period from at least one target control.
[0016] The prompting module is used to extract target numbers from the numbers of at least one target control according to the sampling rules that match each time period; and to determine the target control corresponding to the target number as part of the target controls corresponding to the time period.
[0017] The sampling rules include: matching the remainder of the number divided by the total number of time periods with the order of the time periods, or any of the random rules.
[0018] The prompts can be displayed on the screen or given via voice announcement, or at least one of these methods.
[0019] The prompts can be displayed on the display interface, with the text description corresponding to the target control displayed in the associated area of the target control or within the display graphic corresponding to the target control. The associated area is the area in the display interface that is outside the display graphic corresponding to the target control but within a preset distance range of the display graphic corresponding to the target control.
[0020] The determination module is used to determine at least one target control in the display interface in response to the activation of the voice interaction function, or to determine at least one target control in the display interface in response to the activation of the voice interaction function and the activation of the control guidance service.
[0021] The determining module is used to determine at least one target control in the display interface using one of the following methods: reading the attribute information of each control in the display interface, identifying controls with special markers in the attribute information as target controls; obtaining guidance requirement reference information for each control in the display interface, predicting the accuracy probability of each control's description based on the guidance requirement reference information, and selecting the target control from the controls based on the accuracy probability, wherein the guidance requirement reference information includes at least one of the following: the display content on the display interface within a preset range of each control, the historical description of each control, and the user characteristics of the current user, wherein the historical description of the control represents the frequency with which the control was correctly described in historical voice interactions; and / or, the obtaining module is used to obtain the text description content corresponding to each target control from the attribute information of each target control.
[0022] The prompting module is further configured to, before prompting the text description content corresponding to each target control, not prompt the text description content corresponding to the target control if the number of prompts for the text description content corresponding to the target control exceeds a first preset number; and / or, the prompting module is further configured to, after prompting the text description content corresponding to each target control, stop prompting the text description content corresponding to the target control if one of the following conditions is met: the voice interaction function is turned off; the prompting time for the text description content corresponding to the target control exceeds a preset time; the number of voice hits for the target control exceeds a second preset number.
[0023] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide a voice interaction device, including a memory and a processor coupled to each other, wherein the memory stores program instructions; and the processor is used to execute the program instructions stored in the memory to implement the above-mentioned interface prompting method.
[0024] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium for storing program instructions that can be executed to implement the above-mentioned interface prompting method.
[0025] The above solutions may not be suitable for users who are unsure how to describe special target controls, such as graphical controls and controls with special characters, using voice commands. By providing prompts for the text descriptions of some target controls at different times during the display process, users are guided to use text that the terminal can recognize to describe the target controls during voice interaction. This makes it easier for users to describe these difficult-to-express target controls, improving the effectiveness of voice interaction between users and the terminal to some extent. Attached Figure Description
[0026] Figure 1This is a flowchart illustrating an embodiment of the voice-interaction-based interface prompting method provided in this application;
[0027] Figure 2 This is a schematic diagram of the target control provided in this application;
[0028] Figure 3 This is a schematic diagram of the display interface provided in this application;
[0029] Figure 4 This is a schematic diagram of the framework of an embodiment of the voice-interactive interface prompting device provided in this application;
[0030] Figure 5 This is a schematic diagram of the framework of an embodiment of the voice interaction device provided in this application;
[0031] Figure 6 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation
[0032] To make the purpose, technical solution and effects of this application clearer and more explicit, the following describes this application in further detail with reference to the accompanying drawings and embodiments.
[0033] It should be noted that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this article means two or more. Moreover, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0034] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0035] Please see Figure 1 , Figure 1This is a flowchart illustrating an embodiment of the voice-interaction-based interface prompting method provided in this application. This method can be executed by any terminal with voice interaction and display functions; for example, the terminal can be an in-vehicle device, mobile phone, tablet, computer, or wearable device. It should be noted that if substantially the same result is achieved, the method of this invention is not necessarily identical. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown, the method includes the following steps:
[0036] S101: Determine at least one target control in the display interface, wherein the at least one target control includes at least one of a graphical control and a special character control.
[0037] A display interface includes multiple controls, and at least one target control refers to a specific control among the multiple controls on the display interface. For these specific controls, the user may not know how to express them verbally, or the user may have difficulty expressing them verbally. It should be noted that in this embodiment, when the terminal is running a software application, it will include multiple display interfaces in different states; the display interface in this embodiment is one of these multiple display interfaces in different states. Different display interfaces correspond to the same target control, or different display interfaces have at least some different target controls.
[0038] In this embodiment, at least one target control includes at least one of a graphical control and a special character control. For example, at least one target control includes a graphical control, or at least one target control includes a special character control, or at least one target control includes both a graphical control and a special character control. Figure 2 This is a schematic diagram of the target control provided in this application, such as... Figure 2 As shown, the controls corresponding to labels 1 and 2 are graphical controls; the controls corresponding to labels 3 and 4 are special character controls.
[0039] In one embodiment, the terminal's storage unit pre-stores a first correspondence between display interfaces and target controls, wherein different display interfaces correspond to the same target control or at least partially different target controls. Based on the current display interface and the stored first correspondence, the terminal can directly determine at least one target control corresponding to the current display interface.
[0040] In another embodiment, the terminal's storage unit stores multiple controls corresponding to each display interface and their attribute information. The attribute information of each control can be defined at the code level based on the control description rules of the terminal system. For example, the terminal system can be an Android system, etc. To facilitate the identification of the target control from the multiple controls on the display interface, the target control of the display interface can be determined based on whether the attribute information of each control in the display interface contains a special identifier. Specifically, the attribute information of each control on the display interface is read; the control whose attribute information contains a special identifier is identified as the target control. For example, the special identifier can be in the form of numbers, characters, or strings, etc., and this embodiment does not specifically limit this.
[0041] In another embodiment, to more accurately determine the target control requiring prompts, the target control on the display interface can be determined based on the accuracy probability of the descriptions of each control. Specifically, during voice interaction with the user, the terminal acquires guidance requirement reference information for each control on the display interface, predicts the accuracy probability of the descriptions of each control based on this reference information, and selects the target control from among the controls based on the accuracy probability. The guidance requirement reference information includes at least one of the following: the content displayed on the display interface within a preset range for each control, the historical descriptions of each control, and the user characteristics of the current user. The historical descriptions of a control represent the frequency with which the control was correctly described in historical voice interactions. User characteristics include the user's age, skin color, gender, etc. For example, controls with an accuracy probability less than a probability threshold are selected as target controls. When the accuracy probability of a control is less than the probability threshold, it indicates that the user has difficulty accurately describing the control; when the accuracy probability of a control is greater than or equal to the probability threshold, it indicates that the user can accurately describe the control. The probability threshold can be set according to actual conditions.
[0042] Optionally, in this embodiment, when the voice interaction function is activated, at least one control in the display interface is determined. Alternatively, when the voice interaction function is activated and the control guidance service is activated, at least one target control in the display interface is determined. The control guidance service refers to the terminal's accessibility service. When the terminal activates the voice interaction function or activates both the voice interaction function and the control guidance service, one of the above three implementation methods can be used to determine at least one target control in the display interface. In one example, when the terminal receives a voice interaction function activation command, the voice interaction function is activated. When the terminal receives a control guidance service activation command, the control guidance service is activated. Exemplarily, the voice interaction function activation command and the control guidance service activation command are triggered by the user clicking a virtual button in the terminal's display interface. Alternatively, the voice interaction function activation command and the control guidance service activation command are triggered by the user clicking a hardware button on the terminal.
[0043] S102: Obtain the text description content corresponding to at least one target control.
[0044] In one embodiment, the terminal pre-stores a second correspondence between target controls and text descriptions, where different target controls correspond to different text descriptions. After identifying at least one target control in the display interface, the terminal can determine the text description corresponding to each target control based on each target control and the second correspondence.
[0045] In another embodiment, the attribute information of each control on the display interface also includes the text description content corresponding to each control. After the terminal determines at least one target control in the display interface, it can read the attribute information of each target control and obtain the text description content corresponding to each target control from the attribute information of each target control.
[0046] S103: During different time periods when the display interface is displayed, prompts are given for the text descriptions of some target controls. The text descriptions of the target controls are used to guide the user to use text that the terminal can recognize to describe the target controls during voice interaction with the terminal.
[0047] In this embodiment, the display interface's dwell time includes at least one time period, and the duration of each time period may be the same or different. The target controls corresponding to each time period may be the same or different. During each time period of the display interface's dwell time, only the text description content corresponding to a portion of the aforementioned at least one target control is displayed.
[0048] In one implementation, the target controls corresponding to each time period can be predicted based on the user's current voice interaction content during the voice interaction process. Specifically, based on the user's current voice interaction content during the voice interaction process, the target controls that the user may use in the current time period are predicted. The number of target controls that the user may use in the current time period can be one or more. For example, if the terminal is running a music application and the user's current voice interaction content includes playing music, then the target controls that the user may use in the current time period include comment controls, karaoke controls, etc.
[0049] In another implementation, the number of each target control can be obtained first, and then, based on the number of each target control, a portion of the target controls corresponding to each time period can be selected from at least one target control.
[0050] For example, the attribute information of the target control includes the target control's number, which can be obtained from the attribute information of each target control. The number of each target control can be in numeric or character form, etc., and this embodiment does not specifically limit this.
[0051] Specifically, based on the number of each target control, a subset of target controls corresponding to a specific time period is selected from at least one target control. This includes: extracting target numbers from the numbers of at least one target control according to sampling rules matching each time period; and identifying the target controls corresponding to the target numbers as the subset of target controls corresponding to the time period. The sampling rules matching different time periods may be the same or different, and the sampling rules matching different time periods can be set according to actual needs.
[0052] In one example, the sampling rule includes an odd-number rule, which means that an odd-numbered number is extracted from the numbers of at least one target control as the target number. For example, if the numbers of at least one target control include 1, 2, 3, 4, and 5, the odd-numbered numbers 1, 3, and 5 are used as the target numbers.
[0053] In another example, the sampling rule includes an even-number rule, which means that even-numbered numbers are extracted from the numbers of at least one target control as the target number. For example, if the numbers of at least one target control include 1, 2, 3, 4, and 5, then even-numbered numbers 2 and 4 are used as the target numbers.
[0054] In another example, the sampling rule includes matching the remainder of the number divided by the total number of time periods with the ordinal position of each time period. For example, if the ordinal positions of each time period are 1, 2, ..., n, for the time period with ordinal position 1, the number with a remainder of 0 when divided by the total number of time periods is selected as the target number; for the time period with ordinal position 2, the number with a remainder of 1 when divided by the total number of time periods is selected as the target number; and for the time period with ordinal position n, the number with a remainder of n-1 when divided by the total number of time periods is selected as the target number.
[0055] In another example, the sampling rule includes a random rule, which involves randomly selecting a portion of the numbers from at least one target control as the target number. For example, a random function is used to randomly select a portion of the numbers from at least one target control as the target number.
[0056] In this embodiment, the prompting method includes at least one of displaying on the display interface and voice broadcasting. For example, displaying the text description content corresponding to the target control on the display interface; or, broadcasting the text description content corresponding to the target control via voice; or, displaying a portion of the text description content corresponding to the target control on the display interface and broadcasting another portion of the text description content corresponding to the target control via voice. Voice broadcasting prompts refer to the terminal converting the text description content corresponding to the target control into speech and outputting it.
[0057] In one embodiment, the prompting method includes voice broadcasting, that is, at different times during the display process, voice broadcasting is used to provide prompts for the text descriptions corresponding to some target controls. In this way, users can use the voice broadcasts of the text descriptions corresponding to each target control to provide voice descriptions for those target controls that are not easy to express, thereby improving the effect of voice interaction between users and the terminal.
[0058] Optionally, when the text descriptions corresponding to each target control are read aloud, the target control being read aloud is highlighted to distinguish it from other target controls. For example, the target control being read aloud can be displayed in a set color, which can be set according to actual needs. Alternatively, the graphic of the target control being read aloud can be enlarged or reduced. By highlighting the target control being read aloud, users can easily and intuitively identify the specific target control for the current voice prompt, thereby further improving the effectiveness of the voice prompt.
[0059] In another embodiment, the prompts are displayed on the display interface, specifically, at different times during the display process, showing text descriptions corresponding to a portion of the target control. To facilitate user identification of the specific target control corresponding to each text description, the text description is displayed within the associated area of the target control or within the corresponding display graphic. The associated area is defined as the region on the display interface outside the display graphic of the target control but within a preset distance range of the graphic. For example, the text description might be displayed within a preset distance range above the target control. The preset distance range is set according to actual needs. Figure 3 This is a schematic diagram of the display interface provided in this application, such as... Figure 3 As shown, the display interface includes target control a, target control b, and target control c. The text description for target control a is the original text, the text description for target control b is a list, and the text description for target control c is the anchor. The corresponding text description is displayed above the graphics of each of the target controls.
[0060] When there are many target controls on the display interface, simultaneously displaying the text descriptions of all target controls would make the interface appear crowded, cluttered, and monotonous. In this embodiment, the text descriptions of a subset of the target controls are displayed on the interface at different times during the display process. That is, only a portion of the text descriptions of the target controls are displayed at each time period. On the one hand, displaying only a subset of the target controls at each time period makes the display interface cleaner, and it makes it easier for users to obtain the text descriptions of each target control for voice interaction. On the other hand, when the subset of target controls differs at different times, displaying the text descriptions of a subset of target controls at different times allows for dynamic display of the text descriptions of each target control, thereby improving the display's prompting effect.
[0061] Optionally, in this embodiment, when the text description content corresponding to the target control meets certain conditions, it is assumed that the user already knows how to describe the target control after being prompted, and therefore the prompting of the text description content corresponding to the target control is stopped or not prompted. For example, when the terminal prompts the text description content corresponding to each target control through voice broadcast, the voice broadcasting of the text description content corresponding to the target control is not performed or is stopped. When the terminal prompts the text description content corresponding to each target control through display on the display interface, the text description content corresponding to the target control is not displayed or is hidden. When the text description content corresponding to the target control does not meet the conditions, it is assumed that the user may not yet know how to describe the target control, and the prompting of the text description content corresponding to the target control is still required.
[0062] In one embodiment, the terminal can record the number of times the text description content corresponding to each target control is prompted. Before prompting the text description content corresponding to each target control, if the number of prompts for the text description content corresponding to the target control is greater than a first preset number, the text description content corresponding to the target control is not prompted. Conversely, if the number of prompts for the text description content corresponding to the target control is less than or equal to the first preset number, the text description content corresponding to the target control is prompted. The first preset number is set according to the actual situation, for example, the first preset number is 5 times, 10 times, etc., and this embodiment does not specifically limit it.
[0063] In another embodiment, the terminal can also record the prompting time for the text description content corresponding to each target control. After prompting the text description content corresponding to each target control, if the prompting time for the text description content of the target control is greater than a set time, the prompting of the text description content corresponding to the target control stops. Conversely, if the prompting time for the text description content corresponding to the target control is less than or equal to the set time, the prompting of the text description content corresponding to the target control continues. The set time is set according to the actual situation, for example, the set time is 3 seconds, 5 seconds, etc., and this embodiment does not specifically limit it.
[0064] In another embodiment, the terminal can also record the number of voice hits for each target control. The number of voice hits for a target control refers to the number of times the user describes the target control using text that the terminal can recognize. After prompting the text description content corresponding to each target control, if the number of voice hits for a target control exceeds a second preset number, the prompting of the text description content corresponding to the target control stops. Conversely, if the number of voice hits for a target control is less than or equal to the second preset number, the prompting of the text description content corresponding to the target control continues. The second preset number is set according to actual conditions, for example, the second preset number is 3 times, 5 times, etc., and this embodiment does not specifically limit it.
[0065] In another embodiment, after the voice interaction function is activated and prompts the text description content corresponding to each target control, when the voice interaction function is closed, the prompting of the text description content corresponding to the target control is stopped.
[0066] In this embodiment, users may not know how to describe special target controls such as graphical controls and special character controls using voice. By providing prompts for the text descriptions of some target controls at different times during the display process, users are guided to describe the target controls using text that the terminal can recognize during voice interaction. This makes it easier for users to describe these difficult-to-express target controls using voice, thus improving the effectiveness of voice interaction between users and the terminal to some extent.
[0067] Please see Figure 4 , Figure 4 This is a schematic diagram of a framework of an embodiment of the voice-interactive interface prompting device provided in this application. In this embodiment, the voice-interactive interface prompting device 40 includes: a determining module 41, an acquiring module 42, and a prompting module 43.
[0068] The determining module 41 is used to determine at least one target control in the display interface, wherein the at least one target control includes at least one of a graphical control and a special character control. The obtaining module 42 is used to obtain the text description content corresponding to at least one target control. The prompting module 43 is used to provide prompts for the text description content corresponding to some target controls at different times during the display process, wherein the text description content corresponding to the target controls is used to guide the user to describe the target controls using text that the terminal can recognize during voice interaction with the terminal.
[0069] Optionally, the prompting module 43 is used to obtain the number of each target control; based on the number of each target control, it filters out some target controls corresponding to each time period from at least one target control.
[0070] Optionally, the prompting module 43 is used to extract target numbers from the numbers of at least one target control according to the sampling rules that match each time period; and to determine the target control corresponding to the target number as part of the target controls corresponding to the time period.
[0071] Optionally, the sampling rules include: matching the remainder of the number divided by the total number of time periods with the order of the time periods, or any of the random rules.
[0072] Optionally, the prompts may be displayed on a screen or given via voice.
[0073] Optionally, the prompt may be displayed on the display interface, wherein the text description content corresponding to the target control is displayed in the associated area of the target control or within the display graphic corresponding to the target control. The associated area is an area in the display interface that is outside the display graphic corresponding to the target control but within a preset distance range of the display graphic corresponding to the target control.
[0074] Optionally, the determining module 41 is used to determine at least one target control in the display interface in response to the activation of the voice interaction function, or to determine at least one target control in the display interface in response to the activation of the voice interaction function and the activation of the control guidance service.
[0075] Optionally, the determining module 41 is used to determine at least one target control in the display interface using one of the following methods: reading the attribute information of each control in the display interface, and determining the control containing a special mark in the attribute information as the target control; obtaining the guidance requirement reference information of each control in the display interface, predicting the accuracy probability of the description of each control based on the guidance requirement reference information of each control, and selecting the target control from each control based on the accuracy probability of the description, wherein the guidance requirement reference information includes at least one of the display content on the display interface within a preset range of each control, the historical description of each control, and the user characteristics of the current user, and the historical description of the control represents the frequency with which the control was correctly described in historical voice interactions; and / or, the obtaining module 42 is used to obtain the text description content corresponding to each target control from the attribute information of each target control.
[0076] Optionally, the prompting module 43 is further configured to, before prompting the text description content corresponding to each target control, not prompt the text description content corresponding to the target control if the number of prompts for the text description content corresponding to the target control exceeds a first preset number; and / or, the prompting module 43 is further configured to, after prompting the text description content corresponding to each target control, stop prompting the text description content corresponding to the target control if one of the following conditions is met: the voice interaction function is turned off; the prompting time for the text description content corresponding to the target control exceeds a preset time; the number of voice hits for the target control exceeds a second preset number.
[0077] It should be noted that the apparatus of this embodiment can perform the steps in the above method. For detailed descriptions of the relevant content, please refer to the method section above, which will not be repeated here.
[0078] Please see Figure 5 , Figure 5 This is a schematic diagram of a framework of an embodiment of the voice interaction device provided in this application. In this embodiment, the processing device 50 includes a memory 51 and a processor 52.
[0079] Processor 52 can also be referred to as CPU (Central Processing Unit). Processor 52 may be an integrated circuit chip with signal processing capabilities. Processor 52 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. A general-purpose processor can be a microprocessor, or processor 52 can be any conventional processor 52, etc.
[0080] The memory 51 in the processing device 50 is used to store the program instructions required for the processor 52 to run.
[0081] The processor 52 is used to execute program instructions to implement the voice-based interactive interface prompting method in this application.
[0082] Please see Figure 6 , Figure 6 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium provided in this application. The computer-readable storage medium 60 of this embodiment stores program instructions 61, which, when executed, implement the voice-interactive interface prompting method provided in this application. The program instructions 61 can be formed into a program file and stored in the aforementioned computer-readable storage medium 60 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) can execute all or part of the steps of the methods of various embodiments of this application. The aforementioned computer-readable storage medium 60 includes various media capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or terminal devices such as computers, servers, mobile phones, and tablets.
[0083] The above solutions may not be suitable for users who are unsure how to describe special target controls, such as graphical controls and controls with special characters, using voice commands. By providing prompts for the text descriptions of some target controls at different times during the display process, users are guided to use text that the terminal can recognize to describe the target controls during voice interaction. This makes it easier for users to describe these difficult-to-express target controls, improving the effectiveness of voice interaction between users and the terminal to some extent.
[0084] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0085] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0086] In the several embodiments provided in this application, it should be understood that the disclosed methods, apparatuses, and systems can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of apparatuses or units may be electrical, mechanical, or other forms.
[0087] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0088] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0089] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0090] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for prompting an interface based on voice interaction, characterized by, The method includes: Identify multiple target controls in the display interface, wherein the multiple target controls include at least one of graphical controls and special character controls; Obtain the text description content corresponding to the multiple target controls; During different time periods when the display interface is displayed, prompts are given for the text description content corresponding to some of the target controls. The text description content corresponding to the target controls is used to guide the user to use text that the terminal can recognize to describe the target controls during voice interaction with the terminal. The filtering steps for the target controls corresponding to each of the time periods include: obtaining the number of each target control, and filtering out the target controls corresponding to each time period from the plurality of target controls based on the number of each target control.
2. The method according to claim 1, characterized in that, The step of filtering out a portion of the target controls corresponding to each time period from the plurality of target controls based on their numbers includes: The target number is extracted from the number of the multiple target controls according to the sampling rules that match each of the time periods; The target control corresponding to the target number is determined as a portion of the target controls corresponding to the time period.
3. The method according to claim 2, characterized in that, The sampling rules include: the remainder of the number divided by the total number of time periods matching the ordinal position of the time period, or any of the following random rules.
4. The method according to claim 1, characterized in that, The prompting method includes at least one of displaying on the display interface and voice broadcasting.
5. The method according to claim 4, characterized in that, The prompt is displayed on the display interface, and the text description content corresponding to the target control is displayed in the associated area of the target control or within the display graphic corresponding to the target control. The associated area is an area in the display interface that is outside the display graphic corresponding to the target control and within a preset distance range of the display graphic corresponding to the target control.
6. The method according to claim 1, characterized in that, The determination of multiple target controls in the display interface includes: In response to the activation of the voice interaction function, the plurality of target controls in the display interface are determined; or, in response to the activation of the voice interaction function and the activation of the control guidance service, the plurality of target controls in the display interface are determined.
7. The method according to claim 1, characterized in that, The determination of multiple target controls in the display interface includes one of the following methods: Read the attribute information of each control on the display interface, and identify the control whose attribute information contains a special mark as the target control; The system obtains guidance requirement reference information for each control on the display interface, predicts the accuracy probability of the description of each control based on the guidance requirement reference information, and selects the target control from the controls based on the accuracy probability of the description. The guidance requirement reference information includes at least one of the following: the display content on the display interface within a preset range of each control, the historical description of each control, and the user characteristics of the current user. The historical description of the control indicates the frequency with which the control was correctly described in historical voice interactions. And / or, obtaining the text description content corresponding to the plurality of target controls includes: Obtain the text description content corresponding to each target control from the attribute information of each target control.
8. The method according to claim 1, characterized in that, Before prompting the text description content corresponding to some of the target controls, the method further includes: in response to the number of prompts for the text description content corresponding to the target control being greater than a first set number, not prompting the text description content corresponding to the target control; And / or, after prompting the text description content corresponding to each of the target controls, the method further includes: stopping the prompting of the text description content corresponding to the target control in response to one of the following conditions: the voice interaction function is turned off; the prompting time of the text description content corresponding to the target control is greater than a set time; the number of voice hits of the target control is greater than a second set number.
9. A voice-interactive interface prompting device, characterized in that, The device includes: A determination module is used to determine multiple target controls in the display interface, wherein the multiple target controls include at least one of graphical controls and special character controls; The acquisition module is used to acquire the text description content corresponding to the multiple target controls; The prompting module is used to provide prompts for the text descriptions corresponding to some of the target controls at different time periods during the display process on the display interface. The text descriptions corresponding to the target controls are used to guide the user to describe the target controls using text that the terminal can recognize during voice interaction with the terminal. The prompting module is used to obtain the number of each target control and, based on the number of each target control, to filter out some of the target controls corresponding to each time period from the plurality of target controls.
10. A voice interaction device, characterized in that, Including interconnected memory and processor, The memory stores program instructions; The processor is used to execute program instructions stored in the memory to implement the method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program instructions that can be executed to implement the method of any one of claims 1-8.
Citation Information
Patent Citations
Human-computer interaction, control and live broadcast method and device, and storage medium
CN113301361A
Voice control method, device and equipment and computer storage medium
CN114067797A