"say what you see" implementation method and apparatus, and vehicle
By receiving user voice commands and using different solutions to obtain interface information and execution instructions according to the application type of the display interface, the problem of visible and single processing flow in the prior art is solved, and the interface information acquisition method and instruction execution method are diversified, which improves user experience and driving safety.
Patent Information
- Application Number
- PCT/CN2025/071280
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-17
AI Technical Summary
The visible and sensible processing process in the prior art is too single, resulting in high development costs and inability to effectively control the application interface, and the inability to achieve flexible visible and sensible interactions.
By receiving user voice commands, obtain the application information type of the display interface, and use application-level or system-level solutions to obtain interface information and execute instructions, including using preset whitelists, software development toolkits and accessibility assistance services to match and execute operations.
It has realized the diversification of interface information acquisition methods and instruction execution methods, flexibly selected technical solutions, improved the visible and talkable business implementation and expansion capabilities, and enhanced user experience and driving safety.
Smart Images

Figure CN2025071280_17072025_PF_FP_ABST
Abstract
Description
A visible and audible implementation method, device and vehicle
[0001] This application claims priority to the Chinese patent disclosure with application number 202410046952.0 and application name “A visible and tangible implementation method, device and vehicle” filed with the China Patent Office on January 11, 2024, the entire contents of which are incorporated into this application by reference. Technical Field
[0002] The present application relates to, but is not limited to, the field of speech recognition, and more specifically to a visible and audible implementation method, device, and vehicle. Background Art
[0003] With the rapid development of voice recognition technology and automotive intelligence, voice technology has played an important role in the automotive field in terms of smart travel and is an indispensable part of today's automotive field. From the initial voice navigation to today's "visible and speakable", "visible and speakable" provides a variety of new interactive methods including vehicle control, social interaction and entertainment. "Visible and speakable" means that users only need to speak the name of the control displayed on the screen interface through voice, without having to touch the screen, to achieve the purpose of using voice to control the interface controls. This allows the driver's attention to no longer be focused on various complicated settings and buttons, while ensuring driving safety and enhancing the user experience.
[0004] The visual and audible processing process in related technologies generally involves transmitting the application's interface information and hotwords to the voice client. The natural language understanding module then determines whether a visual and audible intent exists. If so, the command is executed to manipulate the target control. However, the solutions for implementing intelligent voice visual and audible control in related technologies are overly simplistic and require significant development costs, failing to effectively implement visual and audible control of applications. Technical Solutions
[0005] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0006] According to one aspect of the present application, a visible and speakable implementation method is provided, which includes: receiving a user's voice command, and obtaining semantic text information based on the voice command; obtaining the current display interface, and determining the interface information of the display interface based on the information type of the application corresponding to the display interface; and outputting an execution instruction based on the interface information and the semantic text information to perform the operation corresponding to the voice command.
[0007] In one embodiment of the present application, determining the interface information of the display interface based on the information type of the application corresponding to the display interface includes: determining whether the information type of the application corresponding to the display interface is a first-category application or a second-category application, wherein the first-category application is an application that is visible and speakable with the help of a preset software development kit, and the second-category application is an application that is visible and speakable with the help of a system service; in response to the information type of the application corresponding to the display interface being the first-category application, obtaining interface information from the display interface based on a first information acquisition method; in response to the information type of the application corresponding to the display interface being the second-category application, obtaining interface information from the display interface based on a second information acquisition method.
[0008] In one embodiment of the present application, outputting an execution instruction based on the interface information and the semantic text information to perform an operation corresponding to the voice instruction includes:
[0009] In response to the fact that the information type of the application corresponding to the display interface is the first type of application, the semantic text information is matched with the interface information, and a first execution instruction is output after the match is successful to execute the operation corresponding to the voice instruction; in response to the fact that the information type of the application corresponding to the display interface is the second type of application, the semantic text information is matched with the interface information, and a second execution instruction is output after the match is successful to execute the operation corresponding to the voice instruction.
[0010] In one embodiment of the present application, the method further includes: before determining whether the information type of the application corresponding to the display interface is a first-category application or a second-category application, determining whether the application corresponding to the display interface is an application in a preset whitelist; after determining that the application corresponding to the display interface is an application in the preset whitelist, determining whether the information type of the application corresponding to the display interface is the first-category application or the second-category application; wherein, the applications in the preset whitelist are applications that support visibility and speech.
[0011] In one embodiment of the present application, each application in the preset whitelist includes a label, and the label is used to identify the application as the first category application or the second category application; determining whether the information type of the application corresponding to the display interface is the first category application or the second category application includes: determining whether the information type of the application corresponding to the display interface is the first category application or the second category application based on the label for each application in the preset whitelist.
[0012] In one embodiment of the present application, the preset whitelist is deployed in the cloud or locally, and the preset whitelist is editable.
[0013] In one embodiment of the present application, the obtaining of interface information from the display interface based on the first information acquisition method includes: obtaining the interface information of the display interface through the software development kit of the application corresponding to the display interface; the outputting of the first execution instruction to execute the operation corresponding to the voice instruction includes: outputting an execution instruction to the application corresponding to the display interface, so that the application corresponding to the display interface executes the operation corresponding to the voice instruction.
[0014] In one embodiment of the present application, the acquiring of interface information from the display interface based on the second information acquisition method includes: acquiring the interface information of the display interface through the accessibility assistance service of the system service where the display interface is located; the outputting of the second execution instruction to execute the operation corresponding to the voice instruction includes: outputting an execution instruction to the accessibility assistance service of the system service where the display interface is located, so that the accessibility assistance service executes the operation corresponding to the voice instruction.
[0015] In one embodiment of the present application, the method further includes: in response to the display interface presenting interfaces of at least two applications, determining whether each of the at least two applications is the first category application or the second category application based on its information type.
[0016] According to another aspect of the present application, a device for implementing visible and sayable is provided, the device including a processor and a memory, the memory storing a computer program executed by the processor, and the computer program, when executed by the processor, causing the processor to execute the above-mentioned visible and sayable implementation method.
[0017] According to another aspect of the present application, a vehicle is provided, comprising the above-mentioned visible and tangible implementation device.
[0018] According to another aspect of the present application, a storage medium is provided, on which a computer program to be executed by a processor is stored. When the computer program is executed by the processor, the processor executes the above-mentioned visible and verbal implementation method.
[0019] The visible and talkable implementation method, device and vehicle of the present application adopt different schemes to obtain interface information and execute instructions of the display interface based on the information type of the application corresponding to the current display interface, thereby realizing the diversification of interface information acquisition methods and instruction execution methods. The corresponding technical scheme can be flexibly selected according to actual conditions, making the visible and talkable business implementation and expansion more flexible and no longer limited to a single scheme. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0021] FIG1 shows a schematic block diagram of an electronic device for implementing the visible and audible implementation method and apparatus according to an embodiment of the present application.
[0022] FIG2 shows a schematic flow chart of a method for implementing the visible and audible method according to an embodiment of the present application.
[0023] FIG3 is a schematic diagram showing an application-level and system-level combination solution of Visible and Speakable according to an embodiment of the present application.
[0024] FIG4 shows a schematic structural block diagram of a device for implementing visible and audible communication according to an embodiment of the present application.
[0025] Implementation Methods of the Application
[0026] In order to make the purpose, technical solutions and advantages of the present application more apparent, example embodiments according to the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application described in this application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this application.
[0027] First, an example electronic device 100 for implementing the visible and audible implementation method and apparatus according to an embodiment of the present application is described with reference to FIG. 1 .
[0028] As shown in FIG1 , electronic device 100 includes one or more processors 102, one or more storage devices 104, an input device 106, and an output device 108. These components are interconnected via a bus system 110 and / or other connection mechanisms (not shown). It should be noted that the components and structure of electronic device 100 shown in FIG1 are merely exemplary and non-limiting. The electronic device may also have other components and structures as needed.
[0029] The processor 102 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 100 to perform desired functions.
[0030] The storage device 104 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 102 may run the program instructions to implement the client functions and / or other desired functions in the embodiments of the present application (implemented by the processor) described below. Various applications and various data may also be stored in the computer-readable storage medium, such as various data used and / or generated by the application.
[0031] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, etc. In addition, the input device 106 may also be any interface for receiving information.
[0032] The output device 108 can output various information (such as images or sounds) to the outside (such as a user), and can include one or more of a display, a speaker, etc. In addition, the output device 108 can also be any other device with an output function.
[0033] For example, example electronic devices for implementing the visible and audible implementation method and apparatus according to the embodiments of the present application can be implemented in smart vehicle terminals, smart homes, smart education, and the like.
[0034] Below, the visible and talkable implementation method 200 according to an embodiment of the present application will be described with reference to FIG2. FIG2 shows a schematic flow chart of the visible and talkable implementation method 200 according to an embodiment of the present application. As shown in FIG2, the visible and talkable implementation method 200 according to an embodiment of the present application may include the following steps:
[0035] In step S210, a user's voice instruction is received, and semantic text information is obtained based on the voice instruction.
[0036] In step S220, the current display interface is acquired, and based on the information type of the application corresponding to the display interface, the interface information of the display interface is determined.
[0037] In step S230, an execution instruction is output based on the interface information and the semantic text information to perform an operation corresponding to the voice instruction.
[0038] In an embodiment of the present application, the user's voice command is first received. When the voice is activated, the current display interface is obtained, and the information type of the application corresponding to the display interface is determined by the obtained display interface. Different applications have different ways of obtaining interface information and executing instructions. In an embodiment of the present application, it can be determined whether the application is a first-class application or a second-class application. Among them, the first-class application refers to an application that is visible and speakable with the help of a preset software development kit (SDK), which is also referred to as an application-level application in this article, or an application that uses an application-level solution. The second-class application refers to an application that is visible and speakable with the help of system services, which is also referred to as a system-level application in this article, or an application that uses a system-level solution.
[0039] When it is determined that the application corresponding to the display interface is a first type of application, the application uses an application-level solution, and obtains the interface information of the display interface through a first information acquisition method (for example, through the software tool development kit of the application). When it is determined that the application corresponding to the display interface is a second type of application, the application uses a system-level solution, and obtains the interface information of the display interface through a second information acquisition method (for example, the accessibility assistance service of the system service where the display interface is located). In a further embodiment, for some applications, the same application can also use application-level and system-level solutions at the same time to obtain more complete interface information. Based on the acquired user's voice instructions, the voice instructions are semantically converted through automatic speech recognition to obtain semantic text information, and the semantic text information is matched with the interface information of the display interface. When the semantic text information matches the interface information of the display interface (that is, there is a visible and speakable intention), the execution instruction is output.
[0040] When it is determined that the application corresponding to the display interface is a first type of application, a first execution instruction is output. The first execution instruction may, for example, instruct the application corresponding to the display interface to perform an operation corresponding to the voice instruction. When it is determined that the application corresponding to the display interface is a second type of application, a second execution instruction is output. The second execution instruction may, for example, instruct the accessibility assistance service of the system service to perform an operation corresponding to the voice instruction.
[0041] Therefore, the visible and talkable implementation method 200 of the present application adopts different schemes to obtain interface information and execution instructions of the display interface based on the information type of the application corresponding to the current display interface, thereby realizing the diversification of interface information acquisition methods and instruction execution methods. The corresponding technical scheme can be flexibly selected according to actual conditions, making the visible and talkable business implementation and expansion more flexible and no longer limited to a single scheme.
[0042] In an embodiment of the present application, step S210 receives a user's voice command and obtains semantic text information based on the voice command. First, the user's voice command, such as "open sunroof," "play next song," and "open map navigation," can be collected through the voice terminal. The received voice command then undergoes front-end voice signal processing, which is then transmitted to an automatic speech recognition module. Finally, the automatic speech recognition module converts the processed voice command into semantic text information.
[0043] In this embodiment, front-end voice signal processing refers to the processing of the user's voice commands to ensure that the processed speech better represents the essential characteristics of the speech. Front-end voice signal processing consists of sub-functions such as endpoint detection, noise suppression, wake-up, and echo cancellation. These sub-functions accurately determine the starting point of the speech signal and then detect valid speech segments from the continuous speech stream. Noise suppression, also known as noise reduction, is crucial because the captured audio typically contains a certain intensity of background sound. High background noise levels can significantly impact the performance of voice applications, such as reducing speech recognition rates and endpoint detection sensitivity. Therefore, noise suppression is essential in front-end voice signal processing. The key to noise suppression is extracting the noise spectrum and then performing an inverse compensation operation on the noisy speech based on the noise spectrum to produce the noise-reduced speech. Wake-up, for example, involves waking up with the wake-up word "Hello Xiaodi," which recognizes the wake-up word and begins collecting and processing audio signal data. Echo cancellation uses algorithms to remove the sound generated by the device itself, which is recorded by the device microphone. For example, if you say "Hello Xiaodi" to wake up the device, it will respond with "Hello" using text-to-speech (TTS) technology. The device will then capture its own "Hello" through the microphone. If the device's own echo cancellation isn't effective, it might mistakenly think the user is speaking "Hello" to it. Upon receiving this fake "Hello," the machine will politely reply "Hello" according to its design. In this way, the device will continue to speak "Hello" to itself. Therefore, front-end voice signal processing can make the signal generated by subsequent voice processing more uniform and smooth, reducing the difficulty of back-end voice recognition or voice wake-up, and improving voice processing quality.
[0044] In this embodiment, the voice instructions after front-end signal processing need to be converted into semantic text information through voice recognition. The voice instructions can be converted through speech recognition technology (Automatic Speech Recognition, ASR), also known as automatic speech recognition, or through other suitable speech recognition methods for semantic text conversion, which is not specifically limited. Automatic speech recognition can convert a natural speech signal into machine-recognizable text information. Because in actual situations, the user's voice instructions obtained by the machine cannot be recognized and read, they need to be converted into a machine-readable language. Therefore, automatic speech recognition is used to complete the conversion of voice instructions. Automatic speech recognition mainly includes the following steps: first, feature extraction is performed to extract feature vectors from the voice instructions after front-end signal processing. Feature vectors can reflect the rhythm, pitch, timbre and other characteristics of the speech; second, acoustic modeling is performed to model the feature vectors using the acoustic model, and the feature vectors are mapped to the phoneme level, and then to the word level; third, a language model is established to use the language model to impose language constraints on the text converted from the speech, so that the output text is more in line with language habits; then recognition is performed to compare the feature vectors processed by acoustic modeling and language model with a pre-trained vocabulary, and the most matching text is output; finally, post-processing is performed on the output text to perform grammar correction, punctuation processing and other post-processing operations to make it more in line with human language expression habits. Therefore, after automatic speech recognition processing, speech can be converted into text, the semantic information of speech can be understood, text-to-speech can be converted, and communication in application scenarios can be achieved, which facilitates communication between people and promotes human-computer communication.
[0045] In an embodiment of the present application, the current display interface is obtained in step S220, and based on the information type of the application corresponding to the display interface, the interface information of the display interface is determined, including: determining whether the information type of the application corresponding to the display interface is a first-category application or a second-category application, wherein the first-category application is an application that is visible and talkable with the help of a preset software development kit, and the second-category application is an application that is visible and talkable with the help of a system service. Before determining whether the information type of the application corresponding to the display interface is a first-category application or a second-category application, determine whether the application corresponding to the display interface is an application in a preset whitelist; after determining that the application corresponding to the display interface is an application in a preset whitelist, determine whether the application corresponding to the display interface is a first-category application or a second-category application; wherein the applications in the preset whitelist are applications that support visible and talkable functions. The main functions of the whitelist are: first, the preset whitelist can be used to determine which applications support visible and talkable functions; second, the whitelist can be used to determine whether the application in the current display interface is a first-category application or a second-category application, the first-category application is the application that uses an application-level solution, and the second-category application is the application that uses a system-level solution.
[0046] In this embodiment, the judgment rule for the preset whitelist setting is to determine whether the application of the current display interface is in the preset whitelist. If the application is in the preset whitelist, it can be determined that the application is a first-category application and uses an application-level solution; if the application is not in the preset whitelist, it can be determined that the application is a second-category application and uses a system-level solution. Of course, it is also possible to set the applications in the preset whitelist as second-category applications and use a system-level solution, and there is no specific limitation on this. For example, applications such as AutoNavi Map Car Edition and Fusion Music are in the preset whitelist. These applications have integrated Visible and Talkable software development kits. These applications are first-category applications and use an application-level solution. Applications such as QQ Music and NetEase Cloud Music are not in the preset whitelist and do not have integrated Visible and Talkable software development kits. These applications are second-category applications and use a system-level solution.
[0047] In an embodiment of the present application, each application in the preset whitelist includes a label, and the label is used to identify the application as a first-category application or a second-category application; determining whether the application corresponding to the display interface is a first-category application or a second-category application includes: determining whether the application corresponding to the display interface is a first-category application or a second-category application based on the label for each application in the preset whitelist. That is, the label of each application is recorded in the whitelist, and the application attributes corresponding to the current display interface are determined by the label. The labels in the whitelist are used to identify applications, which can more conveniently and quickly divide the corresponding applications in the current display interface, which is more conducive to the selection of subsequent application solutions.
[0048] In the embodiments of this application, the preset whitelist is deployed in the cloud or locally and is editable. The applications in the preset whitelist are those that support visibility and speaking. Another important function of this preset whitelist is that the applications in it can be changed at any time. The preset whitelist can be deployed in the cloud or locally, and applications can be deleted or added at any time, allowing for more flexible implementation of the needs of each application.
[0049] In an embodiment of the present application, when the information type of the application corresponding to the display interface is a first type of application, the interface information is obtained from the display interface based on the first information acquisition method. When the information type of the application corresponding to the display interface is a second type of application, the interface information is obtained from the display interface based on the second information acquisition method. For some applications, the same application can also use application-level and system-level solutions at the same time to obtain more complete interface information, which is not specifically limited. Before obtaining interface information from the display interface, it also includes generating hot words through the interface information of the current display interface to form a hot word list, and then transmitting the hot word list to automatic speech recognition for hot word enhancement.
[0050] Exemplarily, hot words are used to enhance recognition by providing some text content to automatic speech recognition. For example, some brand names such as BYD, Wenjie, Ideal, Weilai and Xiaopeng, if there is no hot word enhancement, it is easy for automatic speech recognition to make mistakes after the user says them. Therefore, it is necessary to provide these brand words to automatic speech recognition for recognition enhancement and improvement of recognition accuracy. The hot words in the embodiment of the present application refer to the interface information text in the current display interface, and the interface information text includes the text in the picture, the control name and the text corresponding to the icon, etc.
[0051] Exemplarily, generating hot words through the interface information of the current display interface to form a hot word list includes first extracting text information from the interface information of the current display interface to obtain a text information extraction result, then performing word segmentation processing based on the text information extraction result to generate hot words to form a hot word list, and then transmitting the hot word list to automatic speech recognition for hot word enhancement. Among them, the text information extraction result is subjected to word segmentation processing, and the result after word segmentation processing is a combination of word segmentation text and word segmentation text. For example, the name of a certain instruction is "eat grapes without spitting out the grape skins, do not eat grapes, spit out the grape skins", and the generalized expressions after word segmentation processing may include "eat grapes", "do not spit out grape skins", "do not eat grapes", and "spit out grape skins", etc. These expressions can all be used as hot words corresponding to the instruction operation, and then form a corresponding hot word list. Therefore, uploading the hot word list can improve the matching accuracy of speech recognition.
[0052] In an embodiment of the present application, obtaining interface information from a display interface based on a first information acquisition method includes: obtaining interface information of the display interface through the software development kit of the application corresponding to the display interface; and outputting a first execution instruction to perform an operation corresponding to the voice instruction, including: outputting an execution instruction to the application corresponding to the display interface, so that the application corresponding to the display interface performs the operation corresponding to the voice instruction. Based on the previously set application in the preset whitelist, if the application corresponding to the current display interface is in the preset whitelist, it is determined that the application corresponding to the current display interface is a first-class application. The first-class application is application-level, so the first information acquisition method is also an application-level solution. The application-level solution is to obtain interface information of the display interface through the software development kit of the application corresponding to the display interface. It can also be obtained through other joint development modes, such as the voice end and the application communicating and determining a certain information transmission method, which is not specifically limited. The application-level solution specifically obtains interface information of the display interface by first providing a software development kit on the voice end, and the application corresponding to the current display interface needs to integrate the software development kit provided by the voice end. The software development kit transmits the interface information of the current display interface to the voice end, and then outputs the corresponding execution instruction to perform the corresponding voice instruction operation.
[0053] In this embodiment, when the information type of the application corresponding to the display interface is the first type of application, outputting an execution instruction based on the interface information and the semantic text information in step S230 to perform the operation corresponding to the voice instruction. Matching the semantic text information with the interface information obtained based on the first information acquisition method, and outputting a first execution instruction to perform the operation corresponding to the voice instruction after a successful match, includes: outputting the execution instruction to the application corresponding to the display interface, so that the application corresponding to the display interface performs the operation corresponding to the voice instruction. Because the application corresponding to the current display interface is determined to be a first-class application, the first-class application is determined to be application-level through the previously preset whitelist. The application-level application executes the application-level solution. The application-level solution obtains the interface information of the display interface through the software development kit of the corresponding application, and then transmits the interface information of the display interface to the natural language understanding layer (NLU). Then, the semantic text information is transmitted to the natural language understanding layer. The two are matched in the natural language understanding layer to determine whether there is a visible and audible intent, that is, whether the user's voice matches the interface information in the current interface. When there is a visible and audible intent, that is, the user's voice matches the interface information in the current interface, the execution instruction is output to the instruction execution module.
[0054] For example, natural language understanding refers to the ability of a computer to understand human natural language, including technologies such as speech recognition, language models, syntactic analysis, and semantic analysis. Through natural language understanding, the visible and expressible intentions of the user's voice commands can be matched, and then execution instructions can be issued. Because the first type of application at this time is an application-level application and an application-level solution is used, the instruction module transmits the instruction to the application corresponding to the display interface, and the application corresponding to the display interface executes the instruction operation, that is, the application corresponding to the display interface executes the operation corresponding to the voice instruction. For example, if the user's voice command is "Open AutoNavi Map", the application will automatically open AutoNavi Map without the need to perform any click operations, which improves the user's driving safety and enhances the user experience.
[0055] In an embodiment of the present application, when the application corresponding to the display interface is a second type of application, interface information is obtained from the display interface based on the second information acquisition method, the semantic text information is matched with the interface information, and a second execution instruction is output after the match is successful to execute the operation corresponding to the voice instruction. Before acquiring the interface information from the display interface based on the second information acquisition method, it also includes generating hot words through the interface information of the current display interface to form a hot word list, and then transmitting the hot word list to the automatic speech recognition for hot word enhancement. Generating hot words through the interface information of the current display interface to form a hot word list includes first extracting text information from the interface information of the current display interface to obtain a text information extraction result, then performing word segmentation processing based on the text information extraction result to generate hot words to form a hot word list, and then transmitting the hot word list to the automatic speech recognition for hot word enhancement to improve the matching accuracy of speech recognition. Afterwards, the interface information of the current display interface is matched with the semantic text information converted from the previous user voice, and the second execution instruction is output to execute the corresponding instruction operation.
[0056] In an embodiment of the present application, interface information is obtained from the display interface based on the second information acquisition method, including: obtaining the interface information of the display interface through the accessibility assistance service of the system service where the display interface is located; outputting a second execution instruction to perform the operation corresponding to the voice instruction, including: outputting an execution instruction to the accessibility assistance service of the system service where the display interface is located, so that the accessibility assistance service performs the operation corresponding to the voice instruction. Through the settings of the applications in the previously preset whitelist, it can be determined that if the application corresponding to the current display interface is not in the preset whitelist, then the application corresponding to the current display interface is a second type of application. The second type of application is system-level, so the second information acquisition method is also a system-level solution. The system-level solution is to obtain the interface information of the display interface through the accessibility assistance service of the system service where the display interface is located. The accessibility assistance service obtains the interface information in the current display interface by scanning the current display interface, and then transmits the interface information of the display interface to the natural language understanding layer, and then transmits the semantic text information to the natural language understanding layer. The two are matched at the natural language understanding layer to determine whether there is a visible and audible intent, that is, whether the user's voice matches the interface information on the current interface. If there is a visible and audible intent, that is, if the user's voice matches the interface information on the current interface, an execution instruction is output to the instruction execution module. Because the application corresponding to the currently displayed interface is a second-category application, the instruction execution module uses the visible and audible system-level service to perform the operation, that is, uses the barrier-free auxiliary service to execute the corresponding voice instruction.
[0057] In an embodiment of the present application, the method further includes: when the display interface presents interfaces of at least two applications, determining whether each of the at least two applications is a first-category application or a second-category application based on its information type. If the same screen is split, first-category applications and second-category applications may coexist, that is, application-level solutions and system-level solutions may coexist. In this case, the user's voice command is first received, and the user's voice command is subjected to front-end signal processing, and then the voice after front-end signal processing is transmitted to automatic speech recognition to obtain semantic text information; secondly, the interface information of the display interface is obtained. Because there are both application-level solutions and system-level solutions, the corresponding methods of obtaining the interface information of the display interface are used at the same time, that is, the interface information of the display interface is obtained through the software development kit corresponding to the application-level solution and the accessibility assistance service of the system service corresponding to the system-level solution; then the obtained interface information of the display interface is transmitted to the natural language understanding layer, and the semantic text information obtained from the user's voice command is transmitted to the natural language understanding layer, and the two are matched to determine the visible and speakable intent. When the two match, that is, there is a visible and speakable intent, an execution instruction is issued; finally, the application-level solution has the application corresponding to the display interface execute the operation corresponding to the voice command, and the system-level solution has the accessibility assistance service execute the operation corresponding to the voice command.
[0058] The following summarizes the application-level and system-level combined solutions for Visible and Speakable according to an embodiment of the present application, with reference to Figure 3. As shown in Figure 3, the user activates voice and begins speaking the interface text currently displayed on the screen. When voice is activated, if the application currently displayed on the interface uses an application-level solution, the application transmits the interface information and text to the intelligent voice client app (i.e., the execution subject of method 200 described above) via the SDK or other communication method. When voice is activated, if the application currently displayed on the interface uses a system-level solution, the interface information and text are delivered to the intelligent voice client app through accessibility assistance. The information collection module transmits the collected interface text to the ASR service provider and also passes the interface information and text to the NLU semantic layer. When the semantic algorithm layer matches the Visible and Speakable intent, it issues an execution instruction to the instruction execution module. If the current application uses an application-level solution, the instruction execution module sends the instruction to the corresponding application. The click operation is performed by the third-party application. If the current application uses a system-level solution, the instruction execution module uses the Visible and Speakable system-level service to perform the operation, i.e., the accessibility assistance service performs the click operation.
[0059] Based on the above description, the implementation method of Visible and Sayable according to the embodiment of the present application adopts different schemes to obtain interface information and execute instructions of the display interface based on the information type of the application corresponding to the current display interface, thereby realizing the diversification of interface information acquisition methods and instruction execution methods. The corresponding technical scheme can be flexibly selected according to actual conditions, making the business implementation and expansion of Visible and Sayable more flexible and no longer limited to a single scheme.
[0060] The following describes a visible and talkable implementation device provided according to another aspect of the present application in conjunction with Figure 4. Figure 4 shows a schematic structural block diagram of a visible and talkable implementation device 400 according to an embodiment of the present application. As shown in Figure 4, the visible and talkable implementation device 400 includes a memory 410 and a processor 420, wherein: the memory 410 stores a computer executable program run by the processor 420, and when the computer executable program is run by the processor 420, the processor 420 executes the visible and talkable implementation method 200 described above. Those skilled in the art can understand the structure and specific operation of each module in the visible and talkable implementation device 400 according to the embodiment of the present application in combination with the content described above. For the sake of brevity, they will not be repeated here. Therefore, the visible and talkable implementation device according to the embodiment of the present application adopts different schemes to obtain interface information of the display interface and execute instructions based on the information type of the application corresponding to the current display interface, thereby realizing the diversification of interface information acquisition methods and instruction execution methods. The corresponding technical solution can be flexibly selected according to the actual situation, making the visible and talkable business implementation and expansion more flexible and no longer limited to a single solution.
[0061] In addition, according to an embodiment of the present application, a vehicle is also provided, which may include the visible and audible implementation device 400 described above.
[0062] In addition, the present application also provides a storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor executes the aforementioned visible and tangible implementation method according to the embodiment of the present application. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0063] Based on the above description, the visible and talkable implementation method, device and vehicle of the embodiments of the present application adopt different schemes to obtain interface information and execute instructions of the display interface based on the information type of the application corresponding to the current display interface, thereby realizing the diversification of interface information acquisition methods and instruction execution methods. The corresponding technical scheme can be flexibly selected according to actual conditions, making the visible and talkable business implementation and expansion more flexible and no longer limited to a single scheme.
[0064] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely illustrative and are not intended to limit the scope of the present application. Various changes and modifications may be made therein by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as required by the appended claims.
[0065] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0066] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units described is merely a logical function division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not performing some features.
[0067] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0068] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various disclosed aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this approach of the present application should not be interpreted as reflecting the intention that the application claimed for protection requires more features than those explicitly recited in each claim. More precisely, as reflected in the corresponding claims, it is that the corresponding technical problem can be solved with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present application.
[0069] It will be understood by those skilled in the art that, except where mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or apparatus disclosed herein may be combined in any combination. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0070] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.
[0071] The various component embodiments of the present application can be implemented in hardware, or in a software module running on one or more processors, or in a combination thereof. It will be appreciated by those skilled in the art that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules according to the embodiments of the present application. The application can also be implemented as a part or all of a program (for example, a computer program and a computer program product) for performing the method described herein. Such a program realizing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0072] It should be noted that the above embodiments illustrate rather than limit the present application, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim listing several anomaly detection devices for train traction systems, several of these anomaly detection devices for train traction systems may be embodied by the same hardware item. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0073] The above description is merely a specific embodiment or illustration of a specific embodiment of the present application, and the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. The scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A visible and speakable implementation method, characterized in that, The method includes: Receiving a voice command from a user and obtaining semantic text information based on the voice command; Obtaining a current display interface and determining interface information of the display interface based on the information type of the application corresponding to the display interface; Outputting an execution command based on the interface information and the semantic text information to perform an operation corresponding to the voice command.
2. The method according to claim 1, wherein The determining the interface information of the display interface based on the information type of the application corresponding to the display interface includes: Determining whether the information type of the application corresponding to the display interface is a first type of application or a second type of application, where the first type of application is an application that realizes visible and speakable functions by means of a preset software development kit, and the second type of application is an application that realizes visible and speakable functions by means of a system service; In response to the information type of the application corresponding to the display interface being the first type of application, obtaining interface information from the display interface based on a first information obtaining method; In response to the information type of the application corresponding to the display interface being the second type of application, obtaining interface information from the display interface based on a second information obtaining method.
3. The method according to claim 1 or 2, characterized in that, The outputting an execution command based on the interface information and the semantic text information to perform an operation corresponding to the voice command includes: In response to the information type of the application corresponding to the display interface being the first type of application, matching the semantic text information with the interface information, and after successful matching, outputting a first execution command to perform an operation corresponding to the voice command; In response to the information type of the application corresponding to the display interface being the second type of application, matching the semantic text information with the interface information, and after successful matching, outputting a second execution command to perform an operation corresponding to the voice command.
4. The method according to claim 2, wherein The method further includes: Before determining whether the information type of the application corresponding to the display interface is the first type of application or the second type of application, determining whether the application corresponding to the display interface is an application in a preset white list; After determining that the application corresponding to the display interface is an application in the preset white list, determining whether the information type of the application corresponding to the display interface is the first type of application or the second type of application; Wherein, the applications in the preset white list are applications that support visible and speakable functions.
5. The method according to claim 4, wherein Each application in the preset white list includes a label, and the label is used to identify that the application is the first type of application or the second type of application; The determining whether the information type of the application corresponding to the display interface is the first type of application or the second type of application includes: determining whether the information type of the application corresponding to the display interface is the first type of application or the second type of application based on the label for each application in the preset white list.
6. The method according to claim 4 or 5, characterized in that, The preset white list is deployed in the cloud or locally, and the preset white list is editable.
7. The method according to claim 3, wherein The obtaining interface information from the display interface based on the first information obtaining method includes: obtaining the interface information of the display interface through the software development kit of the application corresponding to the display interface; Output the first execution instruction to perform the operation corresponding to the voice instruction, including: outputting an execution instruction to the application corresponding to the display interface, so that the application corresponding to the display interface performs the operation corresponding to the voice instruction.
8. The method according to claim 3, characterized in that Obtain the interface information from the display interface based on the second information acquisition method, including: obtaining the interface information of the display interface through the accessibility service of the system service where the display interface is located; Output the second execution instruction to perform the operation corresponding to the voice instruction, including: outputting an execution instruction to the accessibility service of the system service where the display interface is located, so that the accessibility service performs the operation corresponding to the voice instruction.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: In response to the display interface presenting the interfaces of at least two applications, determine, for the information type of each of the at least two applications, whether it is the first type of application or the second type of application.
10. A visible and speakable implementation device, characterized in that, The device includes a processor and a memory, and a computer program run by the processor is stored on the memory. When the computer program is run by the processor, the processor is caused to execute the visible and speakable implementation method according to any one of claims 1-9.
11. A vehicle, characterized in that, The vehicle includes the visible and speakable implementation device according to claim 10.
12. A storage medium, characterized in that, A computer program run by a processor is stored on the storage medium. When the computer program is run by the processor, the processor is caused to execute the visible and speakable implementation method according to any one of claims 1-9.
Citation Information
Patent Citations
Voice control method and device for vehicle-mounted APP
CN111681658A
Content recommendation method and device and storage medium
CN112911068A
Voice control method and device and electronic equipment
CN114049892A
Human-computer interaction method, electronic equipment and system
CN114255745A
Vehicle-mounted voice interaction method and system and computer readable medium
CN116149596A