Voice Command Execution Method, Device, Cloud Server, and Storage Medium

By executing voice command operation instructions on cloud servers, the problem of frequent updates of voice assistants is solved, user experience and R&D efficiency is improved, and cumbersome operations of smart terminals are reduced.

CN114446292BActive Publication Date: 2025-07-08SHENZHEN TCL NEW-TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011223513.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-05
Publication Date
2025-07-08
Estimated Expiration
2040-11-05

AI Technical Summary

Technical Problem

In the prior art, frequent updates and repairs of voice assistants lead to excessive consumption of user time and network traffic, affecting user experience, and low R&D efficiency.

Method used

By executing voice command operation instructions on the cloud server, avoiding tedious operations in the smart terminal, and only updating and repairing voice assistants in the cloud server.

Benefits of technology

It improves user experience and R&D efficiency, reduces the need for frequent upgrades of smart terminals, and saves user time and network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114446292B_ABST
    Figure CN114446292B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, an apparatus, a cloud server and a storage medium for executing a voice instruction. The method includes: obtaining voice operation request information generated by a smart terminal based on a voice instruction; determining a service field corresponding to the voice instruction according to the voice operation request information; determining an operation instruction corresponding to the service field, and sending the operation instruction to the smart terminal. By executing the operation instruction of the voice information in the cloud in the embodiments of the present invention, the function of the voice assistant without upgrade can be realized, and the problems existing in the voice assistant can be solved in the first time to improve the R & D efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and particularly to a method, device, cloud server, and storage medium for executing voice commands. Background Art

[0002] Today, with the booming development of natural language processing technologies, voice interaction technologies have become increasingly mature. The application upgrade method of voice assistants requires downloading the latest version for update first. However, with the development of terminals and the improvement of user requirements, voice assistants need to be updated frequently to meet user needs. Sometimes, untimely updates and bug fixes will affect the user experience.

[0003] Therefore, the existing technologies still need to be improved and developed. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for cloud update and real-time intervention of voice assistants in view of the above-mentioned defects of the existing technologies, aiming to solve the problems that when there is a new APK in the existing technologies, the terminal needs to remind users to download and upgrade, and frequent upgrades by users will consume users' time and network traffic, seriously affecting the product experience. In addition, for developers, they hope to quickly fix bugs in voice assistants.

[0005] The technical solution adopted by the present invention to solve the problem is as follows:

[0006] In a first aspect, an embodiment of the present invention provides a method for executing a voice command, including:

[0007] Obtaining voice operation request information generated by an intelligent terminal based on a voice command;

[0008] Determining the service field corresponding to the voice command according to the voice operation request information;

[0009] Determining an operation command corresponding to the service field, and sending the operation command to the intelligent terminal.

[0010] In a second aspect, an embodiment of the present invention further provides a device for executing a voice command, including:

[0011] An obtaining unit, configured to obtain voice operation request information generated by an intelligent terminal based on a voice command.

[0012] A determining unit, configured to determine the service field corresponding to the voice command according to the voice operation request information.

[0013] A sending unit, configured to determine an operation command corresponding to the service field, and send the operation command to the intelligent terminal.

[0014] Thirdly, an embodiment of the present invention further provides a cloud server, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for executing a voice instruction as described in any one of the above is implemented.

[0015] Fourthly, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method for executing a voice instruction as described in any one of the above is implemented.

[0016] Advantages of the present invention: In an embodiment of the present invention, first, voice operation request information is obtained, and the voice operation request information is generated by an intelligent terminal according to a voice instruction; then, a business field is determined according to the voice operation request information, and the business field corresponds to the voice instruction. Next, an operation instruction to be finally executed is determined, and the operation instruction corresponds to the business field. Finally, the operation instruction is sent to the intelligent terminal, realizing the execution of the voice instruction in the cloud, thereby avoiding cumbersome operations in the intelligent terminal. It can be seen that in the embodiment of the present invention, by generating an operation instruction of the voice instruction in the cloud, frequent upgrades of the voice assistant in the intelligent terminal are avoided, the R & D efficiency is improved, and the user experience is good. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 Schematic flowchart of the method for executing a voice instruction provided by an embodiment of the present invention

[0019] Figure 2 Principle block diagram of the device for executing a voice instruction provided by an embodiment of the present invention.

[0020] Figure 3 Internal structure principle block diagram of the cloud server provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The present invention discloses a method for executing a voice instruction, a cloud server, and a storage medium. To make the purpose, technical solution, and effect of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0022] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0023] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0024] In the prior art, since R & D personnel will continuously repair functional defects with the development of intelligent terminals and the improvement of user requirements, new versions of voice assistants will appear every once in a while. If users want to use the latest version, they have to first download the latest version and then install it, resulting in the need to frequently update the voice assistant on the intelligent terminal, which brings great inconvenience to users.

[0025] To solve the problems of the prior art, this embodiment provides a method for executing voice commands. By using the method for executing voice commands in this embodiment, when executing a voice command, it is possible to obtain the voice operation request information generated by the intelligent terminal according to the voice command, analyze and process the voice operation request information based on the received voice operation request information, determine the service area corresponding to the voice command, and obtain the corresponding operation command according to the service area. Finally, the operation command is sent to the intelligent terminal, thus avoiding the cumbersome operations of the intelligent terminal. Moreover, as the user's requirements increase or the intelligent terminal continues to develop, when new requirements or functional defects occur, it is only necessary to repair the operation command and update it in real time on the cloud server, avoiding the frequent upgrade of the voice assistant in the intelligent terminal, improving the R & D efficiency, and providing a good user experience. Specifically, when the cloud server executes a voice command in this embodiment, the intelligent terminal will generate voice operation request information from the user's voice command, and the intelligent terminal sends the voice operation request information to the cloud server. The cloud server obtains the voice operation request information sent by the intelligent terminal, determines the service area corresponding to the user's voice command through the voice operation request information, and then determines the corresponding operation command through the service area. The cloud server sends the operation command to the intelligent terminal. Since the operation commands are all generated on the cloud server, after the intelligent terminal obtains the user's voice command, it only needs to convert the voice command into voice operation request information, and then send the voice operation request information to the cloud server. The cloud server generates the operation command according to the voice command and then sends the operation command to the intelligent terminal. Therefore, when the intelligent terminal receives the user's voice command, by generating the operation command corresponding to the voice command on the cloud server, the voice command can be executed, avoiding the cumbersome operations of the intelligent terminal, improving the R & D efficiency, and providing a good user experience.

[0026] For example, when the cloud server executes a voice command, it obtains a voice operation request message. This voice operation request message is generated by the intelligent terminal according to the user's voice command. Since the intelligent terminals are a wide range of devices distributed in different regions, each intelligent terminal may generate different voice operation request messages at the same time, and all these voice operation request messages need to be processed in a timely manner. In order to process the voice operation request messages in these intelligent terminals, next, according to these voice operation request messages, the business field corresponding to the voice command needs to be obtained. This business field can be various fields such as film and television, weather query, device control, shopping, consumption, etc. For example, according to the voice operation request message "I want to watch 'Call Me by Fire'", its corresponding field can be determined as the film and television field. Then, according to the business field, the operation command corresponding to the business field is determined. That is, after determining that it is the film and television field, a series of operation commands related to the film and television field can be obtained according to the film and television field. The operation command refers to the code segment related to the film and television field in the program segment where the R & D personnel implement the corresponding operation. Finally, the cloud server sends the operation command to the intelligent terminal, and the intelligent terminal directly receives the operation command corresponding to the film and television field. That is, in the embodiment of the present invention, the operation commands related to the voice command in the intelligent terminal are executed on the cloud server, which is equivalent to all intelligent terminals sharing the operation commands related to the voice command on the cloud server, so that the R & D personnel only need to update the operation commands on the cloud server in real time and release the latest version, avoiding updating and upgrading the corresponding operation command version in the intelligent terminal, thus bringing convenience to the user to use the voice command of the intelligent terminal.

[0027] Exemplary method

[0028] This embodiment provides a method for executing a voice command, and this method can be applied to the cloud server of intelligent voice recognition. Specifically as Figure 1 shown, the method includes:

[0029] Step S100, obtain the voice operation request message generated by the intelligent terminal based on the voice command.

[0030] In this embodiment, the intelligent terminal (intelligent TV) sends the voice operation request message generated based on the voice command to the cloud server, and updates and iterates the voice assistant on the cloud server, thus avoiding the real-time upgrade of the voice assistant in the intelligent terminal and saving the user's usage time. The intelligent terminal can be any large-screen system, intelligent TV and other devices that may appear in practice.

[0031] Specifically, when the user uses the recording function of the voice assistant and presses the recording button, the voice assistant will record what the user says. The voice module in the voice assistant calls the voice assistant recognition module to convert what the user says into voice operation request information, and the voice operation request information will be displayed on the intelligent terminal device. The voice operation request information refers to the intention information of the user to perform a certain operation. The voice assistant is an intelligent application that realizes solving user problems through intelligent interaction of intelligent dialogue and instant Q&A. It mainly helps users solve life-related problems. The voice assistant recognition module is used to convert what the user says into text and display the text (after user QUERY) on the terminal device. User QUERY refers to the user's query, which is used to find a specific file, website, record or a series of records in the database, and is the message sent by the search engine or the database. For example, when the user records with the voice assistant: "I want to open iQIYI." The intelligent terminal calls the voice assistant recognition module and displays the voice operation request information on the intelligent terminal. If the voice operation request information displayed on the intelligent terminal is not the content expressed by the user, the user clicks the return button, and the intelligent terminal re-executes the above operation, records the voice instruction and converts it into voice operation request information. When the user confirms that the voice operation request information is the correct voice instruction expressed by the user, the intelligent terminal sends the voice operation request information after the user query to the cloud server, and the cloud server receives the voice operation request information.

[0032] In one implementation, this embodiment provides a method for executing a voice instruction, and this method can be applied to the cloud server for intelligent voice recognition. Specifically, as Figure 1 shown, the method includes:

[0033] S200: Determine the business field corresponding to the voice instruction according to the voice operation request information.

[0034] In this embodiment, the cloud server cannot directly obtain the user's voice command. Therefore, it needs to obtain the voice command through the interconnection and communication with the intelligent terminal. The user sends the voice command to the intelligent terminal through the recording function of the voice assistant, and the intelligent terminal then converts the voice command into voice operation request information, that is, the intention information of the user to perform a certain operation. Since the intelligent terminals are distributed in different regions and the voice operation request information received by each intelligent terminal is different, a large number of voice operation request information will be generated in the same time period. Each intelligent terminal needs to send the voice operation request information to the cloud server. The cloud server obtains the voice operation request information sent by the intelligent terminal. Since each operation request information comes from different intelligent terminals and different users, the meaning represented by each operation request information is different, and its corresponding business field is also different. Therefore, it is necessary to match the business field according to the voice operation request information, and the business field corresponds to the voice command sent by the user. For example, the user uses the voice assistant to record: "I want to open iQIYI". The intelligent terminal converts it into voice operation request information and sends it to the cloud server. The cloud server can then determine that the voice operation request information in the intelligent terminal corresponds to the film and television field.

[0035] In order to more accurately match the voice command and the business field, according to the voice operation request information, determining the business field corresponding to the voice command includes the following steps:

[0036] S201: Analyze the voice operation request information to obtain the text information corresponding to the voice command;

[0037] S202: Determine the business field corresponding to the voice command according to the text information.

[0038] Specifically, since the voice operation request information refers to the intention information of the user to perform a certain operation, the voice operation request information contains various contents. The voice operation request information sent by each intelligent terminal represents different intentions. Therefore, it is necessary to analyze the voice operation request information to obtain the text information. The text information comes from the analyzed voice operation request information and its content corresponds to the voice command. For example, the user uses the voice assistant to record, that is, the voice command: "I want to open iQIYI". The intelligent terminal converts the voice command into voice operation request information and sends it to the cloud server. The cloud server receives the voice operation request information, analyzes it, and obtains the text information that can be recognized by the cloud server: "I want to open iQIYI".

[0039] The cloud server determines the corresponding business area according to the received text information. The text information refers to the literal expression information in the voice operation request information recognized by the cloud server. In practice, at the same time period, different users' text information will be received from different intelligent terminals, and each text information represents a different intention. The cloud server will obtain the corresponding business area according to each text information, and the business area also corresponds to the voice command sent by the intelligent terminal. For example, after the cloud server recognizes the text information: "I want to open iQIYI", it determines that its business area is the film and television field. As can be seen from the above, the film and television field also corresponds to the voice command sent by the user: "I want to open iQIYI".

[0040] In one implementation, the text information also contains different parts. When corresponding the text information to the business area, only the partial information in the text information is needed to obtain its corresponding business area. Therefore, it is necessary to first decompose the text information to obtain field information; then, according to the field information, determine the business area that matches the field information.

[0041] Specifically, the cloud server decomposes the text information to obtain field information. The field information is also the keyword, and the field information refers to the object of the voice command execution operation. For example: for the text information "I want to open iQIYI", the cloud server decomposes it into two parts. One part is "I want to open", and the other part is "iQIYI". At this time, after decomposing the text information, the field information "iQIYI" is obtained.

[0042] In this embodiment, the field information obtained by the cloud server through decomposing the text information. In practice, different users operate different devices and generate various types of text information. Similarly, the field information decomposed by the cloud server is also distributed in different business domains. Therefore, the cloud server matches the corresponding business domain according to different field information. In practice, the cloud server inputs these field information into the cloud server, calculates the confidence level according to artificial intelligence technology, and then performs speech matching to match the field information to the corresponding business domain. Artificial intelligence is the human intelligence demonstrated with machines as the carrier, so artificial intelligence is also called machine intelligence. The confidence level is also called reliability, or confidence level, confidence coefficient. That is, when making an estimate of the population parameter in sampling, according to the randomness of the sample, using the interval estimation method in mathematical statistics, the corresponding probability value generated when the estimated value and the population parameter are within a certain allowable error range. Speech matching is to generate corresponding response content according to the input information. For example, the cloud server decomposes to obtain the field information "iQIYI", inputs the field information "iQIYI" into the artificial intelligence algorithm model, and the artificial intelligence technology will retrieve the business domains related to "iQIYI" in the database, and then match and estimate these business domains with "iQIYI". When the matching probability value between the business domain and "iQIYI" meets the preset value, the corresponding business domain of "iQIYI" can be determined as the film and television domain.

[0043] In another implementation manner, there are the following special situations, such as the situation that requires real-time intervention. Therefore, the field information is rewritten into specified field information, and a specified domain corresponding to the specified field information is set, and the specified domain is the business domain.

[0044] Specifically, when a special application scenario appears in practice, at this time, the field information obtained by decomposing the text information needs to be rewritten into specified field information. For example: During the Two Sessions, it is necessary to block the calls of apps involving the external network and news pushes. The cloud server can remotely operate the user device and call various built-in functions. For example: According to the user's identity information, remotely configure settings for the user, modify "I want to open iQIYI" to the Two Sessions theme content, and send the "Two Sessions theme content" to artificial intelligence for matching to obtain the national politics domain. In addition, when some emergencies occur, such as detecting an earthquake, then "I want to open iQIYI" needs to be modified to an earthquake forecast, and the specified field information "earthquake" is sent to artificial intelligence for matching, and the climate domain is obtained according to the matching result. In addition, when the identity information of the intelligent terminal is the IP address of a suspect, combined with the longitude and latitude information, the cloud server remotely calls the voice assistant to record the intelligent terminal, and sends the specified field information "criminal suspect" to artificial intelligence for matching, and the public security domain is obtained according to the matching result.

[0045] After the field information is rewritten into the specified field information, a one-to-one correspondence is formed between the specified field information and the specified field. To process subsequent similar intervention situations more quickly, the corresponding relationship needs to be saved. Therefore, it is necessary to create a mapping relationship between the specified field information and the specified field and store the mapping relationship.

[0046] Specifically, when a special situation occurs, the field information obtained by decomposing the text information is rewritten into the specified field information, and the field information will correspond to the specified field. For example: map "the theme content of the Two Sessions" to the national politics field in the intervention template, map the specified field information "earthquake prediction" to the earthquake field in the intervention template, and map the specified field information "criminal suspect" to the public security field in the intervention template, and store the mapping relationship in the memory space of the cloud server. In this way, when a similar situation occurs again next time, the specified field information can be quickly mapped to its corresponding business field, and the cloud server can quickly determine the business field corresponding to its mapping relationship according to the specified field information, improving the speed of executing voice command operations and providing a good user experience.

[0047] In one implementation, this embodiment provides a method for executing a voice command, which can be applied to a cloud server for intelligent speech recognition. Specifically as Figure 1 shown, the method includes:

[0048] S300: Determine the operation instruction corresponding to the business field and send the operation instruction to the intelligent terminal.

[0049] Specifically, after the cloud server determines the business field corresponding to each voice command sent by the user to the intelligent terminal, it will determine the operation instruction related to the business field. The operation instruction is a command set for the user to execute the voice command, that is, a code segment written by the R & D personnel for the related operations of the voice command. In practice, in order to centralize the development and optimization work of the voice assistant on the cloud server and reduce the inconvenience brought by the user's real-time upgrade, the cloud server matches the corresponding operation instruction according to the business field. When the cloud server obtains the operation instruction corresponding to the business field, the cloud server sends the corresponding operation instruction to the intelligent terminal, and the intelligent terminal can thus execute the user's voice command. For example: according to actual needs, the intelligent terminal application communicates with the cloud server through a series of general operation interfaces, and determines the operation instruction corresponding to the business field according to the business field. The cloud server and the intelligent terminal can use Internet communication, and the cloud server sends the operation instruction to the intelligent terminal in the form of json data. Json data refers to a lightweight data exchange format. It is based on a subset of the specifications formulated by the European Computer Manufacturers Association and uses a text format completely independent of programming languages to store and represent data. The json data structure is as follows:

[0050] {

[0051] "directives":{

[0052] "action":"App.Open",

[0053] "appName":"Tencent Video"

[0054] },

[0055] "data":{

[0056] "extend":"Free for members",

[0057] "category":"Movie",

[0058] "thumb":

[0059] "http: / / puui.qpic.cn / vcover_vt_pic / 0 / 00jxecd5him5kmn1585271336 / 770",

[0060] "token":"tenvideo2: / / ?action=7&video_id=&video_name=Pirates of the Caribbean: Dead Men Tell No Tales

[0061] &cover_id=00jxecd5him5kmn",

[0062] "publishDate":20170526,

[0063] "tags":

[0064] "Humorous",

[0065] "Disaster",

[0066] "Adventure",

[0067] "Adventure"

[0068] ,

[0069] "resource_name":"Pirates of the Caribbean: Dead Men Tell No Tales"

[0070] },

[0071] }

[0072] In this embodiment, the cloud server obtains the name information of the business domain and needs to perform some processing to determine the operation instruction. Therefore, according to the business domain, determining the operation instruction corresponding to the business domain includes the following steps:

[0073] S301: Obtain the name information of the business domain;

[0074] S302: Determine the operation instruction corresponding to the business domain according to the name information.

[0075] Specifically, each application will be mapped to a business domain, and the name of the application will also have name information corresponding to the business domain. That is to say, the name information of the business domain refers to the name information corresponding to the application name in each business domain in the business domain. Therefore, it is necessary to obtain the name information of the business domain according to the field information, and then obtain the corresponding operation instruction according to the name information of the business domain. For example: for the field information "iQIYI", the decomposed business domain is the film and television domain, and its name in the film and television domain in the cloud server is "iQIYI". Then, according to the name information of "iQIYI", the operation instruction corresponding to the film and television domain can be determined.

[0076] In this embodiment, to obtain the operation instruction corresponding to the field information according to the domain name, some processing is also required. Therefore, it is necessary to obtain the instruction template corresponding to the name information according to the name information; and obtain the application package name corresponding to the field information, and fill the application package name into the instruction template to generate the operation instruction corresponding to the business domain.

[0077] Specifically, when the cloud server obtains the name information of the business domain, it can obtain its corresponding instruction template. In practice, since the voice instructions of users contain a lot of content and the corresponding business domains are also of many types, in order to improve the processing efficiency of the operation instructions corresponding to the voice instructions, the operation instructions will be classified and processed according to a certain category. Therefore, in this embodiment, the cloud server will establish the corresponding relationship between the name information of the business domain and the instruction template. When the cloud server obtains different voice instructions sent by the user on the intelligent terminal, after converting the voice instructions into text information to obtain the field information, it will determine the business domain according to the field information corresponding to the voice instruction, then find the name information of the business domain according to the business domain, and then find the corresponding instruction template according to the name information of the business domain. For example, the cloud server can calculate the confidence level according to artificial intelligence technology, and then perform speech matching, match the field information to the corresponding business domain and determine that its corresponding instruction template is the video control template.

[0078] In this embodiment, the cloud server first obtains the application package name in the field information, and the application package name is the operation object corresponding to the operation instruction. In practice, the intelligent terminal will send a large amount of text information, and these text information contain field information in multiple different business fields. The cloud server will obtain the corresponding application package name according to the field information and fill the application package name into the instruction template corresponding to its business field, so as to generate an operation instruction corresponding to the business field. Specifically, the R & D personnel have actually generated code segments according to the relationship between the business field and the operation instruction, and will leave an interface for the cloud server to execute the corresponding operation instruction according to different business fields. When the cloud server obtains the application package name "Galaxy Kiwi" in the field information, it will fill the application package name "Galaxy Kiwi" into the instruction template (video control template), and generate an operation instruction corresponding to the business field based on the code segments developed by the R & D personnel before.

[0079] In this embodiment, the execution of the operation instruction is a whole. Only through the application package name corresponding to the field information in the field information, the complete operation instruction cannot be obtained. Therefore, it is necessary to first obtain the behavior information corresponding to the field information, and the behavior information is used to reflect the operation behavior corresponding to the operation instruction; then fill the behavior information and the application package name into the instruction template; finally, according to the instruction template, call the instruction generation program to generate an operation instruction corresponding to the business field.

[0080] Specifically, in addition to the field information corresponding to the business field, the field information also includes behavior information, and the behavior information is the action performed on the object (field information) in the text information, that is, it is used to reflect the operation behavior corresponding to the operation instruction. Therefore, the cloud server also needs to obtain the behavior information in the text information.

[0081] In actual application, the field information sent by each user through the intelligent terminal is different, and the business field to which it belongs is also different. The cloud server needs to fill the behavior information and the application package name in the field information corresponding to each user into the instruction template corresponding to the business field at the same time. For example, when the cloud server obtains the text information "I want to open iQIYI", it obtains the behavior information "I want to open" in the text information and the field information "iQIYI" in the text information. The cloud server searches for the application package name "Galaxy Kiwi" corresponding to the field information "iQIYI", so it fills the behavior information "I want to open" and the application package name "Galaxy Kiwi" into the instruction template, and the finally generated operation instruction is as follows:

[0082] {

[0083] "domain":"app_control",

[0084] "actions":[{

[0085] "property":{

[0086] "action":"App.Open",

[0087] "appName":"Galaxy Kiwi"

[0088] },

[0089] "startType":"app",

[0090] "component":{

[0091] "pkg":""

[0092] }

[0093] }]

[0094] In this embodiment, after the behavior information and the application package name are filled into the instruction template, the corresponding instruction template is called to generate operation instructions corresponding to the business field. For example, according to the identified business field, the service is assigned to different execution modules. When the business field is identified as alarm setting, the alarm instruction module is called; when the business field is identified as weather report, the weather instruction module is called; when the business field is identified as music play, the music instruction module is called; when the business field is identified as anything, the any instruction module is called, and finally the execution module gives the corresponding operation instructions.

[0095] Exemplary device

[0096] like Figure 2 As shown in , an embodiment of the present invention provides a voice command execution device, which includes an acquisition unit 401, a determination unit 402 and a sending unit 403, wherein:

[0097] An acquisition unit 401 is used to acquire voice operation request information generated by the intelligent terminal based on the voice instruction;

[0098] A determination unit 402, configured to determine the business domain corresponding to the voice instruction according to the voice operation request information;

[0099] The sending unit 403 is used to determine the operation instruction corresponding to the business field and send the operation instruction to the smart terminal.

[0100] Based on the above embodiments, the present invention further provides a cloud server, whose principle block diagram can be shown as follows: Figure 3As shown. The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected via a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it realizes a method for executing voice instructions. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen. The temperature sensor of the intelligent terminal is pre-set inside the intelligent terminal and is used to detect the operating temperature of internal devices.

[0101] Those skilled in the art can understand that Figure 3 In the schematic diagram, it is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0102] In one embodiment, a cloud server is provided. The cloud server includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it realizes the instructions for the following operations:

[0103] Obtain the voice operation request information generated by the intelligent terminal based on the voice instruction;

[0104] Determine the business field corresponding to the voice instruction according to the voice operation request information;

[0105] Determine the operation instruction corresponding to the business field and send the operation instruction to the intelligent terminal.

[0106] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0107] In summary, the present invention discloses a method for executing a voice instruction, a smart terminal, and a storage medium. The method includes: obtaining voice operation request information generated by the smart terminal based on a voice instruction; determining the service field corresponding to the voice instruction according to the voice operation request information; determining an operation instruction corresponding to the service field, and sending the operation instruction to the smart terminal. By executing the operation instruction of the voice information in the cloud in the embodiments of the present invention, the function of the voice assistant not needing to be upgraded is realized, and at the same time, it can ensure that the problems existing in the voice assistant can be solved in the first time, improving the R & D efficiency. It should be understood that the present invention discloses a method for executing a voice instruction.

[0108] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for executing a voice command, characterized in that, including: Obtain the voice operation request information generated by the intelligent terminal based on the voice instruction; Determine the business field corresponding to the voice instruction according to the voice operation request information; Determine the operation instruction corresponding to the business field, and send the operation instruction to the intelligent terminal; The determining the business field corresponding to the voice instruction according to the voice operation request information includes: Parse the voice operation request information to obtain the text information corresponding to the voice instruction; Determine the business field corresponding to the voice instruction according to the text information; The determining the business field corresponding to the voice instruction according to the text information includes: Decompose the text information to obtain field information; Determine the business field that matches the field information according to the field information; The determining the business field that matches the field information according to the field information includes: Rewrite the field information into specified field information, and set a specified field corresponding to the specified field information, where the specified field is a business field.

2. The method according to claim 1, wherein The setting the specified field corresponding to the specified field information includes: Create a mapping relationship between the specified field information and the specified field, and store the mapping relationship.

3. The method according to claim 2, wherein The determining the operation instruction corresponding to the business field includes: Obtain the name information of the business field; Determine the operation instruction corresponding to the business field according to the name information.

4. The method according to claim 3, wherein The determining the operation instruction corresponding to the business field according to the name information includes: Obtain the instruction template corresponding to the name information according to the name information; Obtain the application package name corresponding to the field information, and fill the application package name into the instruction template to generate an operation instruction corresponding to the business field.

5. The method according to claim 4, wherein The obtaining the application package name corresponding to the field information, and filling the application package name into the instruction template to generate an operation instruction corresponding to the business field includes: Obtain the behavior information in the field information, where the behavior information is used to reflect the operation behavior corresponding to the operation instruction; Fill the behavior information and the application package name into the instruction template; Call an instruction generation program according to the instruction template to generate an operation instruction corresponding to the business field.

6. A voice command execution device, characterized in that, including: An obtaining unit, configured to obtain the voice operation request information generated by the intelligent terminal based on the voice instruction; A determining unit, configured to determine the business field corresponding to the voice instruction according to the voice operation request information; A sending unit, configured to determine the operation instruction corresponding to the business field, and send the operation instruction to the intelligent terminal; The determining unit includes a component for parsing the voice operation request information to obtain the text information corresponding to the voice instruction; decomposing the text information to obtain field information; Rewriting the field information into specified field information, and setting a specified field corresponding to the specified field information, where the specified field is a business field.

7. A cloud server, characterized in that, The cloud server includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of claims 1-5 is implemented.

8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the method described in any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Voice instruction distribution method and device, electronic device and computer readable medium

    CN109918040A

  • Voice interaction method, device and system and storage medium

    CN110021299A