Voice instruction analysis method, terminal, server and management platform

By collecting and analyzing users' historical voice operation records, the system can recommend or correct voice commands, solving the problem that existing technologies cannot automatically sense the environment or user habits, and enabling more efficient voice operation and personalized services in smart home scenarios.

CN115731930BActive Publication Date: 2025-11-25NANNING FUGUI PRECISION IND CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111013335.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2025-11-25
Estimated Expiration
2041-08-31

AI Technical Summary

Technical Problem

Existing voice control systems cannot automatically sense the environment or user preferences in smart home scenarios, resulting in a poor user experience.

Method used

By collecting and analyzing users' historical voice operation records, the system can recommend or correct voice commands, and use terminals, servers, and management platforms to perform real-time voice command analysis to provide timely feedback and improve user experience.

Benefits of technology

It improves the automation of voice operation, reduces voice command recognition time, enhances user experience, and provides personalized services by analyzing popular keywords and applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731930B_ABST
    Figure CN115731930B_ABST
Patent Text Reader

Abstract

A voice instruction analysis method is applied to a terminal, and the method comprises the following steps: the terminal receives a voice input trigger signal; according to the voice input trigger signal, a history voice operation record of a user is compared, a recommended instruction is matched for the user; when the user does not accept the recommended instruction, voice data input by the user is received and voice instruction recognition is performed. The application further provides a terminal, a server and a management platform for implementing the voice instruction analysis method. The application can realize voice recognition and voice instruction analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing, and more particularly to a speech command analysis method, terminal, server, and management platform. Background Technology

[0002] In smart home scenarios, controlling interface devices or searching for multimedia content via voice is already a basic application. However, existing voice operation relies heavily on human intervention; one voice command triggers one action, and it cannot automatically sense the environment or the user's preferences and habits. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a voice command analysis method, terminal, server and management platform, which can enhance the user experience.

[0004] An embodiment of the present invention provides a voice command analysis method applied to a terminal. The method includes: the terminal receiving a voice input trigger signal; the terminal comparing the voice input trigger signal with the user's historical voice operation records to match a recommended command for the user; the terminal determining whether the user accepts the recommended command; and when the terminal determines that the user does not accept the recommended command, the terminal receiving voice data input by the user and recognizing a voice command from the voice data.

[0005] An embodiment of the present invention also provides a voice command analysis terminal, the terminal including a processor and a memory, the memory being used to store at least one instruction, and the processor being used to execute the at least one instruction to implement the voice command analysis method.

[0006] An embodiment of the present invention also provides a voice command analysis server, the server including a processor and a memory, the memory for storing at least one instruction, and the processor for executing the at least one instruction to cause the server to perform at least the following steps: collecting voice operation information of multiple terminals, wherein the voice operation information includes voice interaction information between users and terminals, and information on IoT devices that can be controlled by voice; grouping the multiple terminals into multiple groups according to the voice operation information of the multiple terminals; and sharing voice commands in the groups according to the voice operation information of each terminal in the groups.

[0007] An embodiment of the present invention also provides a voice command analysis and management platform. The management platform includes a processor and a memory. The memory is used to store at least one instruction, and the processor is used to execute the at least one instruction to cause the management platform to perform at least the following steps: collecting voice search records from multiple terminals, wherein the voice search records include operation time, keywords, and applications executed in the foreground; calculating the voice usage rate of the terminals based on the operation time in the voice search records of the multiple terminals, and dividing the multiple terminals into multiple groups based on the voice usage rate of the terminals; assigning different weights to these groups, wherein the group with a higher voice usage rate has a higher weight; and analyzing popular keywords and their corresponding applications within a preset time period based on the voice search records of the multiple terminals and the weights of the groups.

[0008] Compared to existing technologies, the voice command analysis method, terminal, server, and management platform provided by this invention can perform real-time voice command analysis by collecting voice operation records, providing timely feedback to users and effectively improving user experience. Attached Figure Description

[0009] Figure 1 This is an architecture diagram of a voice command analysis system according to an embodiment of the present invention.

[0010] Figure 2 This is a flowchart of a voice command analysis method according to an embodiment of the present invention.

[0011] Figure 3 This is a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0012] Figure 4 This is a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0013] Figure 5 This is a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0014] Figure 6 This is a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0015] Figure 7 This is a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0016] Figure 8 This is a block diagram of a voice command analysis terminal according to an embodiment of the present invention.

[0017] Figure 9 This is a block diagram of a voice command analysis server according to an embodiment of the present invention.

[0018] Figure 10 This is a block diagram of a voice command analysis and management platform according to an embodiment of the present invention.

[0019] Explanation of main component symbols

[0020]

[0021]

[0022] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0023] To facilitate understanding and implementation of this invention by those skilled in the art, the invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that this invention provides many applicable inventive concepts, which can be implemented in various specific forms. Those skilled in the art can utilize the details described in these or other embodiments, as well as other available structural, logical, and electrical variations, to implement the invention without departing from its spirit and scope.

[0024] This specification provides different embodiments to illustrate the technical features of different implementations of the invention. The configuration of elements in the embodiments is for illustrative purposes only and is not intended to limit the invention. Furthermore, the repetition of some reference numerals in the embodiments is for simplification and does not imply any correlation between different embodiments. The same element numbers used in the illustrations and specification represent the same or similar elements. The illustrations in this specification are simplified and not drawn to scale.

[0025] Furthermore, in describing some embodiments of the present invention, the specification describes the method and / or procedure of the present invention in a specific order of steps. However, since the method and procedure are not necessarily performed according to the specific order of steps described, they are not limited to the specific order of steps. Those skilled in the art will understand that other orders are also possible implementations. Therefore, the specific order of steps described in the specification is not intended to limit the scope of the patent application. Moreover, the scope of the present invention for the method and / or procedure is not limited to the order of execution steps written therein, and those skilled in the art will understand that adjusting the order of execution steps does not depart from the spirit and scope of the present invention.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items. Some embodiments of the invention are described in detail below with reference to the accompanying drawings.

[0027] Please see Figure 1 The diagram shown is an architecture diagram of a voice command analysis system 100 according to an embodiment of the present invention. Figure 1 As shown, system 100 includes terminal 110, one or more IoT devices 120, first server 130, second server 140, and management platform 150. The terminal 110, one or more IoT devices 120, first server 130, second server 140, and management platform 150 are all connected to network 160 and can exchange information with each other via wired or wireless means through network 160.

[0028] In one embodiment, when user 101 speaks to terminal 110, and terminal 110 recognizes the voice command, it converts the voice command into control information and transmits the control information to the corresponding IoT device 120. After receiving the control information, IoT device 120 executes the corresponding control operation and sends the operation result back to terminal 110.

[0029] In one embodiment, when terminal 110 cannot recognize a voice command, the voice data is further transmitted to a first server 130 for voice recognition to obtain the corresponding text information. The first server 130 then transmits the text information to a second server 140 to obtain the corresponding voice command and sends the voice command back to terminal 110. Terminal 110 converts the received voice command into control information and transmits the control information to the corresponding IoT device 120. Terminal 110 can also receive the execution result returned by IoT device 120. In this embodiment, although different servers are used for execution based on function, in different embodiments, the first server 130 and the second server 140 can also be deployed on the same server.

[0030] In one embodiment, management platform 150 provides an administrator interface, enabling the administrator of system 100 to manage various devices within system 100 via management platform 150.

[0031] In one embodiment, terminal 110 is a device with voice recognition function, including devices corresponding to mobile phones, tablet computers, wristbands, smart glasses, watches, smart speakers, and remote controls with voice control function, such as set-top boxes, TV boxes, or other devices with network communication and voice recognition functions.

[0032] In one embodiment, one or more IoT devices 120 may be devices that communicate with the network 160, such as smart lights, smart curtains, smart refrigerators, smart meters, smart sockets, smart air conditioners, smart washing machines, and smart water purifiers. User 101 can control the IoT devices 120 via voice through terminal 110.

[0033] Please see Figure 2 The diagram shows a flowchart of a voice command analysis method according to an embodiment of the present invention.

[0034] Step S201: The terminal receives a voice input trigger signal.

[0035] In some embodiments, the voice input trigger signal may be generated by the user pressing or touching an icon on the terminal screen, such as a microphone icon; the voice input trigger signal may also be generated by the user pressing a physical button on the terminal, such as the voice input button on the corresponding remote control; the voice input trigger signal may also be generated by keyword speech, such as "Hi" or "hey".

[0036] In step S202, the terminal compares the user's historical voice operation records and matches a recommended instruction for the user.

[0037] After receiving a voice input trigger signal, the terminal compares the user's historical voice operation records and matches a recommended instruction for the user.

[0038] The historical voice operation record is used to record operation information related to all recognized voice commands of the user.

[0039] In some embodiments, the historical voice operation record includes voice commands, operation time, operating information of the Internet of Things device, and applications running in the foreground.

[0040] Specifically, the terminal records the operating information of the IoT device when the voice command is recognized, as well as the application running in the foreground.

[0041] In some embodiments, the voice command may also include a voice search, and the historical voice operation record may also include keywords from the voice search.

[0042] In some embodiments, when the terminal collects the user's voice operation records, if it finds duplicate historical voice operation records, it accumulates the execution count for the duplicate historical voice operation records.

[0043] In some embodiments, the historical voice operation records may be stored in a voice operation database, which is stored in the terminal's memory.

[0044] In some embodiments, to reduce the amount of storage space occupied, the terminal may only record historical voice operation records within a recent period, such as one month or one week.

[0045] In some embodiments, to reduce the storage space occupied, the terminal may also record only historical voice operation records of successfully controlling IoT devices.

[0046] In some embodiments, the terminal compares the historical voice operation records with the current time, the current operating information of the IoT device, and the application currently running in the foreground to match a voice command as a recommended command for the user.

[0047] In some embodiments, when the terminal matches multiple voice commands, the voice command that is executed most frequently is selected as the recommended command.

[0048] In some embodiments, when the terminal cannot find any voice command, a suitable voice command is selected as the recommended command for the user. For example, a voice command from a historical voice operation record closest to the current time is selected as the recommended command. The time includes the date, day of the week, and specific time.

[0049] Through step S202, recommended instructions can be obtained before the user makes voice input, reducing the running time of voice instruction recognition and improving the user experience.

[0050] Step S203: The terminal determines whether the user accepts the recommendation instruction.

[0051] For example, before a user presses the microphone icon to speak, the terminal compares historical voice operation records and matches the recommended instruction "turn on the air conditioner." The terminal then asks the user via a graphical user interface or voice interface whether they want to turn on the air conditioner. When the terminal determines that the user accepts the recommended instruction, it converts the instruction into control information and transmits the control information to the corresponding IoT device. When the terminal determines that the user does not accept the recommended instruction, step S204 is executed.

[0052] For example, when the terminal compares historical voice operation records and matches a recommended command as voice search, it obtains the keywords for the voice search and asks the user whether they want to search for the keywords via a graphical user interface or a voice interface. When the terminal determines that the user accepts the recommended command, it immediately displays the corresponding search results; when the terminal determines that the user does not accept the recommended command, it executes step S204.

[0053] That is, through the execution of steps S202 and S203, recommended instructions can be obtained before the user makes voice input, reducing the running time of voice instruction recognition and increasing the user experience.

[0054] Step S204: The terminal receives voice data input by the user.

[0055] Step S205: The terminal recognizes voice commands from the voice data.

[0056] In one embodiment, the historical voice operation record further includes voice files corresponding to the voice commands. The terminal records the user-input voice data as an audio stream, converts it into a feature matrix using Mel-frequency cepstral coefficients, and stores the feature matrix of each successful voice command as a voice file. When the terminal receives user-input voice data, it obtains the similarity between the feature matrix of the audio stream of the voice data and each voice file in the historical voice operation record using a dynamic time correction algorithm. When a voice file is determined to have a similarity exceeding a default threshold, the voice command corresponding to that voice file can be used as the voice command of the recognition result.

[0057] For example, when a user inputs the voice command "turn on the air conditioning", the terminal can immediately obtain the recognized voice command "turn on the air conditioning" by comparing the historical voice operation records.

[0058] In one embodiment, when the terminal cannot recognize the voice command, for example, by comparing the feature matrix of the input audio with the similarity of each voice file in the historical voice operation record, and finding no voice file with a similarity exceeding a default threshold, the audio stream of the voice data is transmitted to the cloud server for recognition.

[0059] In one embodiment, the cloud server may include Figure 1 In different embodiments, the cloud server may be a single server, including the first server 130 and the second server 140.

[0060] In one embodiment, the cloud server first identifies the received audio stream to obtain corresponding text information, then converts the text information into a voice command, and finally sends the voice command back to the terminal. The terminal uses the received voice command as the voice command of the recognition result.

[0061] Please see Figure 3 The diagram shows a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0062] Step S301: The terminal receives voice data input by the user.

[0063] Step S302: The terminal obtains voice commands from the voice data, searches for historical voice operation records, and determines whether the voice command has a past operation record.

[0064] Specifically, the voice commands can be obtained through local voice recognition on the terminal or through remote voice recognition on a cloud server.

[0065] Step S303: When the terminal determines that the voice command has a past operation record, the terminal determines the most recent execution result of the voice command. The execution result includes success and failure.

[0066] Specifically, after acquiring the voice command, the terminal converts the voice command into control information and transmits it to the corresponding IoT device. Upon receiving the control information, the IoT device executes the control information and sends the execution result back to the terminal. The execution result includes success and failure. After receiving the response from the IoT device, the terminal records the voice command and its execution result in the historical voice operation record.

[0067] Step S304: When the terminal determines that the most recent execution result of the voice command was unsuccessful, it searches the historical voice operation record and finds the next voice command in the time sequence as the correction command.

[0068] In step S305, the terminal asks the user whether to replace the voice command with the correction command.

[0069] Please see Figure 4 The diagram shows a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0070] Step S401: The terminal receives voice data input by the user.

[0071] Step S402: The terminal obtains voice commands from the voice data, searches for historical voice operation records, and determines whether the voice command has a past operation record.

[0072] Specifically, the voice commands can be obtained through local voice recognition on the terminal or through remote voice recognition on a cloud server.

[0073] In this embodiment, when the terminal receives a voice command, it can simply play the voice command as feedback for the user's voice operation.

[0074] For example, if a user enters "turn on the laziness", the terminal will recognize the result as "turn on the heating". At this point, it can play "Okay, turning on the heating for you".

[0075] Step S403: When the terminal determines that the voice command has a past operation record, it continues to search for historical voice operation records, finds the next voice command in the order of the voice command, and determines whether the operation time interval between the next voice command and the voice command is less than the interval threshold.

[0076] Step S404: When the terminal determines that the time interval between the next voice command and the voice command is less than the interval threshold, the next voice command is used as the suggested command.

[0077] In one embodiment, the interval threshold may be set as the time interval between the next attempt to perform an operation when the user has not performed an operation successfully, for example, 3 seconds.

[0078] In step S405, the terminal asks the user whether to replace the voice command with the suggested command.

[0079] For example, a user wants to turn on the air conditioner's cooling function, but says "turn on the air conditioner," and the recognition result is "turn on the heating." At this time, the terminal checks the user's most recent execution result for "turn on the heating," which was unsuccessful, while the next successfully executed voice command was "turn on the air conditioner." Therefore, "turn on the air conditioner" is used as the correction command, and the user is asked "Do you want to turn on the air conditioner?".

[0080] For example, a user wants to turn on the cooling function of the air conditioner, but says "turn on the air conditioner," and the recognition result is "turn on the heating function." At this time, the terminal checks if the user has a record of repeating the operation within the interval threshold as "turn on the cooling function." Therefore, it asks the user "whether you accept replacing turning on the heating function with turning on the cooling function?" as a suggested instruction.

[0081] Please see Figure 5 The diagram shows a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0082] Step S501: The cloud server collects voice operation information from multiple terminals. This voice operation information includes user information, voice interaction information between the user and the terminal (e.g., historical voice operation records, including voice search records and voice commands), and information about IoT devices that the terminal can control via voice. The voice operation information can be periodically uploaded to the cloud server by one or more terminals, or it can be reported by the cloud server to one or more terminals.

[0083] Step S502: The cloud server divides the multiple terminals into multiple groups based on the voice operation information of the multiple terminals.

[0084] In one embodiment, the cloud server may use a clustering algorithm to cluster the terminals.

[0085] In one embodiment, the cloud server can calculate the similarity of user information, the similarity of voice interaction information between users and terminals, and the similarity of IoT device information that terminals can be controlled by voice using the Euclidean distance formula, and then group the terminals according to these similarities.

[0086] In one embodiment, the cloud server may periodically regroup the multiple terminals.

[0087] In step S503, the cloud server shares voice commands among the groups based on the voice operation information of each terminal in the groups.

[0088] In one embodiment, the cloud server determines the target terminal and the group to which the voice command is transmitted based on the most recently recognized voice command, and finds the next voice command executed by other terminals after executing the same voice command based on the voice operation information of other terminals in the group, as the shared voice command.

[0089] In one embodiment, the cloud server periodically uses the voice command that has been recognized most frequently in these groups as the shared voice command.

[0090] When the terminal receives the share voice command, it can ask the user of the terminal whether to execute the share voice command via a graphical user interface or a voice interface.

[0091] For example, when a user buys a new IoT device and completes the initial setup via the terminal, the cloud server can share the most frequently used voice commands from users who have also purchased the same IoT device with that user as the sharing voice command.

[0092] Please see Figure 6 The diagram shows a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0093] Step S601: The cloud server collects voice operation information from multiple terminals. This voice operation information includes voice interaction information between the user and the terminal, as well as information about IoT devices that can be controlled by voice. The voice interaction information between the user and the terminal includes historical voice operation records, which include voice commands and applications executed in the foreground.

[0094] In step S602, the cloud server performs correlation analysis on the voice operation information of the multiple terminals to establish the correlation between voice commands and applications executed in the foreground.

[0095] Specifically, after the cloud server collects the voice operation information of the multiple terminals, it can group the multiple terminals into multiple groups based on the fact that they have the same IoT devices that can be controlled by voice.

[0096] The cloud server uses correlation analysis to establish the correlation between voice commands and applications running in the foreground for terminals in various groups.

[0097] Step S603: When the cloud server receives a voice recognition request from the source terminal, it analyzes the next possible voice command that the source terminal may request to be recognized based on the result of the request recognition, the voice operation information of the multiple terminals, and the correlation between the voice command and the application executed in the foreground.

[0098] Specifically, when the cloud server determines that the user of the source terminal falls under a specific voice command and the application being executed in the foreground, it compares the correlation between the voice command identified in the request, the voice commands of the group to which the source terminal belongs, and the application being executed in the foreground to determine the next possible voice command to be identified. For example, based on the correlation between the voice commands of the group to which the source terminal belongs and the application being executed in the foreground, it can analyze the combinations of voice commands within a ten-minute interval of the operation time of the voice command identified in the current request, and select the combination of voice commands with the highest cumulative execution count as the next possible voice command to be identified.

[0099] In step S604, the cloud server transmits the result of the request recognition and the next possible voice command for request recognition to the source terminal.

[0100] In one embodiment, upon receiving the next possible voice command, the source terminal can ask the user whether to execute the next possible voice command via a graphical user interface or a voice interface, and then feed the user's choice back to the cloud server for learning.

[0101] Please see Figure 7 The diagram shows a flowchart of a voice command analysis method according to another embodiment of the present invention.

[0102] Step S701: The management platform collects voice search records from multiple terminals, including operation time, keywords, and applications executed in the foreground.

[0103] Step S702: The management platform calculates the voice usage rate of the terminals based on the operation time in the voice search records of the multiple terminals, and divides the multiple terminals into multiple groups based on the voice usage rate of the terminals.

[0104] Specifically, the management platform can organize the recent voice search operation time intervals of these terminals into a sequence, such as {X1, X2, X3, X4, ..., X...}. n Next, the standard deviation of the voice search operation time interval is calculated using the standard deviation formula. Finally, the K-Means algorithm is used to cluster the data to divide the terminals into multiple groups. The initial K-Means cluster value is set as the maximum difference value of the data sequence divided by the standard deviation.

[0105] In step S703, the management platform assigns different weights to these groups, with the group having a higher voice usage rate having a higher weight.

[0106] In step S704, the management platform analyzes popular keywords and their corresponding applications within a preset time period based on the voice search records of these terminals and the weight of these groups.

[0107] The preset time can be the most recent week.

[0108] In one embodiment, the management platform also determines the push information for each terminal based on the popular keywords and their corresponding applications within the preset time period.

[0109] Specifically, these terminals are OTT service providers with multiple video streaming platform applications installed on them, each with its own video program notifications. When the management platform determines that a terminal has an application corresponding to a popular keyword (such as iQiyi) installed, but the popular keyword (such as Nirvana in Fire) is not present in the recent notifications of the video streaming platform, it can push information about the popular keyword to the terminal, which may change the user's video streaming platform usage behavior.

[0110] Please see Figure 8 The diagram shown is a block diagram of a terminal 800 used for voice command analysis in one embodiment of the present invention.

[0111] Terminal 800 includes a processor 801, a memory 802, a communication element 803, a display screen 804, and an audio element 805. Those skilled in the art should understand that... Figure 8 The composition of the terminal 800 shown does not constitute a limitation of the embodiments of the present invention. Figure 8 The terminal 800 shown is simplified for ease of description; in different embodiments, it may include fewer or more components than shown.

[0112] In some embodiments, the processor 801 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 801 is the control unit of the terminal 800, connecting various components of the terminal 800 via various interfaces and lines. It executes computer programs or modules stored in the memory 802 and calls data stored in the memory 802 to perform various functions of the terminal 800 and process data, such as voice command analysis methods.

[0113] In some embodiments, the memory 802 is used to store computer program code and various data, such as multimedia recommendation methods, and to enable high-speed, automatic access to programs or data during the operation of the device 800. The memory 802 includes read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable storage medium capable of carrying or storing data.

[0114] The communication element 803 is used to communicate with external devices via a network. The communication element 803 can receive requests and information from external devices, and can also send requests and information to the external devices. The external devices can be IoT devices, servers, management platforms, or other network devices. The communication element 803 can communicate with external devices through at least one wired or wireless communication protocol.

[0115] The display screen 804 is used to display a graphical user interface (GUI), which may include graphics, text, icons, videos, and any combination thereof. In one embodiment, the display screen may display push notifications, such as recommendation instructions, correction instructions, and suggestion instructions. In another embodiment, the terminal 800 may also be a device without a display screen 404, such as a smart speaker, which plays push notifications via audio element 805.

[0116] The audio element 805 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input to the processor 801 for processing, or input to the communication element 803 to realize voice communication. The speaker is used to convert the electrical signals from the processor 801 or the communication element 803 into sound waves.

[0117] Please see Figure 9 The diagram shown is a block diagram of a server 900 for voice command analysis in one embodiment of the present invention.

[0118] Server 900 includes processor 901, memory 902, and communication element 903. Those skilled in the art should understand that... Figure 9 The composition of the server 900 shown does not constitute a limitation of the embodiments of the present invention. Figure 9 The server 900 shown is simplified for ease of description; in different embodiments, it may include fewer or more components than shown.

[0119] In some embodiments, the processor 901 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 901 is the control unit of the server 900, connecting various components of the server 900 via various interfaces and lines. It executes computer programs or modules stored in the memory 902 and calls data stored in the memory 902 to perform various functions of the server 900 and process data, such as voice command analysis methods.

[0120] In some embodiments, the memory 902 is used to store computer program code and various data, such as multimedia recommendation methods, and to enable high-speed, automatic access to programs or data during the operation of the server 900. The memory 902 includes read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable storage medium capable of carrying or storing data.

[0121] The communication element 903 is used to communicate with external devices via a network. The communication element 903 can receive requests and information from external devices, and can also send requests and information to the external devices. The external devices can be terminals, IoT devices, management platforms, or other network devices. The communication element 903 can communicate with external devices through at least one wired or wireless communication protocol.

[0122] Please see Figure 10 The diagram shown is a block diagram of a management platform 1000 for voice command analysis in one embodiment of the present invention.

[0123] The management platform 1000 includes a processor 1001, a memory 1002, and a communication element 1003. Those skilled in the art should understand that... Figure 10 The composition of the management platform 1000 shown does not constitute a limitation of the embodiments of the present invention. Figure 10 The management platform 1000 shown is simplified for ease of description; in different embodiments, it may include fewer or more components than shown.

[0124] In some embodiments, the processor 1001 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 1001 is the control unit of the management platform 1000, connecting various components of the management platform 1000 via various interfaces and lines. It executes computer programs or modules stored in the memory 1002 and calls data stored in the memory 1002 to perform various functions of the management platform 1000 and process data, such as multimedia recommendation methods.

[0125] In some embodiments, the memory 1002 is used to store computer program code and various data, such as multimedia recommendation methods, and to achieve high-speed and automatic access to programs or data during the operation of the management platform 1000. The management platform 1000 includes read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable storage medium capable of carrying or storing data.

[0126] The communication element 1003 is used to communicate with external devices via a network. The communication element 1003 can receive requests and information from external devices, and can also send requests and information to the external devices. The external devices can be terminals, IoT devices, servers, or other network devices. The communication element 1003 can communicate with external devices through at least one wired or wireless communication protocol.

[0127] In summary, the voice command analysis method, terminal, server, and management platform of this invention can predict a user's voice commands before the user performs any voice operation, and can also correct or suggest appropriate voice commands, saving users time and reducing backend recognition costs, thus speeding up user operations. Furthermore, by analyzing the popularity of voice search keywords, it can serve as a reference for marketing advertisements by audio-visual platform providers or management platform administrators, enhancing the user experience.

[0128] It is worth noting that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A voice command analysis method, applied to a terminal, characterized in that, The terminal is communicatively connected to at least one Internet of Things (IoT) device, which can be controlled via voice input from a user on the terminal. The method includes: The terminal receives a voice input trigger signal; The terminal compares the user's historical voice operation records with the voice input trigger signal and matches a recommended instruction for the user. The terminal determines whether the user accepts the recommendation instruction; When the terminal determines that the user does not accept the recommendation instruction, the terminal receives the voice data input by the user and recognizes the voice instruction from the voice data; The terminal searches the historical voice operation records according to the voice command to determine whether the voice command has a past operation record; When the terminal determines that the voice command has a past operation record, the terminal determines the most recent execution result of the voice command on at least one Internet of Things device, wherein the execution result includes success and failure; When the terminal determines that the most recent execution result of the voice command was unsuccessful, the terminal searches the historical voice operation record and finds the next voice command in chronological order as the correction command. The terminal asks the user whether to replace the voice command with the correction command; The terminal converts the voice commands into control information for the corresponding Internet of Things (IoT) devices. The terminal transmits the control information to the Internet of Things device; The terminal receives a response from the IoT device regarding the execution result of the control information; and The terminal records the voice command and its execution result to the historical voice operation record based on the response.

2. The method as described in claim 1, characterized in that, The historical voice operation records include: voice commands, operation time, IoT device operation information, and applications running in the foreground.

3. The method as described in claim 2, characterized in that, The method further includes: The terminal compares the historical voice operation records with the current time, the current operating information of the IoT device, and the application currently running in the foreground, and matches a voice command from the historical voice operation records as the recommended command for the user.

4. A voice command analysis terminal, characterized in that, The terminal includes a processor and a memory, the memory being used to store at least one instruction, and the processor being used to execute the at least one instruction to implement the voice instruction analysis method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Text data correction method and device

    CN107622054A

  • Multi-person participation-based man-machine interaction method and apparatus

    CN107831903A

  • Method and device for processing information interaction

    CN112908319A

  • Speech recognition method

    KR1020040107232A