Voice processing method, device and computer storage medium
By analyzing user's historical voice operation commands and establishing association relationships, the problem that voice recognition execution operations do not meet user expectations is solved, and speech recognition satisfaction evaluation and optimization are achieved, improving the accuracy and user satisfaction of speech recognition.
Patent Information
- Application Number
- CN202110359571.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-02
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-04-02
AI Technical Summary
In the prior art, the speech recognition input by the user is successful but the operations performed do not meet the user's expectations, affecting the speech recognition satisfaction, and lacking effective satisfaction analysis and optimization methods.
By analyzing the historical voice operation commands before entering the voice operation commands of the target operation object, identifying operations that meet and do not meet the user's needs, using the voice recognition model for training and optimization, and establishing association relationships to improve the accuracy of speech recognition.
It realizes accurate evaluation and optimization of user speech recognition satisfaction, improves the accuracy and user experience of speech recognition, and simplifies the operation process.
Smart Images

Figure CN115171672B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and in particular to a speech processing method, device and computer storage medium. Background Art
[0002] With the rapid development of voice recognition technology, voice recognition is becoming increasingly common in vehicles, and users are using it in a wider range of scenarios. To optimize voice recognition, it's often necessary to analyze user satisfaction with it. Existing methods for analyzing user satisfaction with voice recognition primarily rely on analyzing whether the user's voice input is successfully recognized. However, even if the user's voice input is successfully recognized, the vehicle's actions based on the input may not be what the user intended, which can also affect user satisfaction with voice recognition.
[0003] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0004] An object of the present invention is to provide a speech processing method, device and computer storage medium, which have the advantage of analyzing the user's use of speech recognition to obtain the user's satisfaction with speech recognition, and can provide assistance for speech recognition optimization.
[0005] Another object of the present invention is to provide a voice processing method, device and computer storage medium, which have the advantage of accurately obtaining the voice recognition satisfaction of the corresponding voice operation commands by analyzing the input time of different voice operation commands, which is simple and easy to operate.
[0006] Another object of the present invention is to provide a speech processing method, device and computer storage medium, which have the advantage of further providing assistance for speech recognition optimization by improving speech operation commands whose speech recognition results do not meet user requirements.
[0007] Other advantages and features of the present invention will become more apparent from the following detailed description and will be realized by means of the instrumentalities and combinations particularly pointed out in the appended claims.
[0008] According to one aspect of the present invention, a speech processing method of the present invention that can achieve the aforementioned objects and other objects and advantages includes the following steps:
[0009] Obtaining at least one second voice operation command for a target operation object input by a user before the user inputs a first voice operation command for the target operation object; wherein an operation performed based on the first voice operation command meets the user's needs, and an operation performed based on the second voice operation command does not meet the user's needs;
[0010] A voice recognition satisfaction analysis is performed based on the number of the second voice operation commands to obtain an analysis result for characterizing the user's satisfaction with the voice recognition.
[0011] According to one embodiment of the present invention, the step of obtaining at least one second voice operation command input by the user to the target operation object before the user inputs the first voice operation command to the target operation object comprises the following steps:
[0012] Obtain a set of historical voice operation commands consisting of historical voice operation commands whose input times are adjacent and whose time intervals meet a preset condition;
[0013] The historical voice operation command corresponding to the latest input time in the historical voice operation command set is determined as a first voice operation command, and the historical voice operation commands other than the historical voice operation command corresponding to the latest input time are determined as second voice operation commands.
[0014] According to one embodiment of the present invention, after performing voice recognition satisfaction analysis based on the number of the second voice operation commands to obtain an analysis result representing the user's satisfaction with voice recognition, the following steps are further included:
[0015] The second voice operation command is used as the input of the set voice recognition model, and the operation performed based on the first voice operation command is used as the output of the voice recognition model to train the voice recognition model.
[0016] Accordingly, the present invention provides a device for executing the above-mentioned voice processing method, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the following steps when executing the computer program: obtaining at least one second voice operation command for the target operation object input by the user before the user inputs the first voice operation command for the target operation object; wherein the operation executed based on the first voice operation command meets the user's needs, while the operation executed based on the second voice operation command does not meet the user's needs; performing a voice recognition satisfaction analysis based on the number of the second voice operation commands to obtain an analysis result for characterizing the user's satisfaction with voice recognition.
[0017] Accordingly, the present invention provides a computer storage medium storing a computer program, wherein the computer program implements the steps of the above-mentioned speech processing method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A flowchart of a speech processing method provided by an embodiment of the present invention;
[0019] Figure 2 This is a structural diagram of a speech processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0021] It should be noted that, in this document, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.
[0022] It should be understood that although the terms first, second, third, etc. may be used herein to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, without departing from the scope of this document, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the term "if" as used herein may be interpreted as "at the time of," "when," or "in response to a determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms "comprising" and "including" indicate the presence of the described features, steps, operations, elements, components, items, types, and / or groups, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, types, and / or groups. The terms "or" and "and / or" as used herein are to be interpreted as inclusive, meaning any one or any combination. Thus, “A, B, or C” or “A, B, and / or C” means “any of: A; B; C; A and B; A and C; B and C; A, B, and C.” An exception to this definition occurs only when a combination of elements, functions, steps, or operations are inherently mutually exclusive in some manner.
[0023] It should be understood that, although the various steps in the flowchart in the embodiment of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless clearly stated herein, the execution of these steps is not strictly limited in order, and they can be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and their execution order is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0024] It should be noted that in this article, step codes such as S101 and S102 are used for the purpose of expressing the corresponding content more clearly and concisely, and do not constitute a substantial limitation on the order. When implementing the step, those skilled in the art may execute S102 first and then S101, etc., but these should all be within the scope of protection of this application.
[0025] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0026] In the subsequent description, the use of suffixes such as "module", "component" or "unit" to represent elements is only for the purpose of facilitating the description of the present application and has no specific meaning. Therefore, "module", "component" or "unit" can be used interchangeably.
[0027] See also Figure 1 , which is a flow chart of a speech processing method provided by an embodiment of the present invention. This method can be applied to analyzing user speech recognition satisfaction. This method can be executed by a speech processing device provided by an embodiment of the present invention. The speech processing device can be implemented in software and / or hardware. The speech processing device can specifically be a terminal such as a mobile phone, a vehicle computer, a wearable device, or a server. This embodiment uses the speech processing method applied to a vehicle computer as an example for description. The method includes the following steps:
[0028] Step S101: Obtaining at least one second voice operation command for a target operation object input by a user before inputting a first voice operation command for the target operation object; wherein an operation performed based on the first voice operation command meets the user's needs, while an operation performed based on the second voice operation command does not meet the user's needs;
[0029] The operation performed based on the first voice operation command satisfies the user's needs if it is the operation performed based on the first voice operation command that the user intended to perform or achieve. For example, if the first voice operation command is to turn on the air conditioner, then if the operation performed based on the first voice operation command is to turn on the air conditioner, then the operation performed based on the first voice operation command satisfies the user's needs. If the operation performed based on the first voice operation command is not to turn on the air conditioner, such as to open the windows, then the operation performed based on the first voice operation command does not meet the user's needs. It should be noted that the vehicle computer successfully recognized both the first and second voice operation commands, but the recognition results differed, meaning that the vehicle computer performed different operations based on the first and second voice operation commands, respectively. Furthermore, the time interval between the input of the first and second voice operation commands should be less than a preset duration threshold. The target operation object can be a specific object, such as the air conditioner or windows, or an abstract object, such as a car radio or multimedia application. That is, when the operation performed based on the first voice operation command meets the user's needs, it means that the vehicle computer successfully recognized the first voice operation command and the corresponding operation performed is what the user wants. When the operation performed based on the second voice operation command does not meet the user's needs, it means that the vehicle computer successfully recognized the second voice operation command, but the corresponding operation performed is not what the user wants. It can be understood that when the user performs voice control operations through the vehicle computer, if the user wants to control the vehicle computer to perform operation A, the user can enter the corresponding voice operation command. However, due to technical factors such as recognition accuracy and / or human factors such as accent, the vehicle computer may recognize the operation to be performed by the voice operation command as operation B. At this time, the user may continue to enter the same voice operation command or enter an adjusted voice operation command until the vehicle computer successfully performs operation A based on the input voice operation command, and then stop voice input. At this time, the last voice operation command entered is determined to be the first voice operation command, and the voice operation command entered before the last voice operation command is determined to be the second voice operation command.
[0030] In one embodiment, the step of obtaining at least one second voice operation command input by the user to the target operation object before the user inputs the first voice operation command to the target operation object comprises the following steps:
[0031] Obtain a set of historical voice operation commands consisting of historical voice operation commands whose input times are adjacent and whose time intervals meet a preset condition;
[0032] The historical voice operation command corresponding to the latest input time in the historical voice operation command set is determined as a first voice operation command, and the historical voice operation commands other than the historical voice operation command corresponding to the latest input time are determined as second voice operation commands.
[0033] It is understandable that users often continue to input voice operation commands after the vehicle computer fails to recognize an input voice operation command, resulting in a temporal correlation between the voice operation commands. That is, the time interval between each voice operation command is not too long and is less than the time interval during which the user normally uses the voice recognition function. Therefore, the first and second voice operation commands can be obtained based on the input time and time interval. It should be noted that due to different desired operations performed by the vehicle computer, multiple first and second voice operation commands may be obtained, and accordingly, multiple second voice operation commands may also be obtained. The preset condition may include that the time interval is less than the minimum or average time interval during which the user uses the voice recognition function. Historical data on the user's use of the voice recognition function can be collected and analyzed to determine the user's characteristics or habits regarding voice recognition, such as the maximum, minimum, or average time interval. In this way, the first and second voice operation commands can be obtained based on the input time and time interval. By analyzing the input time of different voice operation commands, the voice recognition satisfaction of the corresponding voice operation commands can be accurately determined, which is simple and convenient to operate.
[0034] In one embodiment, obtaining a set of historical voice operation commands consisting of historical voice operation commands inputted at adjacent times and with a time interval meeting a preset condition comprises the following steps:
[0035] Sort the historical voice operation commands entered by the user within a preset time period in order of input time from the earliest to the latest;
[0036] Determining a target historical voice operation command from the sorted historical voice operation commands, wherein the time interval between the input time of the target historical voice operation command and the input time of the previous historical voice operation command does not meet a preset condition, but the time interval between the input time of the target historical voice operation command and the input time of the next historical voice operation command meets a preset condition;
[0037] Taking the target historical voice operation command as a starting point, historical voice operation commands whose input times are adjacent and whose time intervals meet a preset condition are selected in sequence from the sorted historical voice operation commands and added to the historical voice operation command set.
[0038] The preset duration can be set based on actual needs, such as 30 days, 60 days, etc. When the time interval between the input time of a historical voice operation command and the input time of the previous historical voice operation command does not meet the preset condition, but the time interval between the input time of the next historical voice operation command meets the preset condition, it indicates that the user wants the historical voice operation command and the next historical voice operation command to perform the same operation, that is, the content of the historical voice operation command and the next historical voice operation command should be the same, but the user wants the historical voice operation command and the previous historical voice operation command to perform different operations, that is, the content of the historical voice operation command and the previous historical voice operation command should be different. It is understandable that when the operation performed by the vehicle computer based on the currently input voice operation command does not meet the user's needs, the user will quickly continue to input the next voice operation command. In other words, in order for the vehicle computer to perform the same operation, the input time interval between each two voice operation commands should not be too long. Therefore, starting from the target historical voice operation command, historical voice operation commands with adjacent input times and time intervals that meet the preset condition are selected from the sorted historical voice operation commands and added to the historical voice operation command set. In this way, by analyzing the input time of historical voice operation instructions, the required voice operation instructions can be accurately extracted, further improving the accuracy of obtaining voice recognition satisfaction.
[0039] Step S102: performing a voice recognition satisfaction analysis based on the number of the second voice operation commands to obtain an analysis result for characterizing the user's satisfaction with voice recognition.
[0040] Here, if the number of second voice operation commands is smaller, it means that the user inputs fewer voice operation commands, the vehicle computer's voice recognition accuracy is higher, and the user's satisfaction with voice recognition is higher. If the number of second voice operation commands is larger, it means that the user inputs more voice operation commands, the vehicle computer's voice recognition accuracy is lower, and the user's satisfaction with voice recognition is lower. Therefore, the analysis results obtained by performing voice recognition satisfaction analysis based on the number of second voice operation commands can represent user satisfaction with voice recognition, that is, can determine the voice recognition accuracy, and thus provide assistance for voice recognition optimization.
[0041] In one embodiment, the voice recognition satisfaction analysis is performed based on the number of the second voice operation commands to obtain an analysis result for characterizing the user's satisfaction with voice recognition, including the following steps: based on the number of the second voice operation commands and the correspondence between the different numbers of preset voice operation commands and the satisfaction levels, a target satisfaction level corresponding to the number of the second voice operation commands is determined to obtain an analysis result for characterizing the user's satisfaction with voice recognition. The correspondence between the different numbers of preset voice operation commands and the satisfaction levels can be set according to actual needs. For example, when the number of the second voice operation commands is 0, the level corresponding to the voice recognition satisfaction can be set to level ten; when the number of the second voice operation commands is 1, the level corresponding to the voice recognition satisfaction can be set to level nine, and so on. In addition, the voice recognition satisfaction can also be scored based on the number of the second voice operation commands. In this way, the user's satisfaction with voice recognition can be quickly evaluated, which improves the efficiency of analysis.
[0042] In summary, in the speech processing method provided in the above embodiment, by analyzing the user's use of speech recognition and obtaining the user's satisfaction with speech recognition, it can provide assistance for speech recognition optimization.
[0043] In one embodiment, after performing a speech recognition satisfaction analysis based on the number of the second voice operation commands and obtaining an analysis result for characterizing the user's satisfaction with speech recognition, the method further includes the following steps: using the second voice operation command as the input of a set speech recognition model, using the operation performed based on the first voice operation command as the output of the speech recognition model, and training the speech recognition model. It is understandable that since the operation performed based on the first voice operation command meets the user's needs, while the operation performed based on the second voice operation command does not meet the user's needs, it can be considered that the recognition result of the set speech recognition model for the second voice operation command is incorrect. There may be many factors that affect the recognition result, such as differences in speaking accents of people in different regions, homophones, polyphones, etc. Therefore, using the second voice operation command as the input of the set speech recognition model, using the operation performed based on the first voice operation command as the output of the speech recognition model, and training the speech recognition model to improve the adaptability of the speech recognition model, thereby correspondingly improving the recognition accuracy. It should be noted that the speech recognition model can be established based on the historical voice operation commands and corresponding recognition results of different users using artificial intelligence algorithms such as genetic algorithms, neural network algorithms, etc. In this way, by training the speech recognition model using the voice operation commands actually input by the user, the adaptability of the speech recognition model can be effectively improved, thereby correspondingly improving the recognition accuracy.
[0044] In one embodiment, after performing voice recognition satisfaction analysis based on the number of the second voice operation commands to obtain an analysis result representing the user's satisfaction with voice recognition, the method further includes the following steps:
[0045] performing semantic recognition on the second voice operation command to obtain at least one keyword;
[0046] An association relationship between the at least one keyword and an operation executed based on the first voice operation command is established, and the association relationship is stored in a set voice command library.
[0047] It is understood that by performing semantic recognition on the second voice operation command, at least one keyword contained in the second voice operation command can be obtained. During the voice recognition process, the accuracy of the voice recognition results is affected by whether the keyword contained in the voice operation command is correctly recognized. Due to differences in user accents, the use of homophones or polyphones in the voice operation command, and other factors, the same word may be misrecognized or fail to be recognized during the voice recognition process due to differences between the user's pronunciation and the preset pronunciation. To improve voice recognition satisfaction, the keyword contained in the second voice operation command whose executed operation does not meet the user's requirements can be associated with the operation performed based on the first voice operation command. This association can be stored in a preset voice command library. This association ensures that the operation performed based on the second voice operation command meets the user's requirements when the user subsequently enters the second voice operation command. For example, if the user mispronounces a polyphonetic word during voice input, such as pronouncing the fourth tone instead of the second tone, the operation performed by the vehicle computer based on the voice input may not meet the user's requirements. Therefore, the word corresponding to the fourth tone of the polyphonetic word can be associated with the operation correctly triggered by the second tone of the polyphonetic word, thereby improving the user's voice recognition satisfaction. In this way, by analyzing different voice operation commands for the target operation object, when the subsequent input operation does not meet the voice operation command of the user's needs, it can also correctly trigger the execution of the operation that meets the user's needs, thereby further improving the user's voice recognition satisfaction.
[0048] In one embodiment, after storing the association relationship in a set voice command library, the following steps are further included:
[0049] Output a prompt message, where the prompt message is used to indicate that inputting the second voice operation command may execute an operation based on the first voice operation command.
[0050] It is understandable that after different users input the second voice operation command into the car computer, if the operation performed by the car computer based on the second voice operation command does not meet the user's needs, some users may continue to try to input voice, while some users may not continue to use the voice function. Therefore, after establishing an association between the at least one keyword and the operation performed based on the first voice operation command, a prompt message can be output to indicate that the operation performed based on the first voice operation command can be performed by inputting the second voice operation command, so as to encourage users to use the voice function, thereby improving the convenience and accuracy of the voice recognition function. Suppose that a user once input a voice "blow the glass window", and the car computer replied with a voice "I don't understand what you are saying". At this time, the user may no longer use the voice function and feel that the voice is not good. After a month, the car computer may prompt "You can say blow the glass window to turn on the defrost function."
[0051] Based on the same inventive concept as the above embodiments, this embodiment describes in detail the technical solutions of the speech processing methods provided by the above embodiments through different examples.
[0052] Example 1
[0053] The purpose of the voice processing method provided in this example is to provide an overall indicator representing the maturity of voice recognition capabilities by statistically analyzing the effectiveness of user voice operations. It also establishes a data model for voice functions to monitor and improve user satisfaction.
[0054] The implementation principle of the speech processing method provided in this example is as follows:
[0055] First, when speech recognition is started, the digital signal of the speech is collected and recorded and saved as a speech binary file;
[0056] Next, the recognition results are recorded and analyzed. If an operation function is triggered accordingly, it does not mean that the recognition of the user's voice is correct, and the correctness of the function operation of the voice can be judged in combination with manual operation. For example, assuming that after voice input of the navigation destination, the user directly starts navigation after recognition, it means that the operation performed based on the voice command is correct. On the contrary, the user is very likely to continue to voice input the navigation destination, or even multiple times, or the user may input text, which means that the operation performed based on the voice command is incorrect, and the recognition rate is very poor at this time. Or, assuming that the user uses voice to control song playback, if the corresponding song is played normally after the user inputs the voice, but the user does not perform any further operation, it means that the recognition of the song contained in the voice is correct. However, if the user continues to input voice instructions related to song playback, it means that the user's needs are not met, that is, the recognition of the song contained in the voice is incorrect.
[0057] Finally, the score is calculated based on the number of voice commands. For example, if the voice command is successfully completed in one attempt, the score is 100 points; if the second attempt is successful, the score is 90 points, and so on. Of course, the success rate of the user's voice command operation can also be counted, ranging from 1 success to 10 successes.
[0058] Example 2
[0059] The purpose of the voice processing method provided in this example is to: analyze the time of voice commands to detect whether the user is still using the voice function; improve the cloud model for previously unrecognized voices and unexecuted commands, remotely upgrade them to the command library, and prompt the user that they can use it. Different users will have different improvements to voice input and different prompts.
[0060] The implementation principle of the speech processing method provided in this example is as follows:
[0061] First, count the time and frequency of a user's voice recognition use and establish user voice commands sorted by time;
[0062] Next, the user's voice commands in the time series are analyzed to obtain the maximum interval time, the last time, and the previous time;
[0063] Next, count the frequency of the user's previous voice input. If the time interval is within a few seconds, it means that the user is continuously inputting voice, indicating that the recognition accuracy of the user's voice command is low, and it is determined that the user is dissatisfied at this time;
[0064] Next, perform speech analysis and statistics on the original user speech files, detect and analyze the speech that was not correctly recognized and make improvements;
[0065] Finally, these users are prompted individually that voices that were not recognized before can now be recognized, and these quick user voice commands are added.
[0066] In this way, the maturity of the user's voice function can be monitored to analyze and improve the user's satisfaction with voice commands and the user's intentions; it improves the user experience, facilitates statistics on the user's main intentions for using voice, and improves the ease and accuracy of using these functions; it provides a basis for diagnosing the causes of user command failures, facilitating debugging and improvement.
[0067] Based on the same inventive concept as the above embodiments, the embodiment of the present invention provides a speech processing device, such as Figure 2 As shown, the speech processing device includes: a processor 110 and a memory 111 for storing a computer program that can be run on the processor 110; wherein, Figure 2The processor 110 shown in the figure is not used to indicate that the number of processors 110 is one, but is only used to indicate the positional relationship of the processor 110 relative to other devices. In actual applications, the number of processors 110 may be one or more; similarly, Figure 2 The memory 111 shown in the figure has the same meaning, that is, it is only used to refer to the position relationship of the memory 111 relative to other devices. In actual application, the number of memories 111 can be one or more. Among them, the processor 110 is used to implement the steps of the above-mentioned speech processing method when running the computer program.
[0068] The speech processing device may also include: at least one network interface 112. The various components in the speech processing device are coupled together via a bus system 113. It is understood that the bus system 113 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 113 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 2 Various buses are labeled as bus system 113 .
[0069] The memory 111 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory may be a magnetic disk or a magnetic tape. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memory 111 described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memories.
[0070] The memory 111 in the embodiment of the present invention is used to store various types of data to support the operation of the voice processing device. Examples of these data include: any computer program for operating on the voice processing device, such as an operating system and an application; contact data; phone book data; messages; pictures; videos, etc. Among them, the operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application program can include various applications, such as a media player (Media Player), a browser (Browser), etc., which are used to implement various application services. Here, the program that implements the method of the embodiment of the present invention can be included in the application program.
[0071] Based on the same inventive concept as the aforementioned embodiment, this embodiment further provides a computer storage medium storing a computer program. The computer storage medium may be a memory such as a ferromagnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface mount memory, an optical disc, or a compact disc read-only memory (CD-ROM); or various devices including one or any combination of the aforementioned memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc. When the computer program is executed by a processor, the steps of the aforementioned speech processing method are implemented.
[0072] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0073] As used herein, the terms "comprises," "comprising," or any other variation thereof, are intended to cover a non-exclusive inclusion of elements other than the listed elements and may also include additional elements not specifically listed.
[0074] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A speech processing method, characterized in that: The method comprises the following steps: Obtaining at least one second voice operation command for a target operation object input by a user before the user inputs a first voice operation command for the target operation object; wherein an operation performed based on the first voice operation command meets the user's needs, and an operation performed based on the second voice operation command does not meet the user's needs; performing a voice recognition satisfaction analysis based on the number of the second voice operation commands to obtain an analysis result for characterizing the user's satisfaction with voice recognition; wherein the number of the second voice operation commands is negatively correlated with the voice recognition satisfaction; The step of obtaining at least one second voice operation command on the target operation object input by the user before the user inputs the first voice operation command on the target operation object comprises the following steps: Obtain a set of historical voice operation commands consisting of historical voice operation commands whose input times are adjacent and whose time intervals meet a preset condition; The historical voice operation command corresponding to the latest input time in the historical voice operation command set is determined as a first voice operation command, and the historical voice operation commands other than the historical voice operation command corresponding to the latest input time are determined as second voice operation commands. 2 . The method according to claim 1 , wherein the preset condition includes that the time interval is less than a minimum time interval or an average time interval of the user using the voice recognition function.
3. The method according to claim 1 or 2, wherein obtaining a set of historical voice operation commands consisting of historical voice operation commands inputted at adjacent times and with a time interval that satisfies a preset condition comprises the following steps: Sort the historical voice operation commands entered by the user within a preset time period in order of input time from the earliest to the latest; Determining a target historical voice operation command from the sorted historical voice operation commands, wherein the time interval between the input time of the target historical voice operation command and the input time of the previous historical voice operation command does not meet a preset condition, but the time interval between the input time of the target historical voice operation command and the input time of the next historical voice operation command meets a preset condition; Taking the target historical voice operation command as a starting point, historical voice operation commands whose input times are adjacent and whose time intervals meet a preset condition are selected in sequence from the sorted historical voice operation commands and added to the historical voice operation command set.
4. The method according to claim 1, further comprising the following steps: performing a voice recognition satisfaction analysis based on the number of the second voice operation commands to obtain an analysis result representing the user's satisfaction with the voice recognition: The second voice operation command is used as the input of the set voice recognition model, and the operation performed based on the first voice operation command is used as the output of the voice recognition model to train the voice recognition model.
5. The method according to claim 1, further comprising the following steps: performing a voice recognition satisfaction analysis based on the number of the second voice operation commands to obtain an analysis result representing the user's satisfaction with the voice recognition; performing semantic recognition on the second voice operation command to obtain at least one keyword; An association relationship between the at least one keyword and an operation executed based on the first voice operation command is established, and the association relationship is stored in a set voice command library.
6. The method according to claim 5, further comprising the following steps after storing the association relationship in a preset voice command library: Output a prompt message, where the prompt message is used to indicate that inputting the second voice operation command may execute an operation based on the first voice operation command.
7. The method according to claim 1, wherein performing voice recognition satisfaction analysis based on the number of the second voice operation commands to obtain an analysis result representing the user's satisfaction with voice recognition comprises the following steps: Based on the number of the second voice operation commands and the correspondence between the different numbers of preset voice operation commands and the satisfaction levels, the target satisfaction level corresponding to the number of the second voice operation commands is determined to obtain an analysis result for characterizing the user's satisfaction with voice recognition.
8. A speech processing device, comprising: a memory configured to store one or more computer programs; and a processor coupled to the memory and configured to execute the one or more computer programs to enable the speech processing apparatus to perform the steps of the speech processing method according to any one of claims 1 to 7.
9. A computer storage medium storing a computer program, wherein: When the computer program is executed by a processor, the steps of the speech processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and device for evaluating a speech recognition system
FR3102603A1
Signal processing apparatus and method of recognizing a voice command thereof
US20100179812A1
User satisfaction detection in a virtual assistant
US20190035386A1
Human machine interface system and method for improving user experience based on history of voice activity
US20190103099A1
Voice interaction method and apparatus, electronic device, and storage medium
WO2023226700A1