Voice processing method, mobile terminal and computer readable storage medium
After detecting the first voice information in the mobile terminal, it is determined whether it matches the subsequently detected second voice information, and performs speech recognition when matching, the problem of inaccurate voice recognition in the prior art is solved, and the recognition accuracy and user experience are improved.
Patent Information
- Application Number
- CN201910485603.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-06-03
AI Technical Summary
Existing mobile terminals can only recognize the currently received voice, resulting in inaccurate voice recognition results and affecting the user experience.
When the first voice information is detected, it is determined whether the second voice information is detected within the preset time interval after the current time. If so, it is determined whether the two match, and voice recognition is performed when the match is made to obtain the recognition result.
It realizes the simultaneous recognition of the associated first voice information and the second voice information, avoids the problem of inaccurate voice recognition, and improves the accuracy and user experience of voice recognition.
Smart Images

Figure CN110148409B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech processing technology, and in particular to a speech processing method, a mobile terminal and a computer-readable storage medium. Background Art
[0002] With the continuous advancement of voice recognition technology, users can perform various operations by inputting voice when using mobile terminals.
[0003] However, the existing voice recognition function can only recognize a continuous voice input by the user. For the current voice, when performing voice recognition, only the currently received voice can be recognized, resulting in inaccurate voice recognition results, affecting the user experience. For example, through the voice assistant of the mobile terminal, when the user says the name of a star, the mobile terminal recognizes the voice containing the star's name and performs related operations, such as searching and outputting the star's recent news or the star's personal information, etc., and then the user inputs the voice again through voice "Is she beautiful (is he handsome)?" The voice assistant of the mobile terminal cannot recognize the voice, which affects the user experience.
[0004] The above contents are only used to assist in understanding the technical solution of the present invention and do not constitute an admission that the above contents are prior art. Summary of the invention
[0005] The main purpose of the present invention is to provide a speech processing method, a mobile terminal and a computer-readable storage medium, aiming to solve the technical problem that the existing mobile terminal only recognizes the currently received speech, resulting in inaccurate speech recognition results.
[0006] To achieve the above object, the present invention provides a speech processing method, which comprises the following steps:
[0007] When the first voice information is detected, determining whether the second voice information is detected within a preset time interval after the current moment;
[0008] If so, determining whether the first voice information matches the second voice information;
[0009] When the first voice information matches the second voice information, performing voice recognition on the first voice information and the second voice information to obtain a first voice recognition result;
[0010] Determine a first voice instruction corresponding to the first voice recognition result, and execute an operation corresponding to the first voice instruction.
[0011] Further, after the step of determining whether the first voice information matches the second voice information, the voice processing method further includes:
[0012] When the first voice information does not match the second voice information, performing voice recognition on the second voice information to obtain a second voice recognition result;
[0013] Obtain a second voice instruction corresponding to the second voice recognition result, and execute an operation corresponding to the second voice instruction.
[0014] Further, the step of determining whether the first voice information matches the second voice information includes:
[0015] Acquire text information corresponding to the first voice information and the second voice information;
[0016] Determine a sentence vector corresponding to the text information;
[0017] Calculating a first similarity between the sentence vector and each preset sentence vector in a preset database;
[0018] Determine whether the first voice information matches the second voice information based on the first similarity.
[0019] Further, the step of determining whether the first voice information matches the second voice information based on the first similarity includes:
[0020] Determine a maximum similarity among the first similarities, and judge whether the maximum similarity is greater than a first preset similarity, wherein when the maximum similarity is greater than the first preset similarity, it is determined that the first voice information matches the second voice information.
[0021] Furthermore, the step of determining the sentence vector corresponding to the text information includes:
[0022] Training the text information based on a word vector model to obtain a word vector corresponding to the text information;
[0023] Calculate the similarity between the word vector in the matching sentence vector and the word vector, and generate a similar word matrix based on the similarity, wherein the elements of each row in the similar word matrix are the similarities between the same word vector and the word vector in the matching sentence vector;
[0024] The sentence vector is generated based on the maximum similarity in each column element of the similar word matrix.
[0025] Further, if yes, the step of determining whether the first voice information matches the second voice information includes:
[0026] If yes, based on the preset time window, the first voice information and the second voice information are sampled respectively at a preset frequency to obtain first sampled data corresponding to the first voice information and second sampled data corresponding to the second voice information;
[0027] generating a first voiceprint feature vector based on the first sampling data, and generating a second voiceprint feature vector based on the second sampling data;
[0028] Based on the first voiceprint feature vector and the second voiceprint feature vector, determining whether the first identity information corresponding to the first voice information is the same as the second identity information corresponding to the second voice information;
[0029] If so, determine whether the first voice information matches the second voice information.
[0030] Furthermore, the first voiceprint feature vector is a first timbre feature vector, and the second voiceprint feature vector is a second timbre feature vector; and the step of determining whether the first identity information corresponding to the first voice information and the second identity information corresponding to the second voice information are the same based on the first voiceprint feature vector and the second voiceprint feature vector comprises:
[0031] Calculating a second similarity between the first timbre feature vector and the second timbre feature vector;
[0032] Determine whether the second similarity is greater than a second preset similarity, wherein if so, it is determined that the first identity information is the same as the second identity information.
[0033] Further, when the first voice information is detected, the step of determining whether the second voice information is detected within a preset time interval after the current moment includes:
[0034] When the first voice information is detected, performing voice recognition on the first voice information to obtain a third voice recognition result;
[0035] Obtaining a third voice instruction corresponding to the third voice recognition result, and executing an operation corresponding to the third voice instruction;
[0036] Determine whether second voice information is detected within a preset time interval after the moment when the first voice information is detected.
[0037] In addition, to achieve the above-mentioned purpose, the present invention also provides a mobile terminal, which includes: a memory, a processor, and a voice processing program stored in the memory and executable on the processor, and the voice processing program implements the steps of the aforementioned voice processing method when executed by the processor.
[0038] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a speech processing program is stored, and when the speech processing program is executed by a processor, the steps of the aforementioned speech processing method are implemented.
[0039] The present invention determines whether second voice information is detected within a preset time interval after the current moment when first voice information is detected, and then, if so, determines whether the first voice information matches the second voice information, and then, when the first voice information matches the second voice information, performs voice recognition on the first voice information and the second voice information to obtain a first voice recognition result, and then determines a first voice instruction corresponding to the first voice recognition result, and executes an operation corresponding to the first voice instruction, thereby achieving simultaneous recognition of the associated first voice information and second voice information, thereby avoiding the problem of inaccurate voice recognition caused by only recognizing the second voice information when the second voice information is associated with the first voice information, improving the accuracy of voice recognition, and thereby improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A schematic diagram of the hardware structure of a terminal for implementing each embodiment of the present invention;
[0041] Figure 2 A communication network system architecture diagram provided for an embodiment of the present invention;
[0042] Figure 3 It is a flowchart of the first embodiment of the speech processing method of the present invention;
[0043] Figure 4 It is a flow chart of the second embodiment of the speech processing method of the present invention;
[0044] Figure 5 A detailed flow chart of the step of determining whether the first voice information matches the second voice information in the third embodiment of the voice processing method of the present invention;
[0045] Figure 6 A detailed flow chart of the step of determining the sentence vector corresponding to the text information in the fourth embodiment of the speech processing method of the present invention;
[0046] Figure 7A detailed flowchart of the step of determining whether the first voice information matches the second voice information in the fifth embodiment of the voice processing method of the present invention is as follows;
[0047] Figure 8 A detailed flow chart of the step of determining whether the first identity information corresponding to the first voice information and the second identity information corresponding to the second voice information are the same based on the first voiceprint feature vector and the second voiceprint feature vector in the sixth embodiment of the voice processing method of the present invention;
[0048] Fig. 9 This is a detailed flowchart diagram of the step of determining whether second voice information is detected within a preset time interval after the current moment when first voice information is detected in the seventh embodiment of the voice processing method of the present invention.
[0049] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0050] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0051] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0052] In the subsequent description, the suffixes such as "module", "component" or "unit" used to represent elements are only used to facilitate the description of the present invention, and have no specific meanings. Therefore, "module", "component" or "unit" can be used in a mixed manner.
[0053] The terminal may be implemented in various forms. For example, the terminal described in the present invention may include mobile terminals such as mobile phones, tablet computers, laptop computers, PDAs, portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.
[0054] The following description will be made by taking a mobile terminal as an example, and those skilled in the art will appreciate that, in addition to components specifically used for mobile purposes, the configuration according to the embodiments of the present invention can also be applied to fixed-type terminals.
[0055] See also Figure 1, which is a schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present invention, the mobile terminal 100 may include: RF (Radio Frequency, radio frequency) unit 101, Wi-Fi module 102, audio output unit 103, A / V (audio / video) input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, processor 110, and power supply 111 and other components. Those skilled in the art can understand that Figure 1 The structure of the mobile terminal 100 shown in the figure does not constitute a limitation of the mobile terminal 100. The mobile terminal 100 may include more or less components than those shown in the figure, or combine certain components, or arrange the components differently.
[0056] Combine the following Figure 1 The various components of the mobile terminal 100 are introduced in detail:
[0057] The radio frequency unit 101 can be used for receiving and sending signals during information transmission or communication. Specifically, after receiving the downlink information of the base station, it is sent to the processor 110 for processing; in addition, the uplink data is sent to the base station. Generally, the radio frequency unit 101 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc. In addition, the radio frequency unit 101 can also communicate with the network and other devices through wireless communication. The above-mentioned wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution) and TDD-LTE (Time Division Duplexing-Long Term Evolution), etc.
[0058] Wi-Fi is a short-range wireless transmission technology. The mobile terminal 100 can help users send and receive emails, browse web pages, and access streaming media through the Wi-Fi module 102. It provides users with wireless broadband Internet access. Figure 1 The Wi-Fi module 102 is shown, but it is understandable that it is not an essential component of the mobile terminal 100 and can be omitted as required without changing the essence of the invention.
[0059] The audio output unit 103 can convert the audio data received by the RF unit 101 or the Wi-Fi module 102 or stored in the memory 109 into an audio signal and output it as sound when the mobile terminal 100 is in a call signal reception mode, a talk mode, a recording mode, a voice recognition mode, a broadcast reception mode, etc. Moreover, the audio output unit 103 can also provide audio output related to a specific function performed by the mobile terminal 100 (for example, a call signal reception sound, a message reception sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.
[0060] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processor (GPU) 1041 and a microphone 1042, and the graphics processor 1041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The processed image frame can be displayed on the display unit 106. The image frame processed by the graphics processor 1041 can be stored in the memory 109 (or other storage medium) or sent via the radio frequency unit 101 or the Wi-Fi module 102. The microphone 1042 can receive sound (audio data) via the microphone 1042 in a telephone call mode, a recording mode, a voice recognition mode, and other operating modes, and can process such sound into audio data. The processed audio (voice) data can be converted into a format output that can be sent to a mobile communication base station via the radio frequency unit 101 in the case of a telephone call mode. The microphone 1042 can implement various types of noise elimination (or suppression) algorithms to eliminate (or suppress) noise or interference generated in the process of receiving and sending audio signals.
[0061] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 1061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1061 and / or the backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can also be configured on the mobile phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be repeated here.
[0062] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0063] The user input unit 107 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile terminal. Specifically, the user input unit 107 may include a touch panel 1071 and other input devices 1072. The touch panel 1071, also known as a touch screen, can collect the user's touch operation on or near it (such as the user's operation on the touch panel 1071 or near the touch panel 1071 using any suitable object or accessory such as a finger, stylus, etc.), and drive the corresponding connection device according to a pre-set program. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 110, and can receive and execute the command sent by the processor 110. In addition, the touch panel 1071 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic waves. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Specifically, the other input devices 1072 may include, but are not limited to, one or more of a physical keyboard, a function key (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, etc., which are not specifically limited here.
[0064] Furthermore, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of the touch event. Then, the processor 110 provides a corresponding visual output on the display panel 1061 according to the type of the touch event. Figure 1 In the figure, the touch panel 1071 and the display panel 1061 are used as two independent components to implement the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to implement the input and output functions of the mobile terminal, which is not limited here.
[0065] The interface unit 108 serves as an interface through which at least one external device can be connected to the mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 108 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the mobile terminal 100 or may be used to transmit data between the mobile terminal 100 and an external device.
[0066] The memory 109 can be used to store software programs and various data. The memory 109 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 109 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0067] The processor 110 is the control center of the mobile terminal 100. It uses various interfaces and lines to connect various parts of the entire mobile terminal. It executes various functions of the mobile terminal 100 and processes data by running or executing software programs and / or modules stored in the memory 109, and calling data stored in the memory 109, so as to monitor the mobile terminal 100 as a whole. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 110.
[0068] In addition, Figure 1 In the mobile terminal shown, the processor 110 is used to call the voice processing program stored in the memory 109 and perform the following operations:
[0069] When the first voice information is detected, determining whether the second voice information is detected within a preset time interval after the current moment;
[0070] If so, determining whether the first voice information matches the second voice information;
[0071] When the first voice information matches the second voice information, performing voice recognition on the first voice information and the second voice information to obtain a first voice recognition result;
[0072] Determine a first voice instruction corresponding to the first voice recognition result, and execute an operation corresponding to the first voice instruction.
[0073] Further, the processor 110 may call the speech processing program stored in the memory 109 and perform the following operations:
[0074] When the first voice information does not match the second voice information, performing voice recognition on the second voice information to obtain a second voice recognition result;
[0075] Obtain a second voice instruction corresponding to the second voice recognition result, and execute an operation corresponding to the second voice instruction.
[0076] Further, the processor 110 may call the speech processing program stored in the memory 109 and perform the following operations:
[0077] Acquire text information corresponding to the first voice information and the second voice information;
[0078] Determine a sentence vector corresponding to the text information;
[0079] Calculating a first similarity between the sentence vector and each preset sentence vector in a preset database;
[0080] Determine whether the first voice information matches the second voice information based on the first similarity.
[0081] Further, the processor 110 may call the speech processing program stored in the memory 109 and perform the following operations:
[0082] Determine a maximum similarity among the first similarities, and judge whether the maximum similarity is greater than a first preset similarity, wherein when the maximum similarity is greater than the first preset similarity, it is determined that the first voice information matches the second voice information.
[0083] Further, the processor 110 may call the speech processing program stored in the memory 109 and perform the following operations:
[0084] Training the text information based on a word vector model to obtain a word vector corresponding to the text information;
[0085] Calculate the similarity between the word vector in the matching sentence vector and the word vector, and generate a similar word matrix based on the similarity, wherein the elements of each row in the similar word matrix are the similarities between the same word vector and the word vector in the matching sentence vector;
[0086] The sentence vector is generated based on the maximum similarity in each column element of the similar word matrix.
[0087] Further, the processor 110 may call the speech processing program stored in the memory 109 and perform the following operations:
[0088] If yes, based on the preset time window, the first voice information and the second voice information are sampled respectively at a preset frequency to obtain first sampled data corresponding to the first voice information and second sampled data corresponding to the second voice information;
[0089] generating a first voiceprint feature vector based on the first sampling data, and generating a second voiceprint feature vector based on the second sampling data;
[0090] Based on the first voiceprint feature vector and the second voiceprint feature vector, determining whether the first identity information corresponding to the first voice information is the same as the second identity information corresponding to the second voice information;
[0091] If so, determine whether the first voice information matches the second voice information.
[0092] Further, the processor 110 may call the speech processing program stored in the memory 109 and perform the following operations:
[0093] Calculating a second similarity between the first timbre feature vector and the second timbre feature vector;
[0094] Determine whether the second similarity is greater than a second preset similarity, wherein if so, it is determined that the first identity information is the same as the second identity information.
[0095] Further, the processor 110 may call the speech processing program stored in the memory 109 and perform the following operations:
[0096] When the first voice information is detected, performing voice recognition on the first voice information to obtain a third voice recognition result;
[0097] Obtaining a third voice instruction corresponding to the third voice recognition result, and executing an operation corresponding to the third voice instruction;
[0098] Determine whether second voice information is detected within a preset time interval after the moment when the first voice information is detected.
[0099] The mobile terminal 100 may also include a power supply 111 (such as a battery) for supplying power to various components. Preferably, the power supply 111 may be logically connected to the processor 110 via a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system.
[0100] although Figure 1 Not shown, the mobile terminal 100 may further include a Bluetooth module, etc., which will not be described in detail here.
[0101] To facilitate understanding of the embodiments of the present invention, the communication network system on which the mobile terminal of the present invention is based is described below.
[0102] See also Figure 2 , Figure 2 A communication network system architecture diagram is provided for an embodiment of the present invention. The communication network system is an LTE system of universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203 and an operator's IP service 204, which are sequentially connected for communication.
[0103] Specifically, UE201 may be the above-mentioned mobile terminal 100, which will not be described in detail here.
[0104] E-UTRAN 202 includes eNodeB 2021 and other eNodeBs 2022 , etc. Among them, eNodeB 2021 can be connected to other eNodeBs 2022 through a backhaul (eg, an X2 interface), and eNodeB 2021 is connected to EPC 203 , and eNodeB 2021 can provide UE 201 with access to EPC 203 .
[0105] EPC203 may include MME (Mobility Management Entity) 2031, HSS (Home Subscriber Server) 2032, other MMEs 2033, SGW (Serving Gate Way) 2034, PGW (PDN Gate Way) 2035 and PCRF (Policy and Charging Rules Function) 2036, etc. Among them, MME2031 is a control node that processes signaling between UE201 and EPC203, and provides bearer and connection management. HSS2032 is used to provide some registers to manage functions such as home location register (not shown in the figure), and save some user-specific information such as service features and data rates. All user data can be sent through SGW2034, PGW2035 can provide IP address allocation and other functions for UE 201, PCRF2036 is the policy and charging control policy decision point for service data flow and IP bearer resources, which selects and provides available policy and charging control decisions for the policy and charging execution function unit (not shown in the figure).
[0106] The IP service 204 may include the Internet, an intranet, an IMS (IP Multimedia Subsystem) or other IP services.
[0107] Although the above introduction takes the LTE system as an example, those skilled in the art should know that the present invention is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA and future new network systems, etc., which are not limited here.
[0108] Based on the above terminal hardware structure and communication network system, various embodiments of the speech processing method of the present invention are proposed.
[0109] The present invention also provides a speech processing method, referring to Figure 3 , Figure 3 FIG. 4 is a flow chart of a first embodiment of a speech processing method according to the present invention.
[0110] In this embodiment, the speech processing method includes the following steps:
[0111] Step S100, when the first voice information is detected, determining whether the second voice information is detected within a preset time interval after the current moment;
[0112] In this embodiment, the user can perform a recording operation by turning on the recording function of the mobile terminal to monitor the voice information in the environment where the mobile terminal is located in real time, and then obtain the first voice information and the second voice information. Alternatively, the user can use professional equipment to record and upload the voice information obtained by the recording to the mobile terminal.
[0113] When detecting voice information obtained by a recording operation, or when receiving voice information sent by other terminals or devices, the mobile terminal determines that the first voice information is detected, and the mobile terminal starts a timing operation to determine whether a second voice information is detected within a preset time interval after the current moment.
[0114] Among them, when the first voice information is detected, if the recording function is not started in the mobile terminal, the recording function is turned on to record to obtain the second voice information. Alternatively, the user can turn on or off the recording function of the mobile terminal by himself. For example, the user turns on the recording function, the mobile terminal records to obtain the first voice information, and then turns off the recording function according to the user's instructions, and when the user records again, the recording function is turned on according to the user's instructions, and the mobile terminal records to obtain other voice information. If the time interval between the other voice information and the first voice information is less than the preset time interval, the other voice information is the second voice information.
[0115] The preset time interval may be reasonably set, for example, the preset time interval may be set to 30S.
[0116] Step S200, if yes, determining whether the first voice information matches the second voice information;
[0117] In this embodiment, if a second voice message is detected within a preset time interval after the current moment, it is determined whether the first voice message and the second voice message match. Specifically, it can be determined whether the first voice message and the second voice message are the voices of the same user, or whether the language logic of the first voice message and the second voice message is similar, and then it is determined whether the first voice message and the second voice message match.
[0118] Step S300, when the first voice information matches the second voice information, performing voice recognition on the first voice information and the second voice information to obtain a first voice recognition result;
[0119] In this embodiment, when the first voice information matches the second voice information, voice recognition is performed on the first voice information and the second voice information to obtain a first voice recognition result. Specifically, the first voice information and the second voice information are voice recognized as the same voice information, thereby avoiding the problem of inaccurate voice recognition caused by only recognizing the second voice information when the second voice information is associated with the first voice information.
[0120] Step S400: determine a first voice instruction corresponding to the first voice recognition result, and execute an operation corresponding to the first voice instruction.
[0121] In this embodiment, after obtaining the first voice recognition result, a first voice instruction corresponding to the first voice recognition result is determined, and an operation corresponding to the first voice instruction is executed.
[0122] For example, when a user says the name of a star, the mobile terminal recognizes the voice containing the star's name, and based on the recognition result, searches and outputs the star's recent news or personal information, etc. When the user subsequently inputs voice input such as "Is she beautiful (is he handsome)?", the mobile terminal can directly output information such as "beautiful (handsome)".
[0123] The voice processing method proposed in this embodiment determines whether the second voice information is detected within a preset time interval after the current moment when the first voice information is detected, and then, if so, determines whether the first voice information matches the second voice information, and then, when the first voice information matches the second voice information, performs voice recognition on the first voice information and the second voice information to obtain a first voice recognition result, and then determines the first voice instruction corresponding to the first voice recognition result, and executes the operation corresponding to the first voice instruction, thereby realizing the simultaneous recognition of the associated first voice information and the second voice information, thereby avoiding the problem of inaccurate voice recognition caused by only recognizing the second voice information when the second voice information is associated with the first voice information, improving the accuracy of voice recognition, and thereby improving user experience.
[0124] Based on the first embodiment, a second embodiment of the speech processing method of the present invention is proposed. Figure 4 In this embodiment, after step S200, the speech processing method further includes:
[0125] Step S500, when the first voice information does not match the second voice information, performing voice recognition on the second voice information to obtain a second voice recognition result;
[0126] Step S600, obtaining a second voice instruction corresponding to the second voice recognition result, and executing an operation corresponding to the second voice instruction.
[0127] In this embodiment, when the first voice information and the second voice information do not match, for example, the first voice information and the second voice information are the voices of users of different identities, and / or the language logic of the first voice information and the second voice information is different (the first voice information and the second voice information cannot be expressed coherently), voice recognition is performed on the second voice information to obtain a second voice recognition result, and a second voice instruction corresponding to the second voice recognition result is obtained to execute the operation corresponding to the second voice instruction.
[0128] The voice processing method proposed in this embodiment performs voice recognition on the second voice information to obtain a second voice recognition result when the first voice information does not match the second voice information, then obtains a second voice instruction corresponding to the second voice recognition result, and executes an operation corresponding to the second voice instruction. When the second voice information is not associated with the first voice information, only the second voice information can be recognized, thereby further improving the accuracy of voice recognition.
[0129] Based on the first embodiment, a third embodiment of the speech processing method of the present invention is proposed, referring to Figure 5 In this embodiment, step S200 includes:
[0130] Step S210, obtaining text information corresponding to the first voice information and the second voice information;
[0131] Step S220, determining a sentence vector corresponding to the text information;
[0132] Step S230, calculating a second similarity between the sentence vector and each preset sentence vector in a preset database;
[0133] Step S240: Determine whether the first voice information matches the second voice information based on the second similarity.
[0134] In this embodiment, if a second voice information is detected within a preset time interval after the current moment, voice recognition is performed on the first voice information and the second voice information to obtain text information corresponding to the first voice information and the second voice information, that is, the text information includes the content of the first voice information and the second voice information, and the sentence vector corresponding to the text information is determined.
[0135] Among them, when the sentence vector corresponding to the text information is obtained, the first similarity between the sentence vector and each preset sentence vector in the preset database is calculated. Specifically, each preset sentence vector is traversed, and the similarity between the currently traversed preset sentence vector and the sentence vector corresponding to the text information is calculated until the preset sentence vector traversal is completed, and the first similarity is obtained, wherein the first similarity is the cosine value between the sentence vector and the preset sentence vector, which is specifically calculated using the cosine formula.
[0136] When the first similarity is obtained, it is determined whether the first voice information matches the second voice information based on the first similarity. For example, it is determined whether there is a similarity in the first similarity that is greater than the first preset similarity, or it is determined whether the maximum similarity in the first similarity is greater than the first preset similarity. If so, it is determined that the first voice information matches the second voice information.
[0137] Further, step S240 includes: determining the maximum similarity among the first similarities, and judging whether the maximum similarity is greater than a first preset similarity, wherein when the maximum similarity is greater than the first preset similarity, it is determined that the first voice information matches the second voice information.
[0138] Specifically, each similarity in the first similarities is compared to determine the maximum similarity in the first similarities, and it is determined whether the maximum similarity is greater than a first preset similarity. If so, it is determined that the first voice information matches the second voice information.
[0139] The first preset similarity can be reasonably set, for example, the first preset similarity is 80%.
[0140] Among them, each preset sentence vector in the preset database is a sentence vector of various sentences that conform to language logic, and the various sentences that conform to language logic include at least two sentences.
[0141] Further, in one embodiment, the step S230 includes: traversing the preset sentence vector to obtain the language information corresponding to the currently traversed preset sentence vector; traversing each word in the text information, calculating the Tf value and idf value between the currently traversed word and the language information, and calculating the product of the Tf value and the idf value; when the traversal of each word in the text information is completed, calculating the average of the products of the Tf value and the idf value corresponding to each word to obtain the Tf-idf value between the text information and the language information; when the traversal of the preset sentence vector is completed, the preset sentence vector corresponding to the largest Tf-idf value of a preset number of each Tf-idf value is used as the target sentence vector; calculating the first similarity between the sentence vector and the target sentence vector.
[0142] It should be noted that the Tf value refers to the frequency of a given word appearing in the file. The Tf value is the number of times the word appears in the preset database divided by the sum of the number of times all words appear in the preset database.
[0143] The speech processing method proposed in this embodiment obtains text information corresponding to the first speech information and the second speech information; then determines the sentence vector corresponding to the text information, and then determines whether the first speech information matches the second speech information based on the first similarity between the sentence vector and each preset sentence vector in a preset database, and then determines whether the first speech information matches the second speech information based on the first similarity. This realizes determining whether the first speech information matches the second speech information based on the first similarity, and then accurately judging whether the two voices have similar language logic based on the similarity, thereby further improving the accuracy of speech recognition.
[0144] Based on the third embodiment, a fourth embodiment of the speech processing method of the present invention is proposed, referring to Figure 6 In this embodiment, step S220 includes:
[0145] Step S221, training the text information based on a word vector model to obtain a word vector corresponding to the text information;
[0146] Step S222, calculating the similarity between the word vector in the matching sentence vector and the word vector, and generating a similar word matrix based on the similarity, wherein the elements of each row in the similar word matrix are the similarities between the same word vector and the word vector in the matching sentence vector;
[0147] Step S223, generating the sentence vector based on the maximum similarity in each column element of the similar word matrix.
[0148] In this embodiment, when obtaining text information, the text information is trained based on a word vector model to obtain a word vector corresponding to the text information, wherein the word vector is a vector corresponding to each word in the text information.
[0149] When determining the word vector of the text information, a matching sentence vector is obtained, wherein the matching sentence vector is a sentence vector composed of pre-set words, and the elements of the matching sentence vector are the word vectors of the words in the vocabulary. The matching sentence vector is an M-dimensional vector, and M is the length of the vocabulary, that is, the number of words in the vocabulary. For example, M is 100,000, which is the number of words corresponding to the matching sentence vector, wherein the words in the vocabulary are all words that may appear in the text information.
[0150] When the matching sentence vector is obtained, the similarity between the word vector in the matching sentence vector and the word vector is calculated, and a similar word matrix is generated based on the similarity, wherein the elements of each row in the similar word matrix are the similarity between the same word vector and the word vector in the matching sentence vector. The similar word matrix is an M*N matrix, wherein M is the length of the vocabulary, and N is the number of similar words, i.e., the number of words in the text information. Then, based on the maximum similarity in each column element of the similar word matrix, the sentence vector is generated. When the similar word matrix of the text information is obtained, the elements of each column in the similar word matrix are compared respectively to determine the maximum similarity of each column element, and the maximum similarity of each column is used as an element of a one-dimensional vector, which is the sentence vector of the text information.
[0151] The speech processing method proposed in this embodiment trains the text information based on a word vector model to obtain the word vector corresponding to the text information, then calculates the similarity between the word vector in the matching sentence vector and the word vector, generates a similar word matrix based on the similarity, and then generates the sentence vector based on the maximum similarity in each column element of the similar word matrix. The sentence vector of the text information can be accurately obtained according to the similar word matrix, thereby further improving the accuracy of speech recognition.
[0152] Based on the first embodiment, a fifth embodiment of the speech processing method of the present invention is proposed, referring to Figure 7 In this embodiment, step S200 includes:
[0153] Step S250: If yes, then based on the preset time window, the first voice information and the second voice information are sampled at a preset frequency to obtain first sampled data corresponding to the first voice information and second sampled data corresponding to the second voice information;
[0154] Step S260, generating a first voiceprint feature vector based on the first sampling data, and generating a second voiceprint feature vector based on the second sampling data;
[0155] Step S270, determining whether the first identity information corresponding to the first voice information is the same as the second identity information corresponding to the second voice information based on the first voiceprint feature vector and the second voiceprint feature vector;
[0156] Step S280: If yes, determine whether the first voice information matches the second voice information.
[0157] In this embodiment, if the second voice information is detected within a preset time interval after the current moment, the first voice information and the second voice information are sampled based on the preset time window and at a preset frequency to obtain first sampling data corresponding to the first voice information and second sampling data corresponding to the second voice information, and a first voiceprint feature vector is generated based on the first sampling data, and a second voiceprint feature vector is generated based on the second sampling data. Specifically, the first voice information and the second voice information are first windowed according to the preset time window to obtain the first voice information within the preset time window and the second voice information within the preset time window, and the first voice information within the preset time window and the second voice information within the preset time window are sampled according to a preset frequency (for example, 8KHz) to obtain multiple first sampling data and second sampling data, and a voiceprint feature vector is generated based on the first sampling data and the second sampling data, that is, each sampling point data is taken as an element of a vector to obtain the voiceprint feature vector.
[0158] Based on the first voiceprint feature vector and the second voiceprint feature vector, determine whether the first identity information corresponding to the first voice information is the same as the second identity information corresponding to the second voice information. Specifically, calculate the similarity between the first voiceprint feature vector and the second voiceprint feature vector. When the similarity is greater than a second preset similarity, determine that the first voice information is the same as the second voice information.
[0159] It should be noted that since the speech signal is a short-term stable signal and a long-term non-stationary signal, its long-term non-stationary characteristics are caused by the change in the physical movement process of the vocal organs. However, the movement of the vocal organs has a certain inertia, so in a short period of time, the speech signal is similar to a stable signal. The short time generally ranges from 10 to 30 milliseconds. Therefore, the preset time window can be set to a time window of 15-20 milliseconds.
[0160] The speech processing method proposed in this embodiment, if yes, then based on a preset time window, the first voice information and the second voice information are sampled respectively according to a preset frequency to obtain first sampling data corresponding to the first voice information and second sampling data corresponding to the second voice information; then a first voiceprint feature vector is generated based on the first sampling data, and a second voiceprint feature vector is generated based on the second sampling data, and then based on the first voiceprint feature vector and the second voiceprint feature vector, it is determined whether the first identity information corresponding to the first voice information and the second identity information corresponding to the second voice information are the same; then if yes, it is determined whether the first voice information matches the second voice information, by first determining the user identity corresponding to the first voice information and the second voice information, and when the user identity corresponding to the first voice information and the second voice information are the same, determining whether the first voice information matches the second voice information, thereby ensuring subsequent processing of the voice information of the same user, and further improving the accuracy of speech recognition.
[0161] Based on the fifth embodiment, a sixth embodiment of the speech processing method of the present invention is proposed, referring to Figure 8 In this embodiment, the first voiceprint feature vector is a first timbre feature vector, and the second voiceprint feature vector is a second timbre feature vector. Step S270 includes:
[0162] Step S271, calculating a second similarity between the first timbre feature vector and the second timbre feature vector;
[0163] Step S272, determining whether the second similarity is greater than a second preset similarity, wherein if it is greater, determining that the first identity information is the same as the second identity information.
[0164] Since timbre is the attribute that best reflects a person's identity information, when a person is in a bad mood, the loudness and pitch of the voice information will obviously decrease compared to when the person is in a good mood, and when the person is in a good mood, the loudness and pitch of the voice information will obviously increase, while the timbre of the voice information will not change significantly when the person is in different moods. Therefore, the timbre feature vector is used to determine the identity of the user.
[0165] In this embodiment, a second similarity between the first timbre feature vector and the second timbre feature vector is calculated, and the second similarity is the cosine value between the first timbre feature vector and the second timbre feature vector; and it is determined whether the second similarity is greater than a second preset similarity, wherein if it is greater, it is determined that the first identity information is the same as the second identity information.
[0166] The speech processing method proposed in this embodiment calculates the second similarity between the first timbre feature vector and the second timbre feature vector; then determines whether the second similarity is greater than a second preset similarity, wherein if greater than, it is determined that the first identity information is the same as the second identity information, thereby accurately determining whether the first identity information is the same as the second identity information based on the timbre feature vector, and further improving the accuracy of speech recognition.
[0167] Based on the above embodiments, a seventh embodiment of the speech processing method of the present invention is proposed, referring to Fig. 9 In this embodiment, step S100 includes:
[0168] Step S110, when the first voice information is detected, performing voice recognition on the first voice information to obtain a third voice recognition result;
[0169] Step S120, obtaining a third voice instruction corresponding to the third voice recognition result, and executing an operation corresponding to the third voice instruction;
[0170] Step S130, determining whether second voice information is detected within a preset time interval after the moment when the first voice information is detected.
[0171] In this embodiment, when the first voice information is detected, voice recognition is performed on the first voice information to obtain a third voice recognition result, and then a third voice instruction corresponding to the third voice recognition result is obtained, and the operation corresponding to the third voice instruction is executed to complete the operation corresponding to the first voice information.
[0172] The voice processing method proposed in this embodiment performs voice recognition on the first voice information to obtain a third voice recognition result when the first voice information is detected; then obtains a third voice instruction corresponding to the third voice recognition result, executes an operation corresponding to the third voice instruction, and then determines whether the second voice information is detected within a preset time interval after the moment when the first voice information is detected. By completing the operation corresponding to the first voice information first, the user experience is further improved.
[0173] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a speech processing program is stored. When the speech processing program is executed by a processor, the following operations are implemented:
[0174] When the first voice information is detected, determining whether the second voice information is detected within a preset time interval after the current moment;
[0175] If so, determining whether the first voice information matches the second voice information;
[0176] When the first voice information matches the second voice information, performing voice recognition on the first voice information and the second voice information to obtain a first voice recognition result;
[0177] Determine a first voice instruction corresponding to the first voice recognition result, and execute an operation corresponding to the first voice instruction.
[0178] Furthermore, when the speech processing program is executed by the processor, the following operations are also implemented:
[0179] When the first voice information does not match the second voice information, performing voice recognition on the second voice information to obtain a second voice recognition result;
[0180] Obtain a second voice instruction corresponding to the second voice recognition result, and execute an operation corresponding to the second voice instruction.
[0181] Furthermore, when the speech processing program is executed by the processor, the following operations are also implemented:
[0182] Acquire text information corresponding to the first voice information and the second voice information;
[0183] Determine a sentence vector corresponding to the text information;
[0184] Calculating a first similarity between the sentence vector and each preset sentence vector in a preset database;
[0185] Determine whether the first voice information matches the second voice information based on the first similarity.
[0186] Furthermore, when the speech processing program is executed by the processor, the following operations are also implemented:
[0187] Determine a maximum similarity among the first similarities, and judge whether the maximum similarity is greater than a first preset similarity, wherein when the maximum similarity is greater than the first preset similarity, it is determined that the first voice information matches the second voice information.
[0188] Furthermore, when the speech processing program is executed by the processor, the following operations are also implemented:
[0189] Training the text information based on a word vector model to obtain a word vector corresponding to the text information;
[0190] Calculate the similarity between the word vector in the matching sentence vector and the word vector, and generate a similar word matrix based on the similarity, wherein the elements of each row in the similar word matrix are the similarities between the same word vector and the word vector in the matching sentence vector;
[0191] The sentence vector is generated based on the maximum similarity in each column element of the similar word matrix.
[0192] Furthermore, when the speech processing program is executed by the processor, the following operations are also implemented:
[0193] If yes, based on the preset time window, the first voice information and the second voice information are sampled respectively at a preset frequency to obtain first sampled data corresponding to the first voice information and second sampled data corresponding to the second voice information;
[0194] generating a first voiceprint feature vector based on the first sampling data, and generating a second voiceprint feature vector based on the second sampling data;
[0195] Based on the first voiceprint feature vector and the second voiceprint feature vector, determining whether the first identity information corresponding to the first voice information is the same as the second identity information corresponding to the second voice information;
[0196] If so, determine whether the first voice information matches the second voice information.
[0197] Furthermore, when the speech processing program is executed by the processor, the following operations are also implemented:
[0198] Calculating a second similarity between the first timbre feature vector and the second timbre feature vector;
[0199] Determine whether the second similarity is greater than a second preset similarity, wherein if so, it is determined that the first identity information is the same as the second identity information.
[0200] Furthermore, when the speech processing program is executed by the processor, the following operations are also implemented:
[0201] When the first voice information is detected, performing voice recognition on the first voice information to obtain a third voice recognition result;
[0202] Obtaining a third voice instruction corresponding to the third voice recognition result, and executing an operation corresponding to the third voice instruction;
[0203] Determine whether second voice information is detected within a preset time interval after the moment when the first voice information is detected.
[0204] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0205] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0206] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0207] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A speech processing method, characterized in that: The speech processing method comprises the following steps: When the first voice information is detected, determining whether the second voice information is detected within a preset time interval after the current moment; If so, determining whether the first voice information matches the second voice information; When the first voice information matches the second voice information, performing voice recognition on the first voice information and the second voice information to obtain a first voice recognition result; Determining a first voice instruction corresponding to the first voice recognition result, and executing an operation corresponding to the first voice instruction; The step of determining whether the first voice information matches the second voice information comprises: Acquire text information corresponding to the first voice information and the second voice information; Determine a sentence vector corresponding to the text information; Calculating a first similarity between the sentence vector and each preset sentence vector in a preset database; Determining whether the first voice information matches the second voice information based on the first similarity; The step of determining whether the first voice information matches the second voice information based on the first similarity comprises: Determining a maximum similarity among the first similarities, and judging whether the maximum similarity is greater than a first preset similarity, wherein when the maximum similarity is greater than the first preset similarity, judging that the first voice information matches the second voice information; The step of determining the sentence vector corresponding to the text information comprises: Training the text information based on a word vector model to obtain a word vector corresponding to the text information; Calculate the similarity between the word vector in the matching sentence vector and the word vector, and generate a similar word matrix based on the similarity, wherein the elements of each row in the similar word matrix are the similarities between the same word vector and the word vector in the matching sentence vector; The sentence vector is generated based on the maximum similarity in each column element of the similar word matrix.
2. The speech processing method according to claim 1, characterized in that: After the step of determining whether the first voice information matches the second voice information, the voice processing method further includes: When the first voice information does not match the second voice information, performing voice recognition on the second voice information to obtain a second voice recognition result; Obtain a second voice instruction corresponding to the second voice recognition result, and execute an operation corresponding to the second voice instruction.
3. The speech processing method according to claim 1, characterized in that: If so, the step of determining whether the first voice information matches the second voice information includes: If yes, based on the preset time window, the first voice information and the second voice information are sampled respectively at a preset frequency to obtain first sampled data corresponding to the first voice information and second sampled data corresponding to the second voice information; generating a first voiceprint feature vector based on the first sampling data, and generating a second voiceprint feature vector based on the second sampling data; Based on the first voiceprint feature vector and the second voiceprint feature vector, determining whether the first identity information corresponding to the first voice information is the same as the second identity information corresponding to the second voice information; If so, determine whether the first voice information matches the second voice information.
4. The speech processing method according to claim 3, characterized in that: The first voiceprint feature vector is a first timbre feature vector, and the second voiceprint feature vector is a second timbre feature vector; and the step of determining whether the first identity information corresponding to the first voice information and the second identity information corresponding to the second voice information are the same based on the first voiceprint feature vector and the second voiceprint feature vector comprises: Calculating a second similarity between the first timbre feature vector and the second timbre feature vector; Determine whether the second similarity is greater than a second preset similarity, wherein if so, it is determined that the first identity information is the same as the second identity information.
5. The speech processing method according to any one of claims 1 to 4, characterized in that: The step of determining whether the second voice information is detected within a preset time interval after the current moment when the first voice information is detected includes: When the first voice information is detected, performing voice recognition on the first voice information to obtain a third voice recognition result; Obtaining a third voice instruction corresponding to the third voice recognition result, and executing an operation corresponding to the third voice instruction; Determine whether second voice information is detected within a preset time interval after the moment when the first voice information is detected.
6. A mobile terminal, characterized in that: The mobile terminal comprises: a memory, a processor, and a voice processing program stored in the memory and executable on the processor, wherein the voice processing program implements the steps of the voice processing method according to any one of claims 1 to 5 when executed by the processor.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a speech processing program, and when the speech processing program is executed by a processor, the steps of the speech processing method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Voice recognition method and electronic device
CN103594089A
Speech information receiving method and device and speech information analyzing method and device
CN108847236A