A speech recognition method and related products
By improving the speech recognition method, combining image information and speech feature extraction, electronic scales provide preliminary and confirmation results when identifying products, solving the problem of cumbersome selection of products in the prior art, and improving the recognition efficiency and accuracy.
Patent Information
- Application Number
- CN202011373516.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-11-30
AI Technical Summary
When identifying products, existing electronic scales need to manually select the correct products, which is cumbersome and inefficient in operation, especially when speech recognition is inaccurate, and the processing efficiency is inefficient.
By improving the speech recognition method, the user's voice signal is received and recognized, and the preliminary recognition result is provided and the confirmation command or second voice signal is waited for a confirmation command or a second voice signal. Combined with image information and voice feature extraction, the recognition accuracy is improved and the final result is output.
It simplifies user operation processes, improves the processing efficiency and accuracy of product identification, and reduces manual intervention, especially when ambient noise and a variety of products exist.
Smart Images

Figure CN114582334B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer application technologies, and particularly to a voice recognition method and related products. Background Art
[0002] An electronic scale is a tool for measuring the mass of an object by using Hooke's law or the principle of lever balance of forces, and is widely used in daily life scenarios such as supermarkets and farmers' markets. In the process of daily commodity transactions, in addition to measuring the weight of the commodity, the electronic scale also needs to identify the commodity to obtain the unit price of the commodity and get the price of the commodity.
[0003] In the method of identifying a commodity using voice information, the user says the name of the commodity, and the electronic scale receives the voice information containing the above commodity name; the name of the commodity is identified from the above voice information to obtain the unit price information of the commodity; the price of the commodity is obtained by combining the weight of the commodity. However, in the case where the electronic scale fails to identify the correct commodity, the user needs to manually select the correct commodity, which is cumbersome and has low processing efficiency. Summary of the Invention
[0004] The embodiments of the present application disclose a voice recognition method and related products, which improve the voice recognition method to simplify the user's operation process and improve the processing efficiency; at the same time, by extracting the user's voice features, voices with the above voice features can be specifically identified during the voice recognition process, thereby improving the accuracy of voice recognition.
[0005] In a first aspect, the embodiments of the present application disclose a voice recognition method, including:
[0006] Receiving a first voice signal for a target object;
[0007] Identifying the first voice signal to obtain a first recognition result;
[0008] Outputting the first recognition result; waiting for a first confirmation instruction input by the user or a second voice signal for the target object;
[0009] In the case of receiving the first confirmation instruction, taking the first recognition result as the recognition result of the first voice signal;
[0010] In the case of not receiving the first confirmation instruction and receiving a second voice signal for the target object, identifying the second voice signal to obtain a second recognition result;
[0011] Outputting the second recognition result; waiting for a second confirmation instruction input by the user;
[0012] In the case of receiving the above-mentioned second confirmation instruction, use the above-mentioned second recognition result as the recognition result of the above-mentioned first voice signal.
[0013] In yet another possible implementation manner of the first aspect, before receiving the first voice signal for the subject matter, the method further includes:
[0014] Obtain image information including the subject matter;
[0015] In the case where the types of the subject matter included in the above-mentioned image information are greater than or equal to two, output a prompt message indicating an error in the number of subject matter types.
[0016] In yet another possible implementation manner of the first aspect, the recognition of the above-mentioned second voice signal to obtain a second recognition result includes:
[0017] Recognize the above-mentioned second voice signal to obtain a third recognition result;
[0018] Use the name of the subject matter in the database that has the highest matching degree with the above-mentioned third recognition result as the second recognition result, and at least two subject matters are included in the database.
[0019] In yet another possible implementation manner of the first aspect, before recognizing the above-mentioned second voice signal to obtain a second recognition result, the method further includes:
[0020] Receive a third voice signal; extract feature parameters from the above-mentioned third voice signal;
[0021] The recognition of the above-mentioned second voice signal to obtain a second recognition result includes: recognizing the part of the above-mentioned second voice signal that is the same as the above-mentioned feature parameters to obtain a second recognition result.
[0022] In yet another possible implementation manner of the first aspect, the extraction of the above-mentioned feature parameters from the above-mentioned third voice signal includes:
[0023] Perform frame division processing on the above-mentioned third voice signal to obtain a set of voice frames;
[0024] Perform a fast Fourier transform on the voice frame to obtain a first frequency spectrum, where the voice frame is any voice frame in the above-mentioned set of voice frames;
[0025] Filter the above-mentioned first frequency spectrum with a Mel filter bank to obtain a second frequency spectrum;
[0026] Perform cepstrum analysis on the above-mentioned second frequency spectrum to obtain Mel frequency cepstral coefficients; use the above-mentioned Mel frequency cepstral coefficients as the feature parameters of the voice frame.
[0027] In yet another possible implementation of the first aspect, the above-mentioned output of the second recognition result includes:
[0028] Playing the above-mentioned second recognition result;
[0029] Or, displaying the above-mentioned second recognition result on a display screen.
[0030] In yet another possible implementation of the first aspect, after receiving the first voice signal for the subject matter, the method further includes:
[0031] Performing noise reduction processing on the above-mentioned first voice signal to obtain a noise-reduced voice signal;
[0032] The above-mentioned recognition of the above-mentioned first voice signal to obtain a first recognition result includes:
[0033] Recognizing the above-mentioned noise-reduced voice signal to obtain a first recognition result.
[0034] In a second aspect, an embodiment of the present application discloses a voice recognition device, including:
[0035] A receiving unit, configured to receive a first voice signal for a subject matter;
[0036] A recognition unit, configured to recognize the above-mentioned first voice signal to obtain a first recognition result;
[0037] An output unit, configured to output the above-mentioned first recognition result;
[0038] A determination unit, configured to, in the case of receiving the above-mentioned first confirmation instruction, use the above-mentioned first recognition result as the recognition result of the above-mentioned first voice signal;
[0039] The above-mentioned receiving unit is further configured to, after the above-mentioned output unit outputs the above-mentioned first recognition result, wait for a first confirmation instruction input by the user or a second voice signal for the above-mentioned subject matter;
[0040] The above-mentioned recognition unit is further configured to, in the case of not receiving the above-mentioned first confirmation instruction and receiving a second voice signal for the above-mentioned subject matter, recognize the above-mentioned second voice signal to obtain a second recognition result;
[0041] The above-mentioned output unit is further configured to output the above-mentioned second recognition result;
[0042] The above-mentioned receiving unit is further configured to, after the above-mentioned output unit outputs the above-mentioned second recognition result, wait for a second confirmation instruction input by the user;
[0043] The above-mentioned determination unit is further configured to, in the case of receiving the above-mentioned second confirmation instruction, use the above-mentioned second recognition result as the recognition result of the above-mentioned first voice signal.
[0044] In yet another possible implementation of the second aspect, the above-mentioned receiving unit is used to obtain image information including the subject matter;
[0045] The above-mentioned output unit is further used to output a prompt message indicating an error in the number of subject matter types when the number of subject matter types included in the above-mentioned image information is greater than or equal to two.
[0046] In yet another possible implementation of the second aspect, the above-mentioned recognition unit is further used to recognize the above-mentioned second voice signal to obtain a third recognition result;
[0047] The above-mentioned determination unit is further used to use the name of the subject matter with the highest matching degree with the above-mentioned third recognition result in the database as the second recognition result, and at least two subject matters are included in the database.
[0048] In yet another possible implementation of the second aspect, the above-mentioned device further includes:
[0049] The above-mentioned receiving unit is further used to receive a third voice signal;
[0050] An extraction unit is used to extract feature parameters from the above-mentioned third voice signal;
[0051] The above-mentioned recognition unit is further used to recognize the part of the above-mentioned second voice signal that is the same as the above-mentioned feature parameters to obtain a second recognition result.
[0052] In yet another possible implementation of the second aspect, the above-mentioned device further includes:
[0053] A framing unit is used to perform framing processing on the above-mentioned third voice signal to obtain a set of voice frames;
[0054] A transformation unit is used to perform a fast Fourier transform on the voice frame to obtain a first spectrum, and the voice frame is any one voice frame in the above-mentioned set of voice frames;
[0055] A filtering unit is used to filter the above-mentioned first spectrum with a Mel filter bank to obtain a second spectrum;
[0056] An inverse spectrum unit is used to perform inverse spectrum analysis on the above-mentioned second spectrum to obtain Mel frequency cepstral coefficients;
[0057] The above-mentioned determination unit is further used to use the above-mentioned Mel frequency cepstral coefficients as the feature parameters of the voice frame.
[0058] In yet another possible implementation of the second aspect, the above-mentioned output unit is further used to play the above-mentioned second recognition result;
[0059] The above output unit is further configured to display the above second recognition result on a display screen.
[0060] In another possible implementation manner of the second aspect, the above device further includes:
[0061] A noise reduction unit, configured to perform noise reduction processing on the above first voice signal to obtain a noise-reduced voice signal;
[0062] The above recognition unit is further configured to recognize the above noise-reduced voice signal to obtain a first recognition result.
[0063] In a third aspect, an embodiment of the present application discloses an electronic device, including: an input device, an output device, a processor, and a memory. Among them, a computer program is stored in the above memory, and the above processor calls the computer program stored in the above memory to execute the method in the first aspect or any one of the possible implementation manners of the first aspect.
[0064] In a fourth aspect, an embodiment of the present application discloses an electronic scale, including: a weight sensor, an input device, an output device, a processor, and a memory. Among them, a computer program is stored in the above memory, and the above processor calls the computer program stored in the above memory to execute the method in the first aspect or any one of the possible implementation manners of the first aspect.
[0065] In a fifth aspect, an embodiment of the present application discloses a computer-readable storage medium, in which a computer program is stored. When the above computer program runs on one or more processors, it executes the method in the first aspect or any one of the possible implementation manners of the first aspect.
[0066] In a sixth aspect, an embodiment of the present application discloses a computer program product, the above computer program product includes program instructions, and when the above program instructions are executed by a processor, the processor is caused to execute the method in the first aspect or any one of the possible implementation manners of the first aspect.
[0067] An embodiment of the present application discloses a voice recognition method and related products. By improving the voice recognition method, the operation process of the user is simplified, and the processing efficiency is improved; at the same time, by extracting the voice features of the user, voices with the above voice features can be specifically recognized during the voice recognition process, thereby improving the accuracy rate of voice recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the following will briefly introduce the drawings required to be used in the embodiments of the present application or the background technology.
[0069] Figure 1It is a schematic structural diagram of a speech recognition system disclosed in an embodiment of the present application;
[0070] Figure 2 It is a schematic flowchart of a speech recognition method disclosed in an embodiment of the present application;
[0071] Figure 3 It is a schematic flowchart of another speech recognition method disclosed in an embodiment of the present application;
[0072] Figure 4 It is a schematic flowchart of another speech recognition method disclosed in an embodiment of the present application;
[0073] Figure 5 It is a schematic structural diagram of a speech recognition device disclosed in an embodiment of the present application;
[0074] Figure 6 It is a schematic structural diagram of an electronic scale disclosed in an embodiment of the present application. Detailed implementation manners
[0075] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described below with reference to the accompanying drawings.
[0076] Terms such as "first" and "second" in the specification, claims and drawings of the present application are only used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device, etc. that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices, etc.
[0077] The phrase "in an embodiment" mentioned herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of the above phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art can explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0078] In this application, "at least one (item)" means one or more, "multiple" means two or more, "at least two (items)" means two, three or more, and "and / or" is used to describe the relationship between related objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (of the following)" or its similar expression means any combination of these items. For example, at least one (of) a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a, b, and c".
[0079] The main function of an electronic scale is to weigh items. However, as electronic scales are increasingly widely used in daily life, in addition to the weighing function, electronic scales often also need to have the function of identifying items. Currently, the methods for electronic scales on the market to identify commodities are mainly divided into the following several types:
[0080] 1. Stick barcodes on commodities in advance, and the barcodes contain relevant information such as the prices of the corresponding commodities. When weighing, scan the barcodes of the commodities, and the unit price information of the commodities can be obtained, thereby obtaining the prices of the commodities; however, the method of identifying commodities by scanning barcodes as described above will consume a large amount of labor costs and time costs, and moreover, some commodities (such as seafood) are not suitable for sticking barcodes in advance.
[0081] 2. Manually input the corresponding coding numbers of the commodities, and identify the corresponding commodities according to the coding numbers. However, for the above method of manually inputting codes, the user needs to remember the corresponding coding numbers of the commodities and cannot make mistakes during the process of inputting the codes. Therefore, the above method requires high labor costs and low efficiency.
[0082] 3. Use video to identify commodities. However, the accuracy rate of video recognition is low, and the accuracy rate of recognition will further decrease after the commodities are bagged.
[0083] 4. Use voice to identify commodities. In the above method, the user says the name of the commodity, and the electronic scale receives the voice information containing the above commodity name; identify the name of the commodity from the above voice information to obtain the unit price information of the commodity; combine the weight of the commodity to obtain the price of the commodity. However, in the case where the electronic scale fails to recognize the correct commodity, the user needs to manually select the correct commodity, which is cumbersome to operate and has low processing efficiency.
[0084] The embodiments of this application disclose a voice recognition method and related products. To more clearly describe the solution of this application, some knowledge related to voice recognition will be introduced first next.
[0085] Speech recognition is to convert a speech signal into corresponding text information. A speech recognition system mainly includes four major parts: a feature extraction module, an acoustic model module, a language model module, and a decoding module. Among them, in order to more effectively extract the features of the speech signal, it is often necessary to perform preprocessing operations such as filtering and framing on the speech signal to be processed before feature extraction; feature extraction mainly converts the speech signal to be processed from the time domain to the frequency domain to obtain a feature vector; the acoustic model mainly calculates the score of the above feature vector on the acoustic features; the language model mainly calculates the probability of the possible corresponding phrase sequence of the above speech signal to be processed according to relevant linguistic theories; finally, according to the existing dictionary, the above phrase sequence is decoded to obtain the final possible text representation, and the above possible text representation is the result of speech recognition of the above speech signal to be processed.
[0086] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the architecture of a speech recognition system disclosed in an embodiment of the present application. As Figure 1 shown, the above system includes:
[0087] A preprocessing module 101, mainly used for filtering and sampling, pre-emphasis processing, framing processing, windowing processing, and endpoint detection processing.
[0088] Filtering and sampling: The speech signal to be processed can be input into a band-pass filter with set upper and lower cut-off frequencies for filtering, and then the filtered speech signal is quantized; the above method can exclude signals with frequencies other than human vocalizations and the interference of the 50Hz power frequency. At the same time, smoothing processing can also be performed on the connection section between the high-frequency part and the low-frequency part; the above method can solve the spectrum under the same signal-to-noise ratio condition, making the analysis of the speech signal more convenient and fast.
[0089] Pre-emphasis processing: The speech signal to be processed is passed through a first-order finite impulse response high-pass filter to make the spectrum of the signal flat and not easily affected by the finite word length effect.
[0090] Framing processing: The speech signal can be assumed to be short-time stationary, that is, within a short period of time (such as 5 - 50ms), the speech signal remains basically unchanged; therefore, the speech signal can be segmented into speech frames and processed in units of frames; at the same time, the above speech frames are generally periodic, and processing each speech frame is equivalent to processing the original speech signal with fixed characteristics; the result of processing each frame can be regarded as a new time-dependent sequence, and the above sequence can be used to describe the features of the speech signal. For example, the speech frame length can be selected as 32ms and the frame overlap as 16ms for framing the speech signal, and the specific parameters can be adjusted according to the actual situation.
[0091] Windowing process: A Hamming window can be used to window the speech frames to reduce the influence of Gibbs effect.
[0092] Endpoint detection process: The start and end points of the speech signal can be determined by short-time energy (the amplitude of signal variation within the same frame) and short-time average zero-crossing rate (the number of times the sampled signal crosses zero within the same frame).
[0093] The acoustic feature extraction module 102 is mainly used to extract the feature parameters of the speech signal.
[0094] The feature parameters that can be used for speech recognition must meet the following requirements: the feature parameters can describe the fundamental features of the speech signal as much as possible; the coupling between parameter components should be minimized to compress the speech signal; the process of calculating the feature parameters is simple and the related algorithms are efficient. Linear prediction cepstral coefficients (LPCC) and Mel frequency cepstral coefficients (MFCC) are both typical feature parameters. In addition, parameters such as pitch period and resonance peak can also be used as feature parameters to characterize the speech characteristics.
[0095] Among them, the pitch period refers to the vibration period of the vocal cord vibration frequency (fundamental frequency), which can effectively characterize the speech signal; the resonance peak is the region where the energy of the speech signal is concentrated. Since the resonance peak characterizes the physical characteristics of the vocal tract and is the main determining condition of the pronunciation quality, the resonance peak is an important feature parameter.
[0096] The acoustic model module 103 is mainly used to distinguish different basic units and recognize the speech signal.
[0097] Speech recognition is essentially a pattern recognition process, and the core of pattern recognition is the problem of classifiers and classification decisions. Generally, in isolated word and medium / small vocabulary recognition, using a dynamic time warping (DTW) classifier will have good recognition results, with fast recognition speed and small system overhead, which is a very successful matching algorithm in speech recognition. However, in large vocabulary and speaker-independent speech recognition, the recognition effect of DTW will drop sharply. The hidden Markov model can be used to depict the internal sub-state changes of a phoneme to solve the problem of the correspondence between the feature sequence and multiple speech basic units.
[0098] The language model module 104 mainly serves as a reference. During the process of speech recognition and decoding, the intra-word transition refers to the pronunciation dictionary, and the inter-word transition refers to the language model. An N-gram language model (N-gram LM) can be used as the reference language model.
[0099] The decoding module 105 is mainly used to decode the speech signal to obtain the most likely word sequence.
[0100] The decoder is the core component in the recognition stage. It decodes the speech through a trained model to obtain the most likely word sequence, or generates a recognition grid based on the intermediate recognition results for subsequent components to process. The core algorithm of the decoder part is the Viterbi algorithm.
[0101] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0102] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a speech recognition method disclosed in the embodiments of the present application. As Figure 2 shown, the above method includes:
[0103] 201: Receive a first speech signal for the target object.
[0104] Wherein, the target object is any item that needs to be recognized by speech, and the present application does not make any restrictions.
[0105] Specifically, before receiving the first speech signal for the target object, an image information including the target object is obtained by using a video device. When the number of types of the target object included in the above image information is greater than or equal to two, a prompt message indicating an error in the number of types of the target object is output. It should be noted that the prompt message indicating an error in the number of types of the target object here does not mean outputting "error in the number of types" as the prompt message, but any information that can prompt the user that the type of the target object should be single. For example, outputting "Please check whether the types of items are single" or "Please determine a single-type item", etc., and the present application does not make any restrictions.
[0106] Specifically, after receiving the first speech signal for the target object, the above first speech signal is denoised to reduce the speech interference in the environment and improve the accuracy of speech recognition. The specific denoising method can be selected according to the actual situation, and the present application does not make any restrictions. For example, denoising can be performed through a noise-canceling microphone. Its basic principle is to assemble two microphones in the same hearing aid according to strict acoustic principles, so that the sound signals arriving at different angles are amplified differently, thereby enhancing the useful signals and relatively reducing the background noise.
[0107] 202: Identify the above first voice signal to obtain a first recognition result.
[0108] Specifically, the steps and related principles for identifying the first voice signal can refer to the above description of the voice recognition system and will not be elaborated here.
[0109] 203: Output the above first recognition result and wait for the first confirmation instruction input by the user or the second voice signal for the above subject matter.
[0110] Among them, the first recognition result can be output to the user by voice playback or by display on a monitor.
[0111] It should be noted that regardless of whether the recognition result of the first voice signal is correct, waiting will be performed: if the above first recognition result is correct, wait to receive the confirmation instruction input by the user; if the above first recognition result is incorrect, wait to receive the second voice signal for the above subject matter.
[0112] 204: In the case of receiving the above first confirmation instruction, use the above first recognition result as the recognition result of the above first voice signal.
[0113] 205: In the case of not receiving the above first confirmation instruction and receiving the second voice signal for the above subject matter, identify the above second voice signal to obtain a second recognition result.
[0114] Among them, after identifying the first voice signal to obtain a first recognition result, in the case where the above first recognition result is incorrect, the confirmation instruction for the above first result will not be received, but the second voice signal for the above subject matter will be received, and then the above second voice signal will be identified to obtain a second recognition result.
[0115] In particular, in addition to directly identifying the above second voice signal to obtain a second recognition result, the above second voice signal can also be identified first to obtain a third recognition result, where the third recognition result is the result of directly identifying the above second voice signal; then match the above third recognition result with the name of the subject matter in the database, and use the name of the subject matter in the database with the highest matching degree with the above third recognition result as the recognition result of the above second voice signal. The above method can improve the efficiency of voice recognition.
[0116] Among them, the information in the above database is input through an input device before speech recognition. It can be input through a keyboard or through speech input, and this application does not impose any restrictions. At the same time, the input information can include other information of the subject matter in addition to the name of the subject matter, such as the price of the subject matter, the storage method of the subject matter, the usage method of the subject matter, etc., which can be adjusted according to different scenarios, and this application does not impose any restrictions.
[0117] 206: Output the above second recognition result and wait for the second confirmation instruction input by the user.
[0118] Among them, the above second recognition result can be output to the user by voice playback or by display on a monitor.
[0119] 207: In the case of receiving the above second confirmation instruction, use the above second recognition result as the recognition result of the above first voice signal.
[0120] Among them, in the case of receiving a confirmation instruction for the above second recognition result, use the above second recognition result as the recognition result of the above first voice signal.
[0121] In order to improve the accuracy of speech recognition, the features of the user's speech can be extracted before speech recognition. During the speech recognition process, only the speech signals of a specific user are recognized, so as to exclude the interference of other speech signals.
[0122] Specifically, before the above step 201, that is, before receiving the first voice signal for the subject matter, receive the third voice signal input by the user. Here, the third voice signal and the first voice signal come from the same user; extract the feature parameters from the above third voice signal; then in steps 202 and 204, only the part of the voice signal that is the same as the above feature parameters is recognized, and the voice signals of other users will be treated as noise interference.
[0123] For the specific process of extracting feature parameters, please refer to Figure 3 , Figure 3 is a schematic flowchart of another speech recognition method disclosed in the embodiments of this application. As Figure 3 shown, the above method includes:
[0124] 301: Perform frame segmentation on the third voice signal to obtain a set of voice frames.
[0125] Among them, in addition to the above frame segmentation, it also includes pre-emphasis, windowing and other preprocessing. The specific principles and processes can refer to the above description of the speech recognition system and will not be elaborated here.
[0126] 302: Perform a fast Fourier transform on the voice frames to obtain a first spectrum.
[0127] Among them, the above voice frame is any voice frame in the above voice frame set.
[0128] 303: Filter the above first spectrum with a Mel filter bank to obtain a second spectrum.
[0129] Among them, the Mel filter bank is a filter bank with a non-linear distribution that is densely distributed in the low-frequency part and sparsely distributed in the high-frequency part. Such a distribution can better meet the auditory characteristics of the human ear. The conversion formula between frequency and Mel frequency is:
[0130] Mel(f) = 2595×log 10 (1 + f / 700)
[0131] Among them, f is the frequency to be converted, and Mel(f) is the converted Mel frequency.
[0132] 304: Perform cepstrum analysis on the above second spectrum to obtain Mel-frequency cepstral coefficients.
[0133] Among them, since the above step 302 has performed spectrum transformation on the voice frame and the above step 303 has converted the frequency domain to the Mel spectrum, the cepstrum analysis here includes:
[0134] 1. Take the logarithm of the spectrum. When generating a voice signal, due to the physical form limitations of the vocal organs, the voice signal is a slowly varying signal. After the voice signal is preprocessed, it inevitably includes slowly varying signals and rapidly varying signals, and the above voice signal can be represented as the product of a low-frequency signal and a high-frequency signal; after taking the logarithm of the spectrum, the low-frequency signal and the high-frequency signal are coupled in an additive manner, which is convenient for extracting the slowly varying signal.
[0135] 2. Take the inverse Fourier transform of the spectrum. The discrete Fourier transform and the inverse transform only differ by a coefficient. Taking the inverse Fourier transform of the spectrum can separate the high-frequency signal and the low-frequency signal coupled in an additive manner. It can be implemented through the discrete cosine transform (DCT), and the first 13 coefficients obtained after the DCT transformation are used as the Mel-frequency cepstral coefficients.
[0136] 305: Use the above Mel-frequency cepstral coefficients as the characteristic parameters of the above voice frame.
[0137] Among them, the 13-dimensional feature vector composed of Mel-frequency cepstral coefficients is the feature of the above voice frame. Each voice frame is represented by its corresponding MFCC vector, and each segment of voice signal can be represented by a matrix composed of voice frames.
[0138] The following describes the specific process of using the speech recognition method provided in this application in combination with two different usage scenarios.
[0139] (1) Usage Scenario 1: The above-mentioned speech recognition method can be well demonstrated in the scenario of storing items during warehouse inbound. Here, taking the storage of cotton quilts as an example.
[0140] Before the warehouse administrator uses the speech recognition device to identify and store the items to be stored in the warehouse, the voice features of the warehouse administrator are first extracted. After the feature extraction, the speech recognition device saves the voice feature information of the warehouse administrator.
[0141] During the specific usage process, the warehouse administrator states the name of the item to be stored in the warehouse, that is, the warehouse administrator says "cotton quilt"; the speech recognition device receives the speech signal with "cotton quilt" and recognizes the above speech signal to obtain the recognition result.
[0142] Particularly, before performing speech recognition, the received speech signal is denoised to reduce environmental interference.
[0143] Particularly, during the speech signal recognition process here, only the part of the speech signal that is the same as the above-mentioned voice feature information of the warehouse administrator is recognized, and the subsequent recognition process is also carried out for the part that is the same as the above-mentioned voice feature information of the warehouse administrator. It will not be elaborated further later.
[0144] Particularly, after obtaining the recognition result for the part that is the same as the above-mentioned voice feature information of the warehouse administrator, the above recognition is matched with the names of the items already existing in the warehouse, and the name of the item with the highest matching degree is used as the final recognition result. The same method is adopted subsequently to determine the recognition result, and it will not be elaborated further later. For example, due to the influence of environmental noise, the speech recognition device only recognizes the word "cotton". At this time, the word "cotton" is matched with all the item names in the database, and the "cotton quilt" with the highest matching degree in the database will be used as the recognition result.
[0145] Then, the above recognition result is output to the warehouse manager in the form of voice playback or display on a monitor, allowing the warehouse manager to select the recognition result: If the above recognition result is correct, that is, the voice recognition device correctly recognizes "cotton-padded quilt", then the warehouse manager confirms the above recognition result. However, since the above voice signal may contain interfering voice signals other than "cotton-padded quilt", resulting in a deviation in the voice recognition result, the warehouse manager needs to repeat "cotton-padded quilt". The voice recognition device receives the voice signal, performs recognition, and then outputs the recognition result to the warehouse manager until the output recognition result is correct, that is, a confirmation instruction from the warehouse manager is received. After that, the storage information about the cotton-padded quilt can be retrieved from the database to determine the storage location of the cotton-padded quilt in the warehouse, or the storage information of the cotton-padded quilt in the database can be modified, etc.
[0146] (2) Use Scenario Two: In the scenario where a salesperson uses an electronic scale to sell goods, the above voice recognition method can be well demonstrated. Among them, the above electronic scale is equipped with a voice recognition module, and the above voice recognition module can implement the above voice recognition method. Please refer to Figure 4 , Figure 4 which is a schematic flowchart of another voice recognition method disclosed in the embodiments of the present application.
[0147] Before the salesperson uses the electronic scale to sell goods, the voice of the salesperson is first subjected to feature extraction. After the feature extraction, the electronic scale saves the above voice feature information of the salesperson.
[0148] During the specific use process, the salesperson places the goods on the electronic scale, and at the same time, the salesperson states the name of the goods to be sold, such as Figure 4 in step 401 of
[0149] For example, the salesperson says "purple sweet potato"; the electronic scale receives the voice signal with "purple sweet potato" and recognizes the above voice signal to obtain a recognition result.
[0150] Specifically, before performing voice recognition, the received voice signal is denoised by a noise-canceling microphone to reduce environmental interference.
[0151] Specifically, before performing voice recognition, the image information about the goods can be obtained through a camera. If the salesperson weighs and sells two kinds of goods at the same time, relevant prompt information, such as "Please confirm that the types of goods are single", etc., is output.
[0152] Specifically, after obtaining the recognition result for the part that is the same as the voice feature information of the above salesperson, match the above recognition with the names of the products already existing in the database, and use the name of the product with the highest matching degree as the final recognition result. The same method will be adopted to determine the recognition result later, and it will not be elaborated further. For example, due to the influence of environmental noise, the voice recognition device only recognizes the character "purple". At this time, match the character "purple" with all the product names in the database, and use the "purple sweet potato" with the highest matching degree in the database as the recognition result.
[0153] Then output the above recognition result to the salesperson in the form of voice playback or display on a monitor. The above recognition method and display method correspond to Figure 4 step 402 in
[0154] Then let the salesperson select the recognition result, corresponding to Figure 4 step 403 in : If the above recognition result is correct, that is, the electronic scale correctly recognizes "purple sweet potato", then the salesperson confirms the above recognition result; however, since the above voice signal may contain interfering voice signals other than "purple sweet potato", resulting in a deviation in the voice recognition result, then the salesperson needs to repeat "purple sweet potato", corresponding to Figure 4 step 404 in ; The electronic scale receives the voice signal and performs recognition, and then outputs the recognition result to the salesperson until the output recognition result is correct, that is, a confirmation instruction from the salesperson is received. Then, price information, preferential information, etc. about purple sweet potatoes can be called from the database, and combined with the weight information on the electronic scale, the price of the product can be obtained, corresponding to Figure 4 step 405 in
[0155] In summary, by improving the voice recognition method, the operation process of the user can be simplified and the processing efficiency can be improved; at the same time, by extracting the voice features of the user, the voice with the above voice features can be specifically recognized during the voice recognition process, thereby improving the accuracy of voice recognition.
[0156] The method of the embodiment of the present application is described in detail above. Now, the device of the embodiment of the present application is provided.
[0157] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a voice recognition device disclosed in the embodiment of the present application. The above voice recognition device 50 may include a receiving unit 501, a recognition unit 502, an output unit 503, and a determination unit 504. Among them, the description of each unit is as follows:
[0158] The receiving unit 501 is used to receive the first voice signal for the subject matter;
[0159] An identification unit 502, configured to identify the first voice signal to obtain a first identification result;
[0160] An output unit 503, configured to output the first identification result;
[0161] A determination unit 504, configured to use the first identification result as the identification result of the first voice signal when receiving the first confirmation instruction.
[0162] The receiving unit 501 is further configured to wait for a first confirmation instruction input by the user or a second voice signal for the subject after the output unit outputs the first identification result;
[0163] The identification unit 502 is further configured to identify the second voice signal to obtain a second identification result when the first confirmation instruction is not received and the second voice signal for the subject is received;
[0164] The output unit 503 is further configured to output the second identification result;
[0165] The receiving unit 501 is further configured to wait for a second confirmation instruction input by the user after the output unit outputs the second identification result;
[0166] The determination unit 504 is further configured to use the second identification result as the identification result of the first voice signal when receiving the second confirmation instruction.
[0167] In a possible implementation manner, the receiving unit 501 is configured to obtain image information including the subject;
[0168] The output unit 503 is further configured to output a prompt message indicating an error in the number of subject types when the number of subject types included in the image information is greater than or equal to two.
[0169] In a possible implementation manner, the identification unit 502 is further configured to identify the second voice signal to obtain a third identification result;
[0170] The determination unit 504 is further configured to use the name of the subject in the database that has the highest matching degree with the third identification result as the second identification result, and the database includes at least two subjects.
[0171] In a possible implementation manner, the device further includes:
[0172] The receiving unit 501 is further configured to receive a third voice signal;
[0173] An extraction unit 505 for extracting feature parameters from the above third speech signal;
[0174] The above recognition unit 502 is further configured to recognize the part of the above second speech signal that is the same as the above feature parameters to obtain a second recognition result.
[0175] In a possible implementation manner, the above device further includes:
[0176] A framing unit 506 for performing framing processing on the above third speech signal to obtain a set of speech frames;
[0177] A transformation unit 507 for performing a fast Fourier transform on the speech frame to obtain a first spectrum, where the speech frame is any speech frame in the above set of speech frames;
[0178] A filtering unit 508 for filtering the above first spectrum with a Mel filter bank to obtain a second spectrum;
[0179] An inverse spectrum unit 509 for performing inverse spectrum analysis on the above second spectrum to obtain Mel frequency cepstral coefficients;
[0180] The above determination unit 504 is further configured to use the above Mel frequency cepstral coefficients as the feature parameters of the speech frame.
[0181] In a possible implementation manner, the above output unit 503 is further configured to play the above second recognition result;
[0182] The above output unit 503 is further configured to display the above second recognition result on a display screen.
[0183] In a possible implementation manner, the above device further includes:
[0184] A noise reduction unit 510 for performing noise reduction processing on the above first speech signal to obtain a noise-reduced speech signal;
[0185] The above recognition unit 502 is further configured to recognize the above noise-reduced speech signal to obtain a first recognition result.
[0186] Please refer to Figure 6 , Figure 6 is a schematic structural diagram of an electronic scale disclosed in an embodiment of the present application. The above electronic scale 60 may include a memory 601 and a processor 602. Further optionally, it may further include a weight sensor 603, an input device 604, and an output device 605.
[0187] Among them, the memory 601, the processor 602, the weight sensor 603, the input device 604, and the output device 605 are communicatively connected to each other through the bus 606. The input device 604 corresponds to the receiving unit 501 of the above-mentioned voice recognition device 50, and the output device 605 corresponds to the output unit 503 of the above-mentioned voice recognition device 50.
[0188] The memory 601 is used to provide a storage space, and data such as an operating system and computer programs can be stored in the storage space. The memory 601 includes but is not limited to a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM).
[0189] The processor 602 is a module for performing arithmetic and logical operations, and can be one or a combination of processing modules such as a central processing unit (CPU), a graphics processing unit (GPU), or a microprocessor unit (MPU).
[0190] A computer program is stored in the memory 601, and the processor 602 calls the computer program stored in the memory 601 to perform the following operations:
[0191] Receive a first voice signal for the subject matter;
[0192] Recognize the above first voice signal to obtain a first recognition result;
[0193] Output the above first recognition result; wait for a first confirmation instruction input by the user or a second voice signal for the above subject matter;
[0194] In the case where the above first confirmation instruction is not received and a second voice signal for the above subject matter is received, recognize the above second voice signal to obtain a second recognition result;
[0195] Output the above second recognition result; wait for a second confirmation instruction input by the user;
[0196] In the case where the above second confirmation instruction is received, use the above second recognition result as the recognition result of the above first voice signal.
[0197] In particular, the above input device 604 includes a printer for printing bills and a keyboard for inputting information.
[0198] The above processor 602 is further configured to add product information, modify and delete the information of the products that have been entered.
[0199] It should be noted that the specific implementation of the electronic scale 60 can also be correspondingly referred to Figure 2 、 Figure 3 and Figure 4 the corresponding descriptions of the method embodiments shown.
[0200] The embodiment of the present application further provides an electronic device, including: an input device, an output device, a processor, and a memory. Among them, the above input device corresponds to the receiving unit of the above voice recognition device, and the above output device corresponds to the output unit of the above voice recognition device. A computer program is stored in the above memory, and when the above processor calls the computer program stored in the above memory, it can implement Figure 2 、 Figure 3 and Figure 4 the voice recognition methods shown.
[0201] The embodiment of the present application further provides a computer-readable storage medium. A computer program is stored in the above computer-readable storage medium. When the above computer program runs on one or more processors, it can implement Figure 2 、 Figure 3 and Figure 4 the voice recognition methods shown.
[0202] In summary, it can be seen that by improving the voice recognition method, the operation process of the user can be simplified and the processing efficiency can be improved; at the same time, by extracting the voice features of the user, the voices with the above voice features can be specifically recognized during the voice recognition process, thereby improving the accuracy of voice recognition.
[0203] Those of ordinary skill in the art can understand all or part of the processes in the above method embodiments. The above processes can be completed by hardware related to computer programs. The above computer programs can be stored in a computer-readable storage medium. When the above computer programs are executed, they can include the processes of the above method embodiments. The foregoing storage media include: various media such as read-only memory ROM or random access memory RAM, magnetic disks, or optical discs that can store computer program codes.
Claims
1. A speech recognition method, characterized in that, including: Receiving a first voice signal for a subject; Identifying the first voice signal to obtain a first identification result; Outputting the first identification result; Waiting for a first confirmation instruction input by the user or a second voice signal for the subject; In the case of receiving the first confirmation instruction, using the first identification result as the identification result of the first voice signal; In the case of not receiving the first confirmation instruction and receiving a second voice signal for the subject, identifying the second voice signal to obtain a second identification result; Outputting the second identification result; Waiting for a second confirmation instruction input by the user; In the case of receiving the second confirmation instruction, using the second identification result as the identification result of the first voice signal; Before identifying the second voice signal to obtain a second identification result, the method further includes: Receiving a third voice signal; Extracting feature parameters from the third voice signal; The identifying the second voice signal to obtain a second identification result includes: Identifying the part of the second voice signal that is the same as the feature parameters to obtain a second identification result.
2. The method according to claim 1, wherein Before receiving the first voice signal for the subject, the method further includes: Obtaining image information including the subject; In the case where the number of subject types included in the image information is greater than or equal to two, outputting a prompt message indicating an incorrect number of subject types.
3. The method according to claim 2, wherein The identifying the second voice signal to obtain a second identification result includes: Identifying the second voice signal to obtain a third identification result; Using the name of the subject in the database that has the highest matching degree with the third identification result as the second identification result, where the database contains at least two subjects.
4. The method according to claim 1, wherein The extracting feature parameters from the third voice signal includes: Performing frame segmentation processing on the third voice signal to obtain a set of voice frames; Performing a fast Fourier transform on the voice frame to obtain a first frequency spectrum, where the voice frame is any voice frame in the set of voice frames; Filtering the first frequency spectrum with a Mel filter bank to obtain a second frequency spectrum; Performing cepstrum analysis on the second frequency spectrum to obtain Mel frequency cepstral coefficients; Using the Mel frequency cepstral coefficients as the feature parameters of the voice frame.
5. The method according to claim 4, wherein The outputting the second identification result includes: Playing the second identification result; Or, displaying the second identification result on a display screen.
6. The method according to any one of claims 1-5, characterized in that, After receiving the first voice signal for the subject, the method further includes: Performing noise reduction processing on the first voice signal to obtain a noise-reduced voice signal; The identifying the first voice signal to obtain a first identification result includes: Identifying the noise-reduced voice signal to obtain a first identification result.
7. A voice recognition device, characterized in that, The apparatus includes: A receiving unit for receiving a first voice signal for a subject; an identifying unit for identifying the first voice signal to obtain a first identification result; An output unit for outputting the first identification result; A determination unit, configured to use the first recognition result as the recognition result of the first voice signal when receiving a first confirmation instruction; The receiving unit is further configured to, after the output unit outputs the first recognition result, wait for a first confirmation instruction input by the user or a second voice signal for the subject matter; The recognition unit is further configured to, when not receiving the first confirmation instruction and receiving a second voice signal for the subject matter, recognize the second voice signal to obtain a second recognition result; The output unit is further configured to output the second recognition result; The receiving unit is further configured to, after the output unit outputs the second recognition result, wait for a second confirmation instruction input by the user; The determination unit is further configured to use the second recognition result as the recognition result of the first voice signal when receiving the second confirmation instruction; Before recognizing the second voice signal to obtain a second recognition result, further comprising: Receiving a third voice signal; Extracting feature parameters from the third voice signal; Recognizing the second voice signal to obtain a second recognition result, comprising: Recognizing a part of the second voice signal that is the same as the feature parameters to obtain a second recognition result.
8. An electronic scale, characterized in that, Comprising: A weight sensor, an input device, an output device, a processor, and a memory, wherein a computer program is stored in the memory, and the processor calls the computer program stored in the memory to execute the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program runs on one or more processors, it executes the method according to any one of claims 1-6.
Citation Information
Patent Citations
Commodity searching method and commodity searching device based on voice recognition
CN105574173A
Voice interaction processing method and device, device and operation system
CN107305769A
Intelligent electronic scale
CN107702778A