Voice signal recognition method and device, storage medium and electronic device
By combining voice signals collected by multiple voice devices in smart homes and performing voice recognition and evaluation on the combined signals, the problem that the voice signals collected by a single device cannot fully reflect the user scenarios, improving the accuracy of voice recognition.
Patent Information
- Application Number
- CN202311674024.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
In the multi-device scenarios of smart homes, the voice signals collected by a single device cannot fully feedback the user's speech, resulting in low speech recognition accuracy, especially in complex interactive scenarios, which is susceptible to environmental noise and device audio quality.
By combining the voice signals collected by multiple voice devices into a first voice signal and performing speech recognition on the signal, a set of candidate voice commands are obtained. Then, the candidate voice command is evaluated based on each voice signal, and the evaluation result is obtained, which is used to represent the probability of each candidate voice command, thereby selecting the target voice command.
By introducing voice signals collected by multiple devices, combining the characteristics of different devices, the accuracy of voice recognition is improved, and the problem that voice signals collected by a single device cannot fully reflect the user scenario.
Smart Images

Figure CN120126485A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of smart home / smart family. Specifically, it relates to a method and device for recognizing voice signals, a storage medium, and an electronic device. Background Art
[0002] Voice recognition technology is one of the important applications in the field of artificial intelligence. In the current mainstream end-to-end voice recognition process, after recognizing a voice signal to obtain N recognition results, the recognition results are re-scored to obtain the final recognition result of the voice signal.
[0003] However, in the above voice signal recognition process, the voice signal recognition and re-scoring processes only involve voice data collected by a single device. That is, in a multi-device smart home scenario, only the device with the highest confidence score is selected as the voice signal input data, and the audio collected by the same device with the highest confidence is also used for input in the re-scoring of the voice recognition result. However, the audio collected by a single device cannot comprehensively reflect the user's speaking scenario. Especially in complex interaction scenarios, when it may not be possible to obtain audio of good quality due to environmental noise or device sound collection, it will have a significant impact on the voice recognition effect.
[0004] It can be seen that in the related art, the voice signal recognition method has the problem of low voice recognition accuracy in the voice signal recognition process. Summary of the Invention
[0005] Embodiments of the present application provide a method and device for recognizing voice signals, a storage medium, and an electronic device, so as to at least solve the problem of low voice recognition accuracy in the voice signal recognition method in the related art.
[0006] According to one aspect of the embodiments of the present application, a method for recognizing voice signals is provided, which is applied to a smart device and includes: combining a group of voice signals into a first voice signal, where each voice signal in the group of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object; performing voice recognition on the first voice signal to obtain a group of candidate voice commands, where each candidate voice command in the group of candidate voice commands is a piece of text information recognized from the first voice signal; evaluating each candidate voice command based on the group of voice signals to obtain an evaluation result corresponding to each candidate voice command, where the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is a voice command issued by the target object; and selecting a target voice command from the group of candidate voice commands based on the evaluation result corresponding to each candidate voice command.
[0007] According to another aspect of the embodiments of the present application, there is also provided an apparatus for recognizing a voice signal, which is applied to an intelligent device and includes: a merging unit configured to merge a group of voice signals into a first voice signal, where each voice signal in the group of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object; a recognition unit configured to perform voice recognition on the first voice signal to obtain a group of candidate voice commands, where each candidate voice command in the group of candidate voice commands is a piece of text information recognized from the first voice signal; a first evaluation unit configured to evaluate each candidate voice command based on the group of voice signals to obtain an evaluation result corresponding to each candidate voice command, where the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is a voice command issued by the target object; and a first selection unit configured to select a target voice command from the group of candidate voice commands based on the evaluation result corresponding to each candidate voice command.
[0008] According to still another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the above-mentioned voice signal recognition method when running.
[0009] According to still another aspect of the embodiments of the present application, there is also provided an electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the above-mentioned processor executes the above-mentioned voice signal recognition method through the computer program.
[0010] In an embodiment of the present application, a method is adopted in which, based on a set of voice signals, multiple voice commands of a combined signal of the recognized set of voice signals are evaluated, and the set of voice signals is combined into a first voice signal. Each voice signal in the set of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object; the first voice signal is subjected to voice recognition to obtain a set of candidate voice commands, where each candidate voice command in the set of candidate voice commands is a type of text information recognized from the first voice signal; each candidate voice command is evaluated based on the set of voice signals to obtain an evaluation result corresponding to each candidate voice command, where the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is a voice command issued by the target object; based on the evaluation results corresponding to each candidate voice command, a target voice command is selected from the set of candidate voice commands. Since voice signals collected by multiple devices are introduced in the recognition and evaluation of multiple voice commands, the determined target voice command combines the characteristics of voice signals collected by multiple different devices, more comprehensively reflects the voice command of the target object, achieves the technical effect of improving the accuracy of voice recognition, and further solves the problem of low accuracy of voice recognition in the related art in the process of recognizing voice signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 is a schematic diagram of the hardware environment of a method for recognizing voice signals according to an embodiment of the present application;
[0014] Figure 2 is a schematic flowchart of an alternative method for recognizing voice signals according to an embodiment of the present application;
[0015] Figure 3 is a schematic flowchart of another alternative method for recognizing voice signals according to an embodiment of the present application;
[0016] Figure 4 is a schematic diagram of an alternative method for recognizing voice signals according to an embodiment of the present application;
[0017] Figure 5 It is a schematic flowchart of another optional method for recognizing a voice signal according to an embodiment of the present application;
[0018] Figure 6 It is a schematic flowchart of another optional method for recognizing a voice signal according to an embodiment of the present application;
[0019] Figure 7 It is a block diagram of the structure of an optional voice signal recognition device according to an embodiment of the present application;
[0020] Figure 8 It is a block diagram of the structure of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0021] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not necessarily limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0023] According to one aspect of the embodiments of the present application, a method for recognizing a voice signal is provided, which can be applied to intelligent devices. The method for recognizing a voice signal is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home appliance ecosystem, and Intelligence House ecosystem. Optionally, in this embodiment, the above-mentioned method for recognizing a voice signal can be applied to, for example, Figure 1 the hardware environment composed of the terminal device 102 and the server 104 as shown. As Figure 1As shown, the server 104 is connected to the terminal device 102 via a network and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data operation services for the server 104.
[0024] The above network can include but is not limited to at least one of the following: wired network, wireless network. The above wired network can include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network can include but is not limited to at least one of the following: WIFI, Bluetooth. The terminal device 102 is not limited to being a PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing device, smart dishwasher, smart projection device, smart TV, smart drying rack, smart curtain, smart audio and video, smart socket, smart speaker, smart sound box, smart fresh air device, smart kitchen and bathroom device, smart bathroom device, smart floor sweeping robot, smart window cleaning robot, smart mopping robot, smart air purification device, smart steam box, smart microwave oven, smart kitchen water heater, smart purifier, smart water dispenser, smart door lock, etc., which are intelligent devices capable of collecting voice signals.
[0025] The method for recognizing the voice signal in the embodiment of the present application can be executed by the server 104 or jointly executed by the server 104 and the terminal device 102. Taking the server 104 to execute the method for recognizing the voice signal in this embodiment as an example, Figure 2 is a schematic flowchart of an optional method for recognizing a voice signal according to an embodiment of the present application, as Figure 2 shown, the process of the method can include the following steps:
[0026] Step S202, combining a group of voice signals into a first voice signal, where each voice signal in the group of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object.
[0027] The method for recognizing the voice signal in this embodiment can be applied to the scenario of performing voice recognition on the collected voice signals. Here, the voice signal can be a signal collected by a voice device, and the voice device can be a device including a voice module, which can be an intelligent device in the aforementioned terminal devices. When the user issues a voice command, the intelligent device near the user can collect the user's voice command in real time, generate a voice signal, and realize the understanding of the user's intention by performing voice recognition on the voice signal.
[0028] The current mainstream end-to-end automatic speech recognition system (E2E ASR) is divided into two parts: an encoder and a decoder. During the recognition process, first, feature extraction is performed on the speech signal, and then the extracted features are input into the encoder of the end-to-end speech recognition system to obtain the posterior probability distribution of the speech signal. By combining the posterior probability with a language model and performing decoding (one-pass decoding), the N most likely recognition results (Nbest) are finally output.
[0029] Since the result with the highest probability among the N recognition results is not necessarily the correct one, a rescore process is usually added to improve the accuracy. That is, the speech signal is input into the encoder of the end-to-end speech recognition system again, and the output speech features are combined with the Nbest recognition results and input into the decoder of the end-to-end speech recognition system, so as to reorder the Nbest recognition results and obtain a more accurate recognition effect.
[0030] As Figure 3 shown, the current speech signal input for speech recognition is all speech data collected by a single device. That is, in the multi-device scenario of smart home, only the device with the highest confidence score is selected as the speech signal input data. In the subsequent process, when performing one-pass decoding and rescore decoding, the audio collected by the same device with the highest confidence is used for input.
[0031] However, considering that the audio collected by a single device cannot fully reflect the user's speaking scenario, especially in complex interaction scenarios, it may be affected by environmental noise or device sound collection, resulting in poor-quality audio and having a greater impact on the speech recognition effect. The inability to obtain more effective speech information from multiple smart home devices will severely limit the speech recognition accuracy.
[0032] To at least partially solve the above problems, in this embodiment, for the audio picked up by multiple smart home devices in the smart home scenario, the speech signals collected by different devices can be merged into a first speech signal, and taking advantage of the fact that rescore decoding of the speech signals collected by different devices does not require merging and alignment, the recognized Nbest is separately rescored in combination with different speech signals, so as to achieve information complementarity of the different information contained in the picked-up sounds of different devices during the rescore process, thereby improving the accuracy of speech recognition.
[0033] In this embodiment, when a voice command issued by a target object is detected, a set of voice signals collected by multiple voice devices and corresponding to the voice command issued by the target object can be merged into a first voice signal. Here, the target object may be the aforementioned user.
[0034] Optionally, each voice signal in the above set of voice signals may correspond to a voice device respectively, that is, one voice device corresponds to one voice signal.
[0035] Step S204: Perform voice recognition on the first voice signal to obtain a set of candidate voice commands, where each candidate voice command in the set of candidate voice commands is a type of text information recognized from the first voice signal.
[0036] After merging the first voice signal, voice recognition can be performed on the first voice signal to obtain a set of candidate voice commands. Here, the set of candidate voice commands may be the recognition result of the first voice signal, that is, the aforementioned Nbest, and each candidate voice command in the set of candidate voice commands may be a type of text information recognized from the first voice signal.
[0037] Due to factors such as many words with similar pronunciations but different meanings, as well as many words with the same pronunciation but different meanings or polyphonic characters, some voices in a voice command can be recognized as multiple texts. For example, "today" is recognized as "money", and "banana" is recognized as "intersect". In this embodiment, there may be only individual word differences between each candidate voice command in the above set of candidate voice commands.
[0038] Optionally, when performing voice recognition on the first voice signal to generate candidate voice commands, a recognition probability corresponding to each candidate voice command can be generated, that is, the probability that each candidate voice command is the voice command issued by the target object.
[0039] Step S206: Evaluate each candidate voice command based on the set of voice signals to obtain an evaluation result corresponding to each candidate voice command, where the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is the voice command issued by the target object.
[0040] In order to more comprehensively feedback the voice characteristics of the target object, avoid the error of the voice signal collected by a single device from affecting the result of voice recognition, and further improve the accuracy of voice recognition. In this embodiment, after obtaining a set of candidate voice commands, each candidate voice command can be evaluated based on the set of voice signals to obtain an evaluation result corresponding to each candidate voice command. Here, the evaluation result corresponding to each candidate voice command can be used to represent the probability that each candidate voice command is the voice command issued by the target object.
[0041] The above evaluation may refer to re-scoring a set of candidate voice commands respectively according to each voice signal in a set of voice signals. For each candidate voice command in a set of candidate voice commands, the re-scoring based on a set of voice signals may correspond to a set of re-scoring results, and a set of re-scoring results may correspond to an evaluation result of a candidate voice command.
[0042] For example, taking a set of voice signals as Voice Signal 1, Voice Signal 2,..., Voice Signal K, and a set of candidate voice commands as Nbest, which contains N speech recognition results as an example, as Figure 4 shown, each speech recognition result may have K re-scoring results, and based on the K re-scoring results, an evaluation result of a speech recognition result can be determined.
[0043] Step S208, based on the evaluation results corresponding to each candidate voice command, select a target voice command from a set of candidate voice commands.
[0044] Since the evaluation results corresponding to each candidate voice command are used to represent the probability that each candidate voice command is a voice command issued by a target object, based on the evaluation results corresponding to each candidate voice command, the candidate voice command with the highest probability can be selected from a set of candidate voice commands according to the magnitude of the probability as the target voice command. Here, the target voice command is the final recognition result of the voice command issued by the target object.
[0045] Through the above steps S202 to S208, a set of voice signals are combined into a first voice signal, where each voice signal in the set of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object; voice recognition is performed on the first voice signal to obtain a set of candidate voice commands, where each candidate voice command in the set of candidate voice commands is a text information recognized from the first voice signal; each candidate voice command is evaluated based on the set of voice signals to obtain an evaluation result corresponding to each candidate voice command, where the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is a voice command issued by a target object; based on the evaluation results corresponding to each candidate voice command, a target voice command is selected from the set of candidate voice commands, which solves the problem of low speech recognition accuracy in the speech signal recognition method in the related art during the speech signal recognition process, and improves the speech recognition accuracy.
[0046] In an exemplary embodiment, combining a set of voice signals into a first voice signal includes:
[0047] S11. Starting from the signal acquisition time corresponding to the wake-up word in each voice signal of a group of voice signals, based on the voice duration corresponding to the second voice signal with the highest confidence level in the group of voice signals, determine the voice signal segments to be merged in each voice signal other than the second voice signal in the group of voice signals, obtaining a group of voice signal segments to be merged;
[0048] S12. Merge the group of voice signal segments to be merged with the second voice signal to obtain the first voice signal.
[0049] Considering that the distance between each voice device and the user may vary, which may lead to different delays in the voice signals collected by each voice device, and the delay will cause the audio of multiple voice signals to not be directly additively merged. In this embodiment, when merging a group of voice signals, starting from the signal acquisition time corresponding to the wake-up word in each voice signal of the group of voice signals, based on the voice duration corresponding to the second voice signal with the highest confidence level in the group of voice signals, determine the voice signal segments to be merged in each voice signal other than the second voice signal in the group of voice signals, obtaining a group of voice signal segments to be merged.
[0050] The confidence level of a group of voice signals can be obtained through relevant calculations on the voice signals after the voice signals are acquired and before voice recognition. The higher the confidence level, the more the collected voice signal matches the actual voice command. Among a group of voice signals, the second voice signal with the highest confidence level is a voice signal with relatively high voice energy and voice integrity in the group of voice signals.
[0051] In human-machine dialogue, usually a relevant wake-up word (e.g., "Xiaoyou, Xiaoyou") is required to wake up the device so that after the device is woken up by the wake-up word, a voice command can be issued to the device. Therefore, in a group of voice signals collected by the voice device, there may be voice signal segments corresponding to the wake-up word. Although there are different delays in the voice signals collected by each voice device, the voice durations of the voice signals corresponding to the voice commands issued by the target object collected by each voice device can be the same. Taking the second voice signal with the highest confidence level as the merging benchmark, that is, based on the voice duration of the second voice signal, determine the voice durations of the voice signal segments to be merged in each of the other voice signals. Then, according to the voice signals corresponding to the wake-up word in each of the other voice signals, determine the starting points of the voice signal segments to be merged, and then intercept the group of voice signal segments to be merged according to the determined voice durations. In addition, the starting time and ending time of the second voice signal can also be used as the starting time and ending time of each voice signal segment to be merged in the group of voice signal segments to be merged, and the other time points in each voice signal segment to be merged are correspondingly converted according to the starting time and ending time.
[0052] After determining a set of speech signal segments to be merged, the set of speech signal segments to be merged can be merged with the second speech signal to obtain the first speech signal.
[0053] Optionally, during the speech merging process, the speech energy corresponding to each speech signal segment can be normalized first. That is, according to the speech energy of the same speech command in different speech signal segments, the weights corresponding to different energies are determined, and then each speech signal segment is processed with the weights respectively. For example, the first speech energy at a certain position in the first speech signal segment is 80, and the second speech energy at the same position as the first speech energy in the second speech signal segment is 40. When merging the first speech signal segment and the second speech signal segment, the signal at the corresponding position in the first speech signal segment can be multiplied by 80 / (80 + 40) first, and the signal at the corresponding position in the second speech signal segment can be multiplied by 40 / (80 + 40), and then the signals at the corresponding positions in the processed first speech signal segment and the second speech signal segment are superimposed and merged.
[0054] Through this embodiment, starting from the signal acquisition time corresponding to the wake-up word in each speech signal, according to the speech duration of the speech signal with the highest confidence, the speech segments to be merged in other speech signals can be determined, which can solve the problem that speech cannot be merged due to different time delays of speech signals collected by different devices.
[0055] In an exemplary embodiment, speech recognition is performed on the first speech signal to obtain a set of candidate speech commands, including:
[0056] S21, input the first speech signal into the target encoder to obtain a feature vector corresponding to the first speech signal, where the target encoder is an encoder for encoding the input speech signal;
[0057] S22, input the feature vector corresponding to the first speech signal into the first decoder to obtain a set of candidate speech commands output by the first decoder, where the first decoder is a decoder for decoding the input feature vector into corresponding text information.
[0058] When performing speech recognition on the first speech signal to obtain a set of candidate speech commands, a set of candidate speech commands can be obtained by encoding and then decoding the first speech signal. That is, the first speech signal is input into the target encoder to obtain a feature vector corresponding to the first speech signal, and then the feature vector corresponding to the first speech signal is input into the first decoder to obtain a set of candidate speech commands output by the first decoder.
[0059] The above-mentioned target encoder can be an encoder for encoding an input voice signal, and the first decoder can be a decoder for decoding an input feature vector into corresponding text information. The target encoder and the first decoder can be corresponding to each other.
[0060] Through this embodiment, the voice signal is converted into a digital signal that can be processed by a computer through an encoder and a decoder, the feature vector therein is extracted, and then the vector is converted into text through decoding, so that the user's voice can be understood by the computer and the human-computer interaction can be improved.
[0061] In an exemplary embodiment, each candidate voice command is evaluated based on a set of voice signals to obtain an evaluation result corresponding to each candidate voice command, including:
[0062] S31. Respectively use each candidate voice command as the current voice command to perform the following evaluation operations to obtain an evaluation result corresponding to each candidate voice command:
[0063] S32. Input the current voice command and the feature vector corresponding to each voice signal into the second decoder to obtain the evaluation result corresponding to each voice signal output by the second decoder, where the second decoder is a decoder for evaluating an input voice command based on the matching degree between the input feature vector corresponding to the voice signal and the input voice command;
[0064] S33. Fuse the evaluation results corresponding to each voice signal to obtain an evaluation result corresponding to the current voice command.
[0065] Since autoregressive decoding has nothing to do with the alignment of sequential audio signals, after obtaining a set of candidate voice commands, the evaluation of a set of candidate voice commands can utilize the advantage that the resampling and decoding of voice signals collected by different devices do not require merging and alignment, and perform resampling and decoding separately according to a set of voice signals.
[0066] In this embodiment, each candidate voice command can be respectively used as the current voice command to perform an evaluation operation to obtain an evaluation result corresponding to each candidate voice command. The evaluation operation can be: input the current voice command and the feature vector corresponding to each voice signal into the second decoder to obtain the evaluation result corresponding to each voice signal output by the second decoder, and then fuse the evaluation results corresponding to each voice signal to obtain an evaluation result corresponding to the current voice command.
[0067] The above-mentioned second decoder can be a decoder for evaluating an input voice command based on the matching degree between the input feature vector corresponding to the voice signal and the input voice command. The second decoder (i.e., the aforementioned resampling and decoding decoder) can be a decoder different from the first decoder, and the two decoders can be relatively independent.
[0068] Optionally, in order to improve the decoding efficiency, the data input into the second decoder each time can be a feature vector of a voice signal and a set of candidate voice commands. Correspondingly, the data output by the second decoder can be the scoring results of a set of candidate voice commands based on a voice signal.
[0069] Optionally, the above-mentioned fusion of the evaluation results corresponding to each voice signal may refer to the calculation of the evaluation results corresponding to a plurality of different voice signals of a candidate voice command according to a preset calculation formula. The calculation result is the evaluation result of the candidate voice command. The calculation method can be to add all the evaluation results, or multiply all the evaluation results, or perform weighted interpolation combination according to the preset weight values. This embodiment does not limit this here.
[0070] Through this embodiment, each voice signal in a group of voice signals is used to perform re-scoring and decoding on the candidate voice command respectively, and then the decoding results corresponding to different voice signals are fused to obtain the evaluation result of the candidate voice command, realizing the signal complementarity of different voice signals and improving the accuracy of speech recognition.
[0071] In an exemplary embodiment, fusing the evaluation results corresponding to each voice signal to obtain the evaluation result corresponding to the current voice command includes:
[0072] S41, performing weighted summation on the evaluation results corresponding to each voice signal according to the weight values corresponding to each voice signal to obtain the evaluation result corresponding to the current voice command, where the magnitude of the weight value corresponding to each voice signal is positively correlated with the confidence level of each voice signal.
[0073] When fusing the evaluation results corresponding to each voice signal, the evaluation results corresponding to each voice signal of a candidate voice command can be fused and calculated in the manner of weighted interpolation combination. That is, performing weighted summation on the evaluation results corresponding to each voice signal according to the weight values corresponding to each voice signal to obtain the evaluation result corresponding to the current voice command.
[0074] Since the distance between each voice device and the target object and the voice pickup ability of each voice device itself may be different, in this embodiment, different weight values can be selected according to the confidence level of each voice signal. The magnitude of the weight value corresponding to each voice signal can be positively correlated with the confidence level of each voice signal, that is, the higher the confidence level, the greater the weight.
[0075] Through this embodiment, different weights are selected according to the confidence levels corresponding to the voice signals, and weighted summation is performed in such a way that the lower the confidence level, the smaller the weight, so as to complete the fusion of the evaluation results corresponding to each voice signal. While realizing the complementary information of different voice signals, it is possible to avoid the voice signals with smaller confidence levels from affecting the speech recognition results, thereby improving the accuracy of the speech recognition results.
[0076] In an exemplary embodiment, before evaluating each candidate voice command based on a set of voice signals, the above method further includes:
[0077] S51, sort a set of voice signals in descending order of confidence levels to obtain the sorting result of the set of voice signals;
[0078] S52, determine the weight value corresponding to each voice signal according to the sorting result of the set of voice signals, where the weight value corresponding to the nth voice signal in the set of voice signals is 1 / 2 n , and the weight value corresponding to the last voice signal in the set of voice signals is 1 / 2 N-1 , N is the number of voice signals included in the set of voice signals, and 1 ≤ n < N.
[0079] In this embodiment, a set of voice signals can be sorted in descending order of confidence levels to obtain the sorting result of the set of voice signals. The sorting can be performed after determining the confidence level of each voice signal, or can be performed when determining the evaluation of each candidate voice command based on a set of voice signals. This embodiment does not make a limitation on this.
[0080] When determining the weight value corresponding to each voice signal, the weight value corresponding to each voice signal can be determined according to the sorting result of the set of voice signals, that is, after arranging different weight values from large to small, the corresponding weight value is selected for each voice signal in turn according to the sorting result of the set of voice signals.
[0081] In this embodiment, the weight value corresponding to the nth voice signal in the set of voice signals can be 1 / 2 n , and the weight value corresponding to the last voice signal in the set of voice signals is 1 / 2 N-1 , N is the number of voice signals included in the set of voice signals, and 1 ≤ n < N.
[0082] Taking a set of voice signals as voice signal 1, voice signal 2,... voice signal K, and a set of candidate voice commands as Nbest, including N speech recognition results as an example, as Figure 5As shown, the merged speech signal passes through an encoder and a single-pass decoder to obtain Nbest. Then, speech signals 1... K are respectively passed through an encoder to obtain high-dimensional feature vectors 1... K. Each high-dimensional feature vector and the Nbest result are used as a group and input into a rescoring decoder to obtain a rescoring result W 1 …W K For the rescoring result W 1 …W K In the order from high to low of the speech signal confidence, interpolation calculations are respectively performed using weights (alpha1…alphaK) of 1 / 2, (1 / 2)^2, (1 / 2)^3, …, [1 - ((1 / 2)^1 + … + (1 / 2)^K - 1)], as shown in formula (1).
[0083] W_ rescore = alpha1*W 1_rescore + … + alphaK*W K_rescore (1)
[0084] Through this embodiment, determining the weight corresponding to each speech signal according to the confidence ranking result of a group of speech signals can improve the selection efficiency of the weights, and further improve the fusion efficiency of the evaluation results.
[0085] In an exemplary embodiment, before selecting the first speech signal from a group of speech signals, the above method further includes:
[0086] S61, obtaining the speech signals collected by each speech device among multiple speech devices to obtain multiple speech signals, where the speech signal collected by each speech device is a speech signal corresponding to a speech command issued by a target object;
[0087] S62, selecting the speech signals with a confidence greater than or equal to a confidence threshold from the multiple speech signals to obtain a group of speech signals.
[0088] Before performing speech recognition, the speech signals collected by each speech device among multiple speech devices can be obtained to obtain multiple speech signals. The speech signal collected by each speech device can all be a speech signal corresponding to a speech command issued by a target object. To improve the accuracy of speech recognition, before speech recognition, for the obtained multiple speech signals, a screening can be performed first.
[0089] The above screening can be performed according to the confidence. According to a preset confidence threshold, when the confidences of multiple speech signals are calculated, the speech signals with a confidence greater than or equal to the confidence threshold are selected from the multiple speech signals to obtain a group of speech signals.
[0090] In this embodiment, the obtained voice signals are screened according to a confidence threshold, and then the voice signals are recognized, which can reduce the computational complexity of voice recognition and improve the voice recognition efficiency.
[0091] In an exemplary embodiment, before selecting the voice signals with a confidence greater than or equal to the confidence threshold from multiple voice signals, the above method further includes:
[0092] S71, performing a confidence evaluation on each voice signal according to a set of voice parameters corresponding to each voice signal to obtain the confidence of each voice signal, where the set of voice parameters includes at least one of the following: voice direction, voice intensity.
[0093] In the process of calculating the confidence of voice signals, considering the differences in factors such as the placement position and physical environment of the voice device and the composition of the voice module, the collected voice signals will contain different voice feature information. In this embodiment, a confidence evaluation can be performed on each voice signal according to a set of voice parameters corresponding to each voice signal to obtain the confidence of each voice signal.
[0094] The above set of voice parameters may include at least one of the following: voice direction, voice intensity. The voice direction may be the direction of the target object relative to the voice device, that is, the user's pronunciation direction. The set of voice parameters may also include distance, ambient noise, the microphone array mode and quantity of the device, etc.
[0095] In this embodiment, by calculating the confidence of voice signals according to a set of voice parameters, the accuracy of the confidence can be improved, and thus the accuracy of the voice signal recognition result can be improved.
[0096] The following explains the voice signal recognition method in the embodiments of the present application in combination with optional examples. In this optional example, the first voice signal is Voice Signal 1, and a set of candidate voice commands is Nbest.
[0097] This optional example provides a voice recognition rescoring method in a smart home multi-network device scenario. After multiple smart home devices pick up sound, the voice signals collected by different devices are merged and then voice recognition is performed to obtain Nbest, and the voice signals collected by different devices are respectively combined with Nbest for rescoring decoding, and weight interpolation is performed on the rescoring results. When performing weight interpolation on the rescoring results, the corresponding weights are respectively selected from high to low for interpolation according to the N voice signals screened according to the confidence. Information complementarity of different devices is performed during the rescoring decoding process, and a better voice recognition effect can be obtained.
[0098] As Figure 6 shown, the process of the voice signal recognition method in this optional example may include the following steps:
[0099] Step S602, after the final control device obtains the voice signals 1…M of all smart home devices, select the top K voice signals according to the confidence threshold, that is, voice signals 1…voice signal K.
[0100] Step S604, after merging the K voice signals, input them into the encoder of the end-to-end voice recognition system to obtain high-dimensional feature vectors.
[0101] Step S606, input the high-dimensional feature vectors into the decoder once to obtain N speech recognition results Nbest according to the probability from high to low.
[0102] Step S608, pass the voice signals 1…voice signal K through the encoder respectively to obtain high-dimensional feature vectors 1…K. Take each high-dimensional feature vector and the Nbest result as a group, and input them into the rescoring decoder respectively to obtain the rescoring results 1…K.
[0103] Step S610, perform interpolation calculation by selecting weights for the rescoring results 1…K in descending order of the confidence of the voice signals.
[0104] Step S612, re-sort the Nbest outputs according to the interpolation results, and output the recognition result sequence with the highest probability as the result of this voice recognition.
[0105] Through this optional example, by integrating the different voice signal information obtained from multiple smart home devices, and using them as the input data when obtaining Nbest and the input data for rescoring respectively, the purpose of improving the recognition accuracy can be achieved, enabling users to obtain a better interaction experience.
[0106] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0107] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware server. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0108] According to another aspect of the embodiments of the present application, there is also provided a voice signal recognition device for implementing the above-mentioned voice signal recognition method. This voice signal recognition device can be applied to intelligent devices. Figure 7 It is a structural block diagram of an optional voice signal recognition device according to an embodiment of the present application, as Figure 7 shown. This device may include:
[0109] A merging unit 702, configured to merge a group of voice signals into a first voice signal, where each voice signal in the group of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object;
[0110] An identification unit 704, connected to the merging unit 702, configured to perform voice recognition on the first voice signal to obtain a group of candidate voice commands, where each candidate voice command in the group of candidate voice commands is a piece of text information recognized from the first voice signal;
[0111] A first evaluation unit 706, connected to the identification unit 704, configured to evaluate each candidate voice command based on the group of voice signals to obtain an evaluation result corresponding to each candidate voice command, where the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is a voice command issued by the target object;
[0112] A first selection unit 708, connected to the first evaluation unit 706, configured to select a target voice command from the group of candidate voice commands based on the evaluation result corresponding to each candidate voice command.
[0113] It should be noted that the merging unit 702 in this embodiment can be used to execute the above step S202, the recognition unit 704 in this embodiment can be used to execute the above step S204, the first evaluation unit 706 in this embodiment can be used to execute the above step S206, and the first selection unit 708 in this embodiment can be used to execute the above step S208.
[0114] Through the above modules, a set of voice signals is merged into a first voice signal, where each voice signal in the set of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object; the first voice signal is subjected to voice recognition to obtain a set of candidate voice commands, where each candidate voice command in the set of candidate voice commands is a piece of text information recognized from the first voice signal; each candidate voice command is evaluated based on the set of voice signals to obtain an evaluation result corresponding to each candidate voice command, where the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is a voice command issued by the target object; based on the evaluation results corresponding to each candidate voice command, a target voice command is selected from the set of candidate voice commands, which solves the problem of low voice recognition accuracy in the voice signal recognition method in the related art and improves the voice recognition accuracy.
[0115] In an exemplary embodiment, the merging unit includes:
[0116] A determination module, configured to use the signal acquisition time corresponding to the wake-up word in each voice signal in a set of voice signals as a starting point, and based on the voice duration corresponding to the second voice signal with the highest confidence in the set of voice signals, determine the voice signal segments to be merged in each of the other voice signals in the set of voice signals except the second voice signal, to obtain a set of voice signal segments to be merged;
[0117] A merging module, configured to merge the set of voice signal segments to be merged with the second voice signal to obtain a first voice signal.
[0118] In an exemplary embodiment, the recognition unit includes:
[0119] A first input module, configured to input the first voice signal into a target encoder to obtain a feature vector corresponding to the first voice signal, where the target encoder is an encoder for encoding the input voice signal;
[0120] A second input module, configured to input the feature vector corresponding to the first voice signal into a first decoder to obtain a set of candidate voice commands output by the first decoder, where the first decoder is a decoder for decoding the input feature vector into corresponding text information.
[0121] In an exemplary embodiment, the first evaluation unit includes:
[0122] An execution module, configured to perform the following evaluation operations on each candidate voice command as the current voice command respectively, so as to obtain an evaluation result corresponding to each candidate voice command:
[0123] Input the current voice command and the feature vector corresponding to each voice signal into a second decoder, so as to obtain an evaluation result corresponding to each voice signal output by the second decoder, where the second decoder is a decoder for evaluating the input voice command based on the matching degree between the input feature vector corresponding to the voice signal and the input voice command;
[0124] Fuse the evaluation results corresponding to each voice signal to obtain an evaluation result corresponding to the current voice command.
[0125] In an exemplary embodiment, the execution module includes:
[0126] A summation sub-module, configured to perform weighted summation on the evaluation results corresponding to each voice signal according to the weight corresponding to each voice signal, so as to obtain an evaluation result corresponding to the current voice command, where the magnitude of the weight corresponding to each voice signal is positively correlated with the confidence of each voice signal.
[0127] In an exemplary embodiment, the above device further includes:
[0128] A sorting unit, configured to sort a group of voice signals in descending order of confidence before evaluating each candidate voice command based on the group of voice signals, so as to obtain a sorting result of the group of voice signals;
[0129] A determination unit, configured to determine the weight corresponding to each voice signal according to the sorting result of the group of voice signals, where the weight corresponding to the nth voice signal in the group of voice signals is 1 / 2 n , and the weight corresponding to the last voice signal in the group of voice signals is 1 / 2 N-1 , N is the number of voice signals included in the group of voice signals, and 1≤n<N.
[0130] In an exemplary embodiment, the above device further includes:
[0131] An acquisition unit, configured to acquire the voice signals collected by each voice device among a plurality of voice devices before selecting a first voice signal from the group of voice signals, so as to obtain a plurality of voice signals, where the voice signal collected by each voice device is a voice signal corresponding to the voice command issued by the target object;
[0132] A first selection unit, configured to select speech signals with a confidence level greater than or equal to a confidence threshold from a plurality of speech signals, so as to obtain a set of speech signals.
[0133] In an exemplary embodiment, the above-mentioned apparatus further includes:
[0134] A second evaluation unit, configured to evaluate the confidence level of each speech signal according to a set of speech parameters corresponding to each speech signal before selecting speech signals with a confidence level greater than or equal to the confidence threshold from the plurality of speech signals, so as to obtain the confidence level of each speech signal, where the set of speech parameters includes at least one of the following: speech direction, speech intensity.
[0135] It should be noted here that the examples and application scenarios implemented by the above-mentioned modules and the corresponding steps are the same, but are not limited to the content disclosed in the above-mentioned embodiments. It should be noted that the above-mentioned modules, as part of the apparatus, can run in the hardware environment as shown in Figure 1 and can be implemented by software or by hardware, where the hardware environment includes a network environment.
[0136] According to another aspect of the embodiments of the present application, there is also provided a storage medium, which can be located on a smart device. Optionally, in this embodiment, the above-mentioned storage medium can be used to execute the program code of any one of the above-mentioned speech signal recognition methods in the embodiments of the present application.
[0137] Optionally, in this embodiment, the above-mentioned storage medium can be located on at least one of a plurality of network devices in the network shown in the above-mentioned embodiment.
[0138] Optionally, in this embodiment, the storage medium is set to store program code for executing the following steps:
[0139] S1. Merge a set of speech signals into a first speech signal, where each speech signal in the set of speech signals is a speech signal collected by a speech device and corresponding to a speech command issued by a target object;
[0140] S2. Perform speech recognition on the first speech signal to obtain a set of candidate speech commands, where each candidate speech command in the set of candidate speech commands is a piece of text information recognized from the first speech signal;
[0141] S3. Evaluate each candidate speech command based on the set of speech signals to obtain an evaluation result corresponding to each candidate speech command, where the evaluation result corresponding to each candidate speech command is used to represent the probability that each candidate speech command is a speech command issued by the target object;
[0142] S4. Select a target voice command from a set of candidate voice commands based on the evaluation results corresponding to each candidate voice command.
[0143] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiment, and will not be elaborated herein.
[0144] Optionally, in this embodiment, the above storage medium may include, but is not limited to, various media such as USB flash drives, ROMs, RAMs, mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0145] According to another aspect of the embodiments of the present application, an electronic device for implementing the above voice signal recognition method is further provided. The electronic device may be a smart device, and the electronic device may be a server, a terminal, or a combination thereof.
[0146] Figure 8 is a structural block diagram of an optional electronic device according to the embodiments of the present application, as Figure 8 shown, including a processor 802, a communication interface 804, a memory 806, and a communication bus 808. Among them, the processor 802, the communication interface 804, and the memory 806 communicate with each other through the communication bus 808. Among them,
[0147] The memory 806 is used to store computer programs;
[0148] When the processor 802 is used to execute the computer program stored in the memory 806, the following steps are implemented:
[0149] S1. Merge a set of voice signals into a first voice signal, where each voice signal in the set of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object;
[0150] S2. Perform voice recognition on the first voice signal to obtain a set of candidate voice commands, where each candidate voice command in the set of candidate voice commands is a piece of text information recognized from the first voice signal;
[0151] S3. Evaluate each candidate voice command based on the set of voice signals to obtain an evaluation result corresponding to each candidate voice command, where the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is a voice command issued by the target object;
[0152] S4. Select a target voice command from the set of candidate voice commands based on the evaluation results corresponding to each candidate voice command.
[0153] Optionally, the communication bus may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 8 only a thick line is used to represent it in Figure 8 , but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0154] The memory may include a RAM, and may also include a non-volatile memory, for example, at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0155] As an example, the above-mentioned memory 806 may but is not limited to include the merging unit 702, the recognition unit 704, the first evaluation unit 706, and the first selection unit 708 in the recognition device of the above-mentioned voice signal. In addition, it may also include but is not limited to other module units in the recognition device of the above-mentioned voice signal, which will not be elaborated in this example.
[0156] The above-mentioned processor may be a general-purpose processor, which may include but is not limited to: a CPU (Central Processing Unit), an NP (Network Processor), etc.; it may also be a DSP (Digital Signal Processing), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0157] Optionally, the specific examples in this embodiment may refer to the examples described in the above-mentioned embodiment, and will not be elaborated here.
[0158] Those of ordinary skill in the art can understand that Figure 8 the structure shown is only schematic. The device for implementing the above-mentioned voice signal recognition method may be a terminal device, which may be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a Mobile Internet Devices (MID), a PAD, and other terminal devices.Figure 8 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as network interfaces, display devices, etc.) than those shown in Figure 8 , or have a different configuration from that shown in Figure 8 .
[0159] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: flash drives, ROM, RAM, magnetic disks, or optical discs, etc.
[0160] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.
[0161] If the integrated unit in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in the storage medium and includes several instructions for causing one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0162] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0163] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the units or modules can be in an electrical or other form.
[0164] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution provided in this embodiment.
[0165] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist separately as individual physical units, or at least two units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0166] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for recognizing a voice signal, characterized in that, it includes: merging a set of voice signals into a first voice signal, where each voice signal in the set of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object; performing voice recognition on the first voice signal to obtain a set of candidate voice commands, where each candidate voice command in the set of candidate voice commands is a piece of text information recognized from the first voice signal; evaluating each candidate voice command based on the set of voice signals to obtain an evaluation result corresponding to each candidate voice command, where the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is a voice command issued by the target object; selecting a target voice command from the set of candidate voice commands based on the evaluation result corresponding to each candidate voice command.
2. The method according to claim 1, characterized in that, the merging of the set of voice signals into the first voice signal includes: starting from the signal acquisition time corresponding to the wake-up word in each voice signal in the set of voice signals, determining the voice signal segments to be merged in each voice signal other than the second voice signal in the set of voice signals based on the voice duration corresponding to the second voice signal with the highest confidence in the set of voice signals, to obtain a set of voice signal segments to be merged; merging the set of voice signal segments to be merged with the second voice signal to obtain the first voice signal.
3. The method according to claim 1, characterized in that, the performing of voice recognition on the first voice signal to obtain a set of candidate voice commands includes: inputting the first voice signal into a target encoder to obtain a feature vector corresponding to the first voice signal, where the target encoder is an encoder for encoding the input voice signal; inputting the feature vector corresponding to the first voice signal into a first decoder to obtain the set of candidate voice commands output by the first decoder, where the first decoder is a decoder for decoding the input feature vector into the corresponding text information.
4. The method according to claim 3, characterized in that, the evaluating of each candidate voice command based on the set of voice signals to obtain an evaluation result corresponding to each candidate voice command includes: using each candidate voice command as the current voice command to perform the following evaluation operation to obtain an evaluation result corresponding to each candidate voice command: inputting the current voice command and the feature vector corresponding to each voice signal into a second decoder to obtain the evaluation result corresponding to each voice signal output by the second decoder, where the second decoder is a decoder for evaluating the input voice command based on the matching degree between the input feature vector corresponding to the voice signal and the input voice command. Fuse the evaluation results corresponding to each of the voice signals to obtain an evaluation result corresponding to the current voice command.
5. The method according to claim 4, wherein, the fusing the evaluation results corresponding to each of the voice signals to obtain an evaluation result corresponding to the current voice command includes: performing weighted summation on the evaluation results corresponding to each of the voice signals according to the weights corresponding to each of the voice signals to obtain an evaluation result corresponding to the current voice command, wherein the magnitude of the weight corresponding to each of the voice signals is positively correlated with the confidence level of each of the voice signals.
6. The method according to claim 5, wherein, before evaluating each candidate voice command based on the set of voice signals, the method further includes: sorting the set of voice signals in descending order of confidence level to obtain a sorting result of the set of voice signals; Determine the weight corresponding to each of the voice signals according to the sorting result of the set of voice signals, where the weight corresponding to the nth voice signal in the set of voice signals is 1 / 2 n , and the weight corresponding to the last voice signal in the set of voice signals is 1 / 2 N-1 , N is the number of voice signals included in a set of voice signals, and 1 ≤ n < N.
7. The method according to any one of claims 1 to 6, wherein, before merging a set of voice signals into a first voice signal, the method further includes: acquiring voice signals collected by each of a plurality of voice devices to obtain a plurality of voice signals, wherein the voice signal collected by each of the voice devices is a voice signal corresponding to a voice command issued by the target object; selecting, from the plurality of voice signals, voice signals with a confidence level greater than or equal to a confidence threshold to obtain the set of voice signals.
8. The method according to claim 7, wherein, before selecting, from the plurality of voice signals, voice signals with a confidence level greater than or equal to a confidence threshold, the method further includes: evaluating the confidence level of each of the voice signals according to a set of voice parameters corresponding to each of the voice signals to obtain the confidence level of each of the voice signals, wherein the set of voice parameters includes at least one of the following: voice direction, voice intensity.
9. An identification device for voice signals, wherein, comprising: a merging unit, configured to merge a set of voice signals into a first voice signal, wherein each voice signal in the set of voice signals is a voice signal collected by a voice device and corresponding to a voice command issued by a target object; an identification unit, configured to perform voice recognition on the first voice signal to obtain a set of candidate voice commands, wherein each candidate voice command in the set of candidate voice commands is a text information identified from the first voice signal; a first evaluation unit, configured to evaluate each candidate voice command based on the set of voice signals to obtain an evaluation result corresponding to each candidate voice command, wherein the evaluation result corresponding to each candidate voice command is used to represent the probability that each candidate voice command is a voice command issued by the target object; a first selection unit, configured to select a target voice command from the set of candidate voice commands based on the evaluation results corresponding to each candidate voice command.
10. A computer-readable storage medium, wherein, The computer-readable storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 8.
11. An electronic device, comprising a memory and a processor, characterized in that the memory stores a computer program, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.