Speech recognition method, training method, device, electronic device and storage medium

By introducing semantic difference parameters into the speech recognition model and correcting and editing quantitative data, the problem of difficult to measure semantic bias in the prior art is solved, and the accuracy of performance evaluation and user experience of the speech recognition model is improved.

CN115359799BActive Publication Date: 2025-08-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210995296.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2025-08-08
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

Existing speech recognition technologies are difficult to differentiately measure the semantic bias caused by predicted text to user understanding, resulting in the recognition text being mistakenly considered to be the correct text output.

Method used

Using improved performance evaluation indicators, by introducing semantic difference parameters, correcting edited quantized data, measuring the error rate of predicted text relative to standard text, including identifying correct quantized data and editing operations, ensuring that performance evaluation indicators reflect the impact of semantic bias on user understanding.

Benefits of technology

It improves the accuracy of performance evaluation of the speech recognition model, avoids recognition text with large semantic deviations being mistaken for correct text output, and improves user understanding experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359799B_ABST
    Figure CN115359799B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speech recognition method, training method, device, electronic device, and storage medium. The method includes obtaining a target speech, determining a recognized text of the target speech based on a speech recognition model, and training the model in the following manner: using the model to predict the predicted text of the speech sample, and updating the model parameters of the model when the performance evaluation index of the model meets the iteration condition. The index parameters include the recognition quantization parameters and semantic difference parameters of the statistical objects contained in the predicted text. When the statistical object is correctly recognized, the recognition quantization parameters include the recognized correct quantization data. When the recognition is incorrect, the recognition quantization parameters include the edited quantization data of the editing operation, and the semantic difference parameters are used to correct the edited quantization data. In this case, the performance evaluation index can differentially measure the impact of the predicted text on the user's understanding, avoiding the recognition text with large semantic deviation from being mistakenly output as the correct text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech recognition method, training method, device, electronic equipment and storage medium. Background Art

[0002] In Automatic Speech Recognition (ASR) technology, performance indicators such as Character Error Rate (CER) and Word Error Rate (WER) are mainly used to evaluate the performance of ASR algorithms.

[0003] When training a speech recognition model, the model can predict text based on the recognized speech sample, and then compare the predicted text with the reference text to obtain evaluation index parameters. Based on the evaluation index parameters, the performance evaluation indicators of the ASR algorithm are determined. However, these performance evaluation indicators are difficult to differentiate and measure the impact of the predicted text on user understanding, resulting in recognized text with large deviations being mistakenly output as correct text. Summary of the Invention

[0004] According to one aspect of the present disclosure, there is provided a speech recognition method, comprising:

[0005] Obtain a target speech, and determine the recognized text of the target speech based on a speech recognition model, wherein the speech recognition model is trained in the following manner:

[0006] The speech recognition model is used to predict the predicted text of the speech sample. In response to the performance evaluation index of the speech recognition model satisfying the iteration condition, the model parameters of the speech recognition model are updated. The performance evaluation index is used to measure the error rate of the predicted text of the speech sample relative to the standard text. The parameters of the performance evaluation index include the recognition quantization parameter and semantic difference parameter of the statistical object contained in the predicted text. When the statistical object is correctly recognized, the recognition quantization parameter includes the correct recognition quantization data. When the statistical object is incorrectly recognized, the recognition quantization parameter includes the editing quantization data of the editing operation. The semantic difference parameter is used to correct the editing quantization data.

[0007] According to another aspect of the present disclosure, a training method is provided, comprising:

[0008] Use speech recognition models to predict text from speech samples;

[0009] In response to the performance evaluation index of the speech recognition model satisfying an iteration condition, updating a model parameter of the speech recognition model;

[0010] Among them, the performance evaluation index is used to measure the error rate of the predicted text of the speech sample relative to the standard text. The parameters of the performance evaluation index include the recognition quantization parameters and semantic difference parameters of the statistical objects contained in the predicted text. When the statistical object is recognized correctly, the recognition quantization parameters include the correct recognition quantization data. When the statistical object is recognized incorrectly, the recognition quantization parameters include the editing quantization data of the editing operation, and the semantic difference parameters are used to correct the editing quantization data.

[0011] According to another aspect of the present disclosure, there is provided a speech recognition device, comprising:

[0012] An acquisition module is used to acquire target speech;

[0013] The recognition module is used to determine the recognized text of the target speech based on the speech recognition model, wherein the speech recognition model is trained in the following manner:

[0014] The speech recognition model is used to predict the predicted text of the speech sample. In response to the performance evaluation index of the speech recognition model satisfying the iteration condition, the model parameters of the speech recognition model are updated. The performance evaluation index is used to measure the error rate of the predicted text of the speech sample relative to the standard text. The parameters of the performance evaluation index include the recognition quantization parameter and semantic difference parameter of the statistical object contained in the predicted text. When the statistical object is correctly recognized, the recognition quantization parameter includes the correct recognition quantization data. When the statistical object is incorrectly recognized, the recognition quantization parameter includes the editing quantization data of the editing operation. The semantic difference parameter is used to correct the editing quantization data.

[0015] According to another aspect of the present disclosure, there is provided a training device comprising:

[0016] A prediction module, used to predict the predicted text of the speech sample using the speech recognition model;

[0017] An updating module is used to update the model parameters of the speech recognition model in response to the performance evaluation index of the speech recognition model satisfying the iteration condition, wherein the performance evaluation index is used to measure the error rate of the predicted text of the speech sample relative to the standard text, and the parameters of the performance evaluation index include the recognition quantization parameter and semantic difference parameter of the statistical object contained in the predicted text. When the statistical object is correctly recognized, the recognition quantization parameter includes the correct recognition quantization data; when the statistical object is incorrectly recognized, the recognition quantization parameter includes the editing quantization data of the editing operation, and the semantic difference parameter is used to correct the editing quantization data.

[0018] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0019] processor; and,

[0020] Memory for storing programs;

[0021] The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to the exemplary embodiment of the present disclosure.

[0022] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to the exemplary embodiments of the present disclosure.

[0023] In one or more technical solutions provided in the exemplary embodiments of the present disclosure, when the speech recognition model is in the training phase, when the statistical object recognition error occurs, the recognition quantification parameter includes the editing quantification data of the editing operation, and the semantic difference parameter is used to correct the editing quantification data. Therefore, when determining the performance evaluation index based on the editing quantification data and the recognition correct quantification data, the semantic difference data can be used to correct the editing quantification data, so that the corrected editing quantification data can not only reflect the objective difference between the predicted text and the standard text, but also reflect the subjective difference between the predicted text and the standard text from the perspective of semantic understanding. Based on this, after determining the performance evaluation index based on the corrected editing quantification data and the recognition correct quantification data, when the performance evaluation index is used to evaluate the error rate of the predicted text of the speech sample relative to the standard text, the performance evaluation index can distinguish the impact of semantic deviation on the recognition result. Therefore, the performance evaluation index used by the method of the exemplary embodiment of the present disclosure can differentially measure the impact of the predicted text on the user's understanding, and avoid the recognition text with a large semantic deviation from being mistakenly output as the correct text. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0025] Figure 1 A schematic diagram illustrating an example system of various methods described in exemplary embodiments of the present disclosure;

[0026] Figure 2 An example flow chart showing a training method according to an exemplary embodiment of the present disclosure;

[0027] Figure 3 An example flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure is shown;

[0028] Figure 4 A schematic diagram showing the influence of sentiment shift and semantic shift on word error rate according to an exemplary embodiment of the present disclosure is shown;

[0029] Figure 5A schematic block diagram of functional modules of a speech recognition device according to an exemplary embodiment of the present disclosure is shown;

[0030] Figure 6 shows a schematic block diagram of functional modules of a training device according to an exemplary embodiment of the present disclosure;

[0031] Figure 7 shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure;

[0032] Figure 8 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0033] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0034] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0035] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0036] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0037] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0038] Before introducing the embodiments of the present disclosure, the following definitions are given for the relevant terms involved in the embodiments of the present disclosure:

[0039] Speech recognition technology, also known as Automatic Speech Recognition (ASR), aims to convert the vocabulary content in human speech into computer-readable input.

[0040] Word segmentation refers to the process of recombining a continuous sequence of characters into a word sequence according to certain specifications. This includes both English and Chinese word segmentation. In English word segmentation, spaces are used as natural delimiters between words, while words in Chinese sentences do not have formal delimiters. Chinese word segmentation is the process of breaking a sequence of Chinese characters into meaningful words. The result of word segmentation in the exemplary embodiments of the present disclosure is called word segmentation.

[0041] The edit distance, proposed by Soviet mathematician Vladimir Levenstein in 1965, describes the difference between two strings by calculating the minimum number of edits required to convert them into each other. Edit operations include substitution, deletion, and insertion. It is currently widely used in areas such as word error rate calculation, deoxyribonucleic acid (DNA) sequence alignment, and spelling detection.

[0042] The Character Error Rate (CER) is used to evaluate the word error rate between the predicted text and the standard text. The Word Error Rate (WER) is an important metric for evaluating ASR performance. It is used to evaluate the word error rate between the predicted text and the standard text.

[0043] Part-of-speech tagging, also known as part-of-speech tagging or simply tagging, is the process of assigning the correct part of speech to each word in the segmentation results. This process involves determining whether each word is a noun, verb, adjective, or another part of speech. This process is accomplished by the part-of-speech tagger algorithm.

[0044] Sentiment Analysis, also known as opinion mining, refers to the process of analyzing, processing, and extracting subjective texts with emotional connotations using natural language processing and text mining techniques.

[0045] The backpropagation algorithm is a method that uses the gradient descent method to optimize the network parameters of the neural network. It calculates the value of the loss function based on the value calculated by the neural network and the expected value, and then calculates the partial derivative of the loss function with respect to the model parameters, and finally updates the network parameters.

[0046] The model parameters include weight parameters and bias parameters. The weight parameters represent the slope of the hyperplane, and the bias parameters represent the intercept of the hyperplane.

[0047] The exemplary embodiments of the present disclosure provide a speech recognition method and a training method, wherein the performance evaluation index of the speech recognition model used therein during the training phase can not only measure the recognition error rate of the predicted text of the speech sample relative to the standard text, but also differentially reflect the impact of semantic interference on the recognition error rate, thereby ensuring differentiated measurement of the impact of the predicted text on user understanding and avoiding the recognized text with large semantic deviation from being mistakenly output as correct text.

[0048] Figure 1 FIG. 1 shows a schematic diagram of a system architecture according to an exemplary embodiment of the present disclosure. Figure 1 As shown, the system architecture 100 provided by the exemplary embodiment of the present disclosure includes: a user device 110 , an execution device 120 and a data storage system 130 .

[0049] like Figure 1 As shown, the user equipment 110 can communicate with the execution device 120 via a communication network. The communication network can be a wired communication network or a wireless communication network. The wired communication network can be a communication network based on power line carrier technology, and the wireless communication network can be a local area wireless network or a wide area wireless network. The local area wireless network can be a Wi-Fi wireless network, a Zigbee wireless network, a mobile communication network, or a satellite communication network.

[0050] like Figure 1 As shown, the user device 110 may include a computer, a mobile phone, or an intelligent terminal such as an information processing center. The user device 110 may serve as the initiator of voice recognition and initiate a request to the execution device 120. The execution device 120 may be a server with data processing capabilities such as a cloud server, a network server, an application server, and a management server, which is used to implement the voiceprint recognition method. A deep learning processor may be configured in the server. The deep learning processor may be a single-core deep learning processor (Deep Learning Processor-Singlecore, DLP-S) or a multi-core deep learning processor (Deep Learning Processor-Multicore, DLP-M). DLP-M is a multi-core extension based on DLP-S, which interconnects multiple DLP-S through a network-on-chip (Network-on-chip, Noc) and uses protocols such as multicast and inter-core synchronization for inter-core communication.

[0051] like Figure 1As shown, the above-mentioned data storage system 130 can be a general term, including local storage and a database for storing historical data. The database can be on the execution device 120, on other network servers, or on the data storage system 130. The data storage system 130 can be separate from the execution device 120 or integrated into the execution device 120. The data storage system 130 can not only input data uploaded by the user device 110, but also store program instructions, neuron data, etc. These neuron data can be trained data. In addition, the data storage system 130 can also store the processing results obtained by the execution device 120 (such as pre-processing the target speech to be recognized, intermediate processing results or recognition text) into the data storage system 130.

[0052] In practical applications, such as Figure 1 As shown, the user device 110 may have a voice collection function, allowing the user device 110 to not only initiate a request to the execution device 120 through the interactive interface, but also collect the target voice to be recognized and send the target voice to the execution device 120 via the communication network. Based on this, when the execution device 120 implements the voice recognition method, the recognized target voice can be obtained not only from the data storage system 130, but also from the user device 110 via the communication network. In addition, when the execution device 120 implements the voice recognition method, its voiceprint recognition results can not only be fed back to the user device 110 via the communication network, but can also be stored in the data storage system 130.

[0053] In related technologies, speech recognition technology has been widely used in various voice interaction scenarios. Speech recognition algorithms often use performance evaluation indicators such as character error rate and word error rate to measure the performance of speech recognition models. The smaller these performance evaluation indicators are, the better the performance of the speech recognition model.

[0054] Due to language differences, the same performance evaluation metric may differ between speech recognition models in different languages. For example, for English, the character error rate and word error rate are equivalent, both using words as units to evaluate the error rate of the predicted text relative to the standard text. For Chinese, the character error rate evaluates the error rate of the predicted text relative to the standard text using Chinese characters. The word error rate evaluates the error rate of the predicted text relative to the standard text using Chinese word segmentation units.

[0055] In practical applications, the formulas for calculating character error rate and word error rate are the same. Both can use the edit distance to predict the edit quantification parameter between the text and the standard text, that is, the number of edit operations. For example, the number of edit operations can include the number of replacement operations, the number of deletion operations, and the number of insertion operations. Then, the performance evaluation index can be determined using Formula 1.

[0056] The above-mentioned performance evaluation indicator can be a character error rate or a word error rate. In the Chinese speech recognition scenario, when the performance evaluation indicator is a character error rate, a single Chinese character can be used as the statistical object (a punctuation mark can be considered as a statistical object with Chinese characters as the unit), and the editing quantitative parameters between the predicted text and the standard text can be counted; when the performance evaluation indicator is a word error rate, a Chinese phrase can be used as the statistical object (a punctuation mark can be considered as a statistical object with phrases as the unit), and the editing quantitative parameters of the editing operation of the predicted text relative to the standard text can be counted.

[0057]

[0058] Among them, WER in the exemplary embodiment of the present disclosure specifies a performance evaluation index, and does not distinguish between performance evaluation indicators of statistical objects with characters as the smallest unit or words as the smallest unit. S represents the minimum number of replacements (that is, the number of replacement operations) that occur when the predicted text is transcribed into standard text, or the minimum number of statistical objects that need to be replaced when the predicted text is edited into standard text. D represents the minimum number of deletions (that is, the number of deletion operations) that occur when the predicted text is transcribed into standard text, or the minimum number of statistical objects that need to be deleted when the predicted text is edited into standard text. I represents the minimum number of insertions (that is, the number of insertion operations) that occur when the predicted text is transcribed into standard text, or the minimum number of statistical objects that need to be inserted when the predicted text is edited into standard text. C represents the number of words correctly recognized by the predicted text, or the maximum number of statistical objects that do not need to be changed in the predicted text.

[0059] When the statistical object is a word, S can be expressed as the minimum number of words that need to be replaced when editing the predicted text into the standard text, D can be expressed as the minimum number of words that need to be deleted when editing the predicted text into the standard text, I can be expressed as the minimum number of words that need to be inserted when editing the predicted text into the standard text, and C can be expressed as the maximum number of words that do not need to be changed in the predicted text.

[0060] As can be seen from Formula 1, when using Formula 1 to evaluate performance evaluation indicators such as character error rate and word error rate, the performance evaluation indicators can reflect the statistical object recognition error rate contained in the predicted text. However, in some cases, the performance evaluation indicators are difficult to "truly and fairly" represent the error rate of the predicted text of the speech sample relative to the standard text.

[0061] When the statistical objects contained in the predicted text are incorrectly identified, some recognition errors can be ignored, while some recognition errors may cause the user to have subjective understanding deviations. These understanding deviations may cause the user to misunderstand the meaning of the predicted text and even cause the user to have unnecessary negative emotions, thereby affecting the user's understanding experience.

[0062] The performance evaluation index used in the training phase of the speech recognition model involved in the speech recognition method provided by the exemplary embodiment of the present disclosure is an improved performance evaluation index, which introduces a semantic difference parameter in accordance with the subjective judgment of the user, so that the performance evaluation index can fully reflect the performance impact of the semantic difference on the speech recognition model, thereby ensuring that the trained semantic recognition model can differentially measure the impact of the predicted text on the user's understanding, and avoid the recognized text with large semantic deviation being mistaken for the correct text output.

[0063] The speech recognition model involved in the speech recognition method provided by the exemplary embodiment of the present disclosure can be trained by the server in the aforementioned execution device. Based on this, the exemplary embodiment of the present disclosure also provides a training method, which can be executed by the server or a chip in the server. For ease of understanding, the exemplary embodiment of the present disclosure uses the interaction between the user device and the server to describe the training method of the exemplary embodiment of the present disclosure.

[0064] Figure 2 An example flow chart of a training method according to an exemplary embodiment of the present disclosure is shown. Figure 2 As shown, the method of the exemplary embodiment of the present disclosure may include:

[0065] Step 201: The user device sends a voice sample to the server. The user device can act as a voice sample collection terminal to collect voice samples and upload the collected voice samples to the server included in the execution device through the communication network.

[0066] Step 202: The server uses the speech recognition model to predict the predicted text of the speech sample. For example, the server reads the neuron data of the speech recognition model stored in the data storage system, and uses the speech sample as the input of the neuron data to finally obtain the predicted sample of the speech sample.

[0067] Step 203: In response to the performance evaluation index of the speech recognition model meeting the iteration condition, the server updates the model parameters of the speech recognition model. The performance evaluation index of the exemplary embodiment of the present disclosure can be used to measure the error rate of the predicted text of the speech sample relative to the standard text, and can be obtained by the following example:

[0068] The server performs performance evaluation indicator parameter statistics based on the predicted text and the standard text of the speech sample, obtains the parameters of the performance evaluation indicator, and determines the performance evaluation indicator based on the parameters of the performance evaluation indicator.

[0069] The standard text of the exemplary embodiment of the present disclosure is the reference text of the speech sample, which can be regarded as the real text of the speech sample. It can be a manually annotated standard text or a recognized text output by a trained speech recognition model, and it is proven that the recognized text matches the input speech.

[0070] In actual applications, in response to the performance evaluation indicator satisfying an iteration condition, the model parameters of the speech recognition model are updated. The iteration condition may be that the performance evaluation indicator is less than a preset performance evaluation parameter. If the performance evaluation indicator is greater than or equal to the preset performance evaluation indicator, it indicates that the performance evaluation indicator satisfies the iteration condition, and a backpropagation algorithm may be used to update the model parameters of the speech recognition model. The model parameters may be weights, may include offset values, or may include both weights and offset values.

[0071] Exemplarily, the parameters of the performance evaluation indicators of the exemplary embodiments of the present disclosure may include recognition quantification parameters and semantic difference parameters of the statistical objects contained in the predicted text. The semantic difference parameters can correct the recognition quantification parameters, thereby ensuring that the performance evaluation indicators determined based on the corrected recognition quantification parameters can differentially measure the impact of the predicted text on user understanding, and avoid the recognition text with large semantic deviation being mistaken for correct text output.

[0072] For example, when the statistical object is correctly identified, the identification quantization parameter includes identifying correct quantization data; when the statistical object is incorrectly identified, the identification quantization parameter includes edited quantization data of the editing operation, and the semantic difference parameter is used to correct the edited quantization data.

[0073] The speech recognition method of the exemplary embodiment of the present disclosure can be executed by an electronic device or a chip in the electronic device, and the electronic device can be a server or a user device. When the electronic device is a server, the speech recognition model is deployed on the server side, and when the electronic device is a user device, the speech recognition model is deployed on the user device. The method of the present disclosure is described below with reference to the accompanying drawings. It should be understood that the relevant contents involved in the training method and speech recognition method of the exemplary embodiment of the present disclosure can be referenced to each other. As for the specific process of the training method and the speech recognition method, reference can be made to the relevant technology and will not be described as a key point.

[0074] Figure 3 FIG. 1 shows a flow chart of a speech recognition method according to an exemplary embodiment of the present disclosure. Figure 3 As shown, the speech recognition method of the exemplary embodiment of the present disclosure may include:

[0075] Step 301: Acquire the target voice. The target voice can be collected by the voice collector of the user device. The user device can pre-process the collected target voice. The pre-processing may include: sampling, quantizing and encoding the analog signal of the target voice at equal intervals to obtain the digital signal of the target voice. At the same time, before converting the analog signal of the target voice into a digital signal, the target voice can be filtered to avoid introducing unnecessary interference to the voice recognition. On this basis, the digital signal of the target voice can be processed by pre-emphasis, framing, windowing and endpoint detection.

[0076] Step 302: Determine the recognized text of the target speech based on the speech recognition model. The speech recognition model can be trained in the following manner: use the speech recognition model to predict the predicted text of the speech sample, and in response to the performance evaluation index of the speech recognition model meeting the iteration condition, update the model parameters of the speech recognition model. The performance evaluation index can be used to measure the error rate of the predicted text of the speech sample relative to the standard text. The parameters of the performance evaluation index include the recognition quantification parameter and semantic difference parameter of the statistical objects contained in the predicted text. It should be understood that the speech recognition model can be various models that can realize speech recognition, such as the Chinese language model (also known as the n-gram model), recurrent neural network, etc. The specific architecture can refer to the relevant technology and is not described in detail here.

[0077] The predicted text of the exemplary embodiment of the present disclosure may include one or more statistical objects, which is specifically related to the length of the predicted text.

[0078] When the number of statistical objects is one, the predicted text may include the identification quantification parameter and semantic difference parameter of one statistical object. When the number of statistical objects is multiple, the predicted text may include the identification quantification parameters and semantic difference parameters of multiple statistical objects.

[0079] Taking the Chinese speech recognition scenario as an example, the predicted text during the training phase is "Have you eaten?" When the statistical unit is Chinese characters, the statistical objects, that is, the number of Chinese characters, are 7, namely "today", "day", "you", "eat", "have", "may", and "?". Each of these Chinese characters or punctuation marks includes recognition quantification parameters and semantic difference parameters. When the statistical unit is phrases, the statistical objects, that is, the number of phrases, are 5, namely "today", "you", "eat", "may", and "?". Each of these phrases or punctuation marks includes recognition quantification parameters and semantic difference parameters.

[0080] During the training phase, statistical objects contained in the predicted text are either correctly or incorrectly recognized. When the statistical object is correctly recognized, the recognition quantization parameter may include correct recognition quantization data. When the statistical object is incorrectly recognized, the recognition quantization parameter may include edit quantization data of an edit operation, and the semantic difference parameter may be used to correct the edit quantization data.

[0081] For example, the recognition correctness quantification data of the exemplary embodiment of the present disclosure can be the number of words correctly recognized during statistical object recognition, or the maximum number of statistical objects that do not need to be changed in the predicted text. The editing quantification data of the exemplary embodiment of the present disclosure can refer to the related description of the editing quantification parameter or the number of editing operations mentioned above.

[0082] On this basis, when determining performance evaluation indicators based on edited quantized data and identified correct quantized data, semantic difference data can be used to correct the edited quantized data, so that the corrected edited quantized data can not only reflect the objective differences between the predicted text and the standard text, but also reflect the subjective differences between the predicted text and the standard text from the perspective of semantic understanding. Based on this, after determining the performance evaluation indicators based on the corrected edited quantized data and identified correct quantized data, when using the performance evaluation indicators to evaluate the error rate of the predicted text of the speech sample relative to the standard text, the performance evaluation indicators can distinguish the impact of semantic deviation on the recognition results. Therefore, the method of the exemplary embodiment of the present disclosure can differentially measure the impact of the predicted text on user understanding, avoiding the recognition text with large semantic deviation from being mistakenly output as correct text.

[0083] In practical applications, the performance evaluation index of the exemplary embodiments of the present disclosure is determined by the first parameter and the second parameter. The performance evaluation index is positively correlated with the first parameter, and negatively correlated with the second parameter. For example, the performance evaluation index can be equal to the ratio of the first parameter to the second parameter. Assuming the first parameter is M and the second parameter is N, the performance evaluation index = M / N, where M and N are both integers greater than or equal to 1.

[0084] The first parameter and the second parameter of the exemplary embodiment of the present disclosure may both be positively correlated with the semantic difference parameter. The first parameter may be determined by the edited quantization data and the semantic difference parameter, and the second parameter may be determined by the edited quantization data, the identified correct quantization data, and the semantic difference parameter.

[0085] In one possible implementation, the semantic difference parameter correction and editing quantization data of the exemplary embodiment of the present disclosure can be an additive method (i.e., a summation method). In this case, the performance evaluation index introduces the semantic difference parameter in an additive manner on the basis of Formula 1. The semantic difference parameter can be used to correct the edited quantization parameter in an additive manner, and then the corrected edited quantization parameter is used to determine the first parameter. The semantic difference parameter can be used to correct the edited quantization data and identify the correct quantization data in an additive manner, and the second parameter is determined based on the corrected edited quantization data and the corrected identified correct quantization data. In this case, the performance evaluation index satisfies Formula 2:

[0086]

[0087] E 总 The sum of the semantic difference parameters of the edited objects representing each editing operation can be viewed as a penalty term in Equation 2.

[0088] In practical applications, considering that it is difficult to determine the semantics expressed by a single Chinese character, therefore, if the statistical object is a statistical object with the character as the smallest unit, the semantic difference parameters of the editing objects belonging to the same word segmentation are shared. For a statistical object with the word segmentation as the smallest unit, the semantic difference parameters of the editing objects belonging to different word segmentations can be independent of each other.

[0089] Exemplarily, the editing object in the exemplary embodiment of the present disclosure may refer to the editing object involved before and after the editing operation. Based on this, the editing object may include the pre-editing object and the post-editing object. The essence of the editing operation may include a replacement operation, a deletion operation, and an insertion operation, and its purpose is to transcribe the predicted text into a standard text. Therefore, the word segmentation to which the editing object belongs may include the word segmentation to which the pre-editing object belongs in the predicted text and the word segmentation to which the post-editing object belongs in the standard sample.

[0090] When the editing operation is a replacement operation, the editing object may include the pre-replacement object and the post-replacement object. The pre-replacement object may be the statistical object that needs to be replaced in the predicted text. At this time, the word segmentation to which the pre-replacement object belongs may be the word segmentation to which the statistical object to be replaced belongs after the word segmentation operation of the predicted text. The post-replacement object may be the replacement content of the statistical object in the predicted text. At this time, the essence of the word segmentation to which the post-replacement object belongs is the word segmentation to which the post-replacement object belongs after the word segmentation operation of the standard text.

[0091] For example, the predicted text hyp = "我说经验我们去哪玩?", its word segmentation result is ['我','说', '经验', '我们', '去', '哪', '玩', '?'], and the standard text ref = "我说今夜我们去哪玩?", its word segmentation result is ['我','说', '今夜', '我们', '去', '哪', '玩', '?'].

[0092] If the statistical object is a statistical object with the character as the smallest unit, by comparing the predicted text hyp and the standard text ref, it can be seen that the pre-replacement objects included in the editing object of the editing operation include "经" and "验", and the post-replacement objects include "今" and "夜". It is necessary to replace "经" with "今" and "验" with "夜". It can be seen that it is necessary to perform two replacement operations to edit the predicted text hyp into the standard text ref. The word segmentation to which the pre-replacement object belongs in the first replacement operation and the second replacement operation is both "经验", and the word segmentation to which the post-replacement object belongs in the first replacement operation and the second replacement operation is both "今夜". Therefore, "经" and "验" share the same semantic difference parameter.

[0093] If the statistical object is a word-based statistical object, comparing the word segmentation results of the predicted text hyp with the word segmentation results of the standard text ref shows that the word "experience" in the predicted text hyp needs to be replaced with "tonight", thereby editing the predicted text hyp to the standard text ref. Based on this, the number of replacement operations is 1, and the editing objects of this replacement operation include the pre-replacement object "experience" and the post-replacement object "tonight". The pre-replacement object is the word segmentation of "experience" and the post-replacement object is the word segmentation of "tonight".

[0094] As for the delete and insert operations, since the delete operation targets the statistical objects contained in the predicted text that need to be deleted, the deleted objects can include the pre-deletion objects, that is, the statistical objects contained in the predicted text that need to be deleted. As for the insert operation, since the insert operation targets the content contained in the standard sample that needs to be inserted into the predicted text, the inserted objects can include the post-insertion objects, that is, the content contained in the standard sample that needs to be inserted into the predicted text.

[0095] For example, when correcting and editing quantized data using semantically different parameters, the exemplary embodiments of the present disclosure can prioritize the correction and editing of the quantized data based on the attributes of the semantically different parameters. Furthermore, when correcting and editing the quantized data, the quantized data can be prioritized from a single perspective or from multiple perspectives.

[0096] In practical applications, the impact of recognition errors can sometimes be ignored, but sometimes cannot be ignored, resulting in serious negative effects on user understanding, such as semantic deviation, sentiment deviation and other semantic understanding differences. Based on this, the semantic difference parameter of the exemplary embodiment of the present disclosure can introduce the impact of semantic deviation for editing quantization parameters from the perspective of syntactic analysis, and can also introduce the impact of sentiment deviation for editing quantization parameters from the perspective of sentiment analysis, and can also introduce the impact of semantic deviation and sentiment deviation for editing quantization parameters from the perspectives of syntactic analysis and sentiment analysis.

[0097] Taking word error rate as an example, Table 1 shows the impact of sentiment shift and semantic shift on word error rate. Figure 4 The following is a schematic diagram showing the effect of sentiment shift and semantic shift on word error rate in an exemplary embodiment of the present disclosure. It should be understood that WER in Table 1 represents word error rate, and the more plus signs after WER, the higher the word error rate.

[0098] Table 1 The influence of sentiment shift and semantic shift on word error rate

[0099]

[0100] Figure 4Curve a in the middle shows the influence of semantic shift on word error rate when there is no emotion shift, curve b shows the influence of semantic shift on word error rate when there is moderate emotion shift, and curve c shows the influence of semantic shift on word error rate when there is extreme emotion shift. Figure 4 From curves a, b, and c, we can see that for phrases with a certain degree of semantic deviation, the greater the degree of sentiment deviation, the higher the corresponding word error rate; and for phrases with a certain degree of sentiment deviation, the greater the degree of semantic deviation, the higher the corresponding word error rate.

[0101] In an optional embodiment, the semantic difference parameter of the exemplary embodiment of the present disclosure may include the editing importance of the editing operation when introducing the impact of semantic deviation from the perspective of syntactic analysis. The editing importance of the editing operation may reflect the degree of misunderstanding caused by the semantic error to the user.

[0102] The editing importance of the above-mentioned editing operation is positively correlated with the semantic qualification level of the segmentation of the editing object. In other words, the greater the semantic qualification of the segmentation of the statistical object to which the error is identified, the higher the importance of the statistical object.

[0103] The semantic qualification level of the exemplary embodiment of the present disclosure can be determined by the part-of-speech of the segmented word to which the statistical object of the recognition error belongs. Based on this, the predicted text can be segmented to obtain segmented words, and the number of segmented words is related to the segmentation method and the length of the predicted text. From a technical point of view, the segmentation method can be a dictionary-based segmentation method, a statistics-based segmentation method, a rule-based segmentation method, etc. From the perspective of the segmentation tool, the segmentation method can be divided into LAC segmentation and Jieba segmentation, etc. At the same time, the segmentation tool can also mark the part-of-speech of the segmented words, so as to obtain the part-of-speech of each segmented word.

[0104] In one example, the relationship between a part-of-speech (POS) and a semantic qualification level can be determined based on a rule. For example, a first mapping relationship between a semantic qualification level and a part-of-speech (POS) can be set. For statistical objects with identification errors, the corresponding semantic qualification level can be found from the first mapping relationship based on the part-of-speech (POS) to which the statistical object belongs. It should be understood that a semantic qualification level can correspond to one POS or multiple POS.

[0105] For example, a 4-Level gradient classification table can be used to set the mapping relationship between the semantic qualification level and the word part of speech included in the first mapping relationship. Table 2 shows a 4-Level gradient classification table of semantic qualification level and word part of speech. The part of speech tags and part of speech meanings shown in Table 2 are exemplary reference to the LAC word part of speech comparison table.

[0106] Table 2. 4-Level Gradient Classification of Semantic Definition Level and Part of Speech

[0107]

[0108] The 4-Level Gradient Classification of Semantic Qualification Levels and Parts of Speech shown in Table 2 defines the correspondence between semantic qualification levels and part of speech. It divides semantic qualification levels into four levels, with each level corresponding to a segmentation category. The lower the semantic qualification level, the smaller the semantic qualification level indicator, and the less semantically qualified the segmentation category.

[0109] For example, the first level of semantic qualifiers corresponds to redundant words, which have almost no semantic qualifiers. Therefore, the first level of semantic qualifiers has a quantified semantic qualifier value of 0. The second level of semantic qualifiers corresponds to weak qualifiers, which have a slight semantic qualifier value. Therefore, the second level of semantic qualifiers has a quantified semantic qualifier value of 1. The third level of semantic qualifiers corresponds to strong qualifiers, which have a relatively strong semantic qualifier value. Therefore, the third level of semantic qualifiers has a quantified semantic qualifier value of 2. The fourth level of semantic qualifiers corresponds to core words, which have a particularly strong semantic qualifier value. Therefore, the fourth level of semantic qualifiers has a quantified semantic qualifier value of 3. Therefore, the higher the semantic qualifier level, the higher the corresponding semantic qualifier value, and the two show a positive correlation. Since the editing importance of an edit operation is positively correlated with the semantic qualifier level of the edited word, the editing importance of the semantic qualifier edit operation can be determined by the quantified semantic qualifier value.

[0110] In one example, the relationship between the word part of speech and the semantic qualification level can be determined based on a neural network model. The neural network model can be a Transformer neural network model based on a self-attention mechanism, or a recurrent neural network model (RNN) or a long short-term memory network (LSTM).

[0111] During the training phase, the neural network model can be fed with the semantics of the word segmentation (i.e., the word segmentation results), the word part of speech, and the quantified values of the semantically qualified level as input. The neural network model is then used to determine the quantitative prediction value of the semantically qualified level based on the word segmentation semantics and word part of speech. A loss function is then used to determine the loss between the quantified prediction value and the quantified value of the annotation. If the loss is less than the preset loss, the neural network model training is complete. Otherwise, the backpropagation algorithm is required to update the model parameters of the neural network model. The loss function can be selected based on actual conditions and is not limited here.

[0112] During the inference phase, the multiple segmentation results (predicted text or standard text) included in the text segmentation results and their segmentation results can be constructed into a segmentation code sequence, where each segmentation code in the segmentation code sequence includes a segmentation and its part of speech. Taking a recurrent neural network or a long short-term memory network as an example, the segmentation code sequence is input into the recurrent neural network, which analyzes the long-term dependencies between the segmentation codes contained in the segmentation code sequence. The long-term dependencies are used to determine a quantitative value of the degree of semantic specificity of the segmentation contained in each segmentation code.

[0113] For example, for "I said where shall we go to play tonight?", its word segmentation results are ['I', 'said', 'tonight', 'we', 'go', 'where', 'play', '? '].

[0114] When determining the relationship between word parts of speech and semantic definiteness levels in a rule-based manner, you can refer to Table 2 to obtain the quantified semantic definiteness values of each word in the word segmentation result ['I', 'say', 'tonight', 'we', 'go', 'where', 'play', '? ']. For specific results, refer to Table 3.

[0115] Table 3 Quantified values of word segmentation levels determined based on rule

[0116] Word content Participle part-of-speech tags Quantified value of semantic qualification level I r 1 explain v 2 Tonight t 3 us r 1 go v 2 where r 1 Play v 2 ? w 0

[0117] When determining the relationship value between the word segmentation part of speech and the semantic qualification level based on the neural network model recognition method, the word segmentation result and the word segmentation result's part of speech can be used as the input value of the neural network model to predict the quantitative value of the semantic qualification level of each word contained in the word segmentation result.

[0118] Taking the recurrent neural network as an example, based on the word segmentation results ['I', 'say', 'tonight', 'we', 'go', 'where', 'play', '? '], a word segmentation encoding sequence {x1, x2, x3, x4, x5, x6, x7, x8} is constructed, where x1 is the word segmentation encoding of "I", x2 is the word segmentation encoding of "say", x3 is the word segmentation encoding of "tonight", x4 is the word segmentation encoding of "go", x5 is the word segmentation encoding of "where", x6 is the word segmentation encoding of "we", x7 is the word segmentation encoding of "play", and x8 is the word segmentation encoding of "?". These word segmentation encodings all include the word segmentation and the corresponding part of speech of the word segmentation. Table 4 shows the quantified value of the semantic qualification level of each word segmentation determined based on the recognition method of the neural network model.

[0119] Table 4 Quantitative values of semantic qualification levels of each word determined based on neural network model recognition

[0120] Word content Participle part-of-speech tags Quantified value of semantic qualification level I r 1 explain v 3 Tonight t 3 us r 1 go v 1 where r 1 Play v 2 ? w 0

[0121] By comparing Table 3 and Table 4, it can be found that when the word segmentation results are the same, there are certain differences in the quantization values of the semantic limitation levels of each word segment determined by the rule-based method and the method based on neural network model recognition. This is because when determining the quantization value of the semantic limitation level of a word segment by the method based on neural network model recognition, not only the semantics and word nature of "say" are considered, but also the long-term dependence relationship of the word segment in the sentence "I said where shall we go tonight?" is considered, so that the quantization value of the semantic limitation level of the word segment is closer to its actual semantic limitation level in the sentence "I said where shall we go tonight?".

[0122] When the editing importance of an editing operation is positively correlated with the semantic limitation level of the word segment to which the statistical object of the recognition error belongs, this positive correlation can be a linear positive correlation or a non-linear positive correlation. For example: P represents the editing importance of an editing operation, lev represents the semantic limitation level of the word segment to which the statistical object of the recognition error belongs. Taking Lev as the independent variable and P as the dependent variable, the relationship between P and Lev satisfies a quadratic function or a higher-order function. Using this quadratic function or higher-order function relationship, the influence of the semantic limitation level on the editing importance can be amplified. Based on this, when correcting the editing quantization data using the editing importance, the editing quantization data can more easily reflect the influence brought by the semantic shift, so as to ensure that when the predicted text has a semantic shift, the semantic shift error can be easily reflected by the performance evaluation index. The editing importance of the editing operation will be described below according to the type of the editing operation.

[0123] When the editing operation is a replacement operation, the editing importance P of the i-th replacement operation Si satisfies Equation 3:

[0124]

[0125] lev(w Si ) represents the quantization value of the semantic limitation level of the word segment to which the object before replacement of the i-th replacement operation belongs, lev(w Si )' represents the quantization value of the semantic limitation level of the object after replacement of the i-th replacement operation, n represents the total number of levels of the semantic limitation level, i is the serial number of the replacement operation required to transcribe the predicted text into the standard text, and i represents greater than or equal to 0 and less than or equal to the total number of replacement operations. It can be seen that the sum of the editing importances of all replacement operations

[0126] In practical applications, the total number of replacement operations can be determined based on the difference between the predicted text and the standard text. For example, when the total number of replacement operations is 0, it means that all statistical objects contained in the predicted text do not need to be replaced. When the total number of replacement operations is equal to the total number of statistical objects contained in the predicted text, it means that all statistical objects contained in the predicted text need to be replaced. From Formula 3, it can be seen that the editing importance P of the i-th replacement operation of the exemplary embodiment of the present disclosure is Si Essentially, the semantic qualification level of the object before and after replacement is taken into account and the two are averaged to more comprehensively evaluate the editing importance P of the i-th replacement operation. Si , thereby improving the accuracy of editing importance of replacement operations.

[0127] When the editing operation is a replacement operation, the editing importance of the j-th deletion operation is P Dj Satisfy formula 4:

[0128]

[0129] lev(w Dj ) represents the quantified value of the semantic qualification level of the word to which the jth deletion object belongs, n represents the total number of semantic qualification levels, j is the number of deletion operations required to transcribe the predicted text into the standard text, and j represents a value greater than or equal to 0 and less than or equal to the total number of deletion operations. It can be seen that the sum of the editing importance of all deletion operations is

[0130] In practical applications, the total number of deletion operations can be determined based on the difference between the predicted text and the standard text. For example, if the total number of deletion operations is 0, it means that all statistical objects contained in the predicted text do not need to be deleted. Since it is unlikely that all statistical objects will be deleted, the total number of deletion operations is less than the total number of statistical objects contained in the predicted text.

[0131] When the editing operation is an insert operation, the editing importance of the kth insert operation is P Ik Satisfy formula 4:

[0132]

[0133] lev(w Ik ) represents the quantified value of the semantic qualification level of the word to which the insertion object of the k-th insertion operation belongs, n represents the total number of semantic qualification levels, k is the insertion operation number required to transcribe the predicted text into the standard text, k represents a value greater than or equal to 0 and less than or equal to the total number of insertion operations. k is the insertion operation number required to transcribe the predicted text into the standard text, k represents a value greater than or equal to 0 and less than or equal to the total number of insertion operations. It can be seen that the sum of the editing importance of all insertion operations is

[0134] When the total number of insert operations is equal to 0, it means that all statistical objects contained in the standard text do not need to be inserted into the predicted text. Considering that it is impossible for all statistical objects contained in the standard text to need to be inserted into the predicted text, the total number of insert operations is less than the total number of statistical objects contained in the standard text.

[0135] In the exemplary embodiment of the present disclosure, when the performance evaluation index measures the error rate of the predicted text of the speech sample relative to the standard text, if the editing importance of the editing operation is used to correct the editing quantization data, the editing importance of the replacement operation can be regarded as a penalty item for the minimum replacement number S, the editing importance of the deletion operation can be regarded as a penalty item for the minimum deletion number D, and the editing importance of the insertion operation can be regarded as a penalty item for the minimum insertion number I. For example, if only the editing importance of the editing operation is considered to correct the editing quantization data, then E in Formula 2 总 It can satisfy formula 6:

[0136]

[0137] Substituting Equation 6 into Equation 2, we can obtain the performance evaluation index WER that satisfies Equation 7 when only considering the edit importance correction of edit quantization data using edit operations:

[0138]

[0139] If the predicted text is transcribed into the standard text without any replacement operation, the formula in formula 7 If no deletion operation is performed, If no insertion operation is performed,

[0140] In an alternative approach, the semantic difference parameter of the exemplary embodiment of the present disclosure introduces the impact of sentiment shift into the editing quantization parameter from the perspective of sentiment analysis. The semantic difference parameter may include sentiment difference data of the editing operation. The sentiment difference data of the editing operation may reflect the negative emotions caused by semantic errors in user understanding.

[0141] The emotional difference data of the editing operation includes the difference between the emotional quantification data of the object before the editing operation and the emotional quantification data of the object after the editing operation. The emotional quantification data can be determined by text sentiment analysis and expressed as an emotional percentage.

[0142] For the corresponding replacement operation, the emotional difference data Pe(Si,Si') of the i-th replacement operation satisfies Formula 8:

[0143] Pe(Si,Si')=|Si-Si'| Formula 8

[0144] Si is the emotional percentage of the object before the i-th replacement operation, and Si' is the emotional percentage of the object after the i-th replacement operation. The values of Si and Si' range from 0 to 1. At this time, when Si = Si', Pe(Si,Si') = 0, and when Si>Si' or Si <Si', Pe(Si,Si') is positive. Based on this, the sum of the emotional difference data of all replacement operations is As for the sentiment difference data between insert and delete operations, it can be ignored.

[0145] On this basis, if we only consider the impact of sentiment deviation on performance evaluation indicators, Substituting it into Formula 2, the obtained performance evaluation index WER satisfies Formula 9:

[0146]

[0147] Considering that semantics will inevitably shift when sentiment shifts, the exemplary embodiment of the present disclosure can use syntactic analysis to determine the semantic qualification level quantization value of the segmentation to which the editing object of the editing operation belongs, and then determine the importance of editing operations of different error types based on the semantic qualification level quantization value, and then introduce the sentiment difference data of the replacement operation, so that E in Formula 2 总 It can satisfy formula 10:

[0148]

[0149] Substituting Equation 10 into Equation 2, we can obtain the performance evaluation index of fusing semantic shift and sentiment shift when only considering the edit importance of the edit operation to correct the edit quantization data. The performance evaluation index WER satisfies Equation 11:

[0150]

[0151] In order to demonstrate that the method of the exemplary embodiment of the present disclosure can not only reflect the objective difference between the predicted text and the standard text, but also reflect the subjective difference between the predicted text and the standard text from the perspective of semantic understanding, the following is illustrated with examples.

[0152] Example 1

[0153] The standard text ref = "I said where should we go tonight?", and its segmentation result is ['I', 'said', 'tonight', 'we', 'go', 'where', 'play', '?']. The first predicted text hyp1 = "I said experience where should we go to play?", and its segmentation result is ['I', 'said', 'experience', 'we', 'go', 'where', 'play', '?']. The second predicted text hyp2 = "I said semen where should we go to play?", and its segmentation result is ['I', 'said', 'semen', 'we', 'go', 'where', 'play', '?']. Table 5 shows the comparative results of different performance evaluation indicators of Example 1.

[0154] Table 5 Comparison results of different performance evaluation indicators of Example 1

[0155] Example 1 Related technologiesWER Fusion Semantic Offset WER Fusion of semantic shift and sentiment shift WER (ref, hyp1) 0.125 0.222 0.232 (ref, hyp2) 0.125 0.222 0.247

[0156] (Ref, hyp1) represents an example of editing the first predicted text hyp1 into the standard text ref, which is referred to as the first case for the convenience of subsequent description; (Ref, hyp2) represents an example of editing the second predicted text hyp2 into the standard text ref, which is referred to as the second case for the convenience of subsequent description.

[0157] Example 2

[0158] The standard text ref = "Jiannan, I have another problem, which is that I hate studying.", and its segmentation result is ['Jiannan', ', ', 'I', 'also', 'a', 'problem', 'that's me', 'hate studying', '.']. The first predicted text hyp1 = "Jiannan, I have another problem, which is that I hate studying.", and its segmentation result is ['Jiannan', ', ', 'I', 'also', 'a', 'problem', 'that's me', 'hate studying', '.']. The second predicted text hyp2 = "Jiannan, I have another problem, which is that I hate studying.", and its segmentation result is ['Jiannan', ', ', 'I', 'also', 'a', 'problem', 'that's me', 'hate studying', '.']. Table 6 shows the comparative results of different performance evaluation indicators of Example 2.

[0159] Table 6 Comparison results of different performance evaluation indicators of Example 2

[0160] Example 2 Related technologiesWER Fusion Semantic Offset WER Fusion of semantic shift and sentiment shift WER (ref, hyp1) 0.111 0.2 0.209 (ref, hyp2) 0.111 0.2 0.230

[0161] (ref, hyp1) represents an example of editing the first predicted text hyp1 into the standard text ref, which is referred to as the first case for the convenience of subsequent description; (ref, hyp2) represents an example of editing the second predicted text hyp2 into the standard text ref, which is referred to as the second case for the convenience of subsequent description.

[0162] In Examples 1 and 2, when a user says the standard text "ref," the speech recognition model may transcribe it into two possible errors: the first predicted text "hyp1" and the second predicted text "hyp2." In both cases, in an office setting, the first predicted text "hyp1" can be considered a common error, with little semantic or emotional impact on the user. However, the second predicted text "hyp2" can cause considerable emotional distress to office workers in public settings, such as meetings, and is considered an emotional error.

[0163] When the WER of the related technology is unable to fully measure the performance of the speech recognition model, for example, in both the first and second cases, the WER of the related technology determined by Formula 1 in Example 1 is 0.125, and in Example 2, the WER of the related technology determined by Formula 1 is 0.111. These results fail to differentiate the semantic and emotional differences caused to users in different situations.

[0164] When the speech recognition model considers the impact of transcription errors on semantic shift on performance evaluation indicators, the performance evaluation indicators incorporate semantic shift, which can differentially reflect the degree of semantic shift. When the speech recognition model considers the impact of transcription errors on semantic shift and emotional shift on performance evaluation indicators, the performance evaluation indicators incorporate semantic shift and emotional shift, which can differentially reflect the degree of semantic shift and emotional shift.

[0165] As shown in Table 5, in Example 1, the fused semantic offset WER of the first case determined by Formula 7 is WER=0.222, and the fused semantic offset and sentiment offset WER of the first case determined by Formula 11 is WER=0.232; the fused semantic offset WER of the second case determined by Formula 7 is WER=0.222, and the fused semantic offset and sentiment offset WER=0.247 determined by Formula 11.

[0166] As shown in Table 6, in Example 2, the fused semantic offset WER of the first case determined by Formula 7 is WER=0.2, and the fused semantic offset and sentiment offset WER=0.209 determined by Formula 11; the fused semantic offset WER of the second case determined by Formula 7 is WER=0.2, and the fused semantic offset and sentiment offset WER=0.230 determined by Formula 11.

[0167] As can be seen from Tables 5 and 6, when the WER incorporates semantic offset but not emotional offset, in the same embodiment, the WER of the first and second cases is the same when the semantic offset is fused. However, there is a difference in the WER when the semantic offset is fused with the emotional offset. The WER of the first case, which is a common error, is smaller than the WER of the second case, which is an emotional error, which is consistent with the actual judgment.

[0168] It can be seen that the fused semantic offset WER can reflect the algorithm performance of the fused semantic calculation, and the fused semantic offset and sentiment offset WER can reflect the algorithm performance of the fusion of semantic syntactic analysis and sentiment calculation. When the semantic offset is not serious, the fused semantic offset WER is close to the related technology WER, while the fused semantic offset WER with extremely serious semantic offset is greater than the related technology WER. When both the semantic offset and the sentiment offset are not serious, the fused semantic offset and sentiment offset WER is close to the fused semantic offset and sentiment offset WER. When both the semantic offset and the sentiment offset are serious, the fused semantic offset and sentiment offset WER is greater than the fused semantic offset and sentiment offset WER. Therefore, the performance evaluation indicators used in the method of the exemplary embodiment of the present disclosure can differentially reflect the algorithm performance of the speech recognition model.

[0169] Example 3

[0170] Standard text ref="The relevant functions have already reached the regression testing stage, and some functional points on the server side are left." Its segmentation result is ['relevant', 'function', 'already', 'in', 'regression', 'testing', 'stage', 'already', ',', 'and', 'still', 'remaining', 'server side', 'of', 'some', 'function', 'points', '.']. First predicted text hyp1="The relevant functions have already reached the regression testing stage, and some functional points on the server side are left." Its segmentation result is ['relevant', 'function', 'already', 'in', 'regression', 'testing', 'stage', 'already', ',', 'still', 'remaining', 'server side', 'of', 'some', 'function', 'points', '.']. The second predicted text hyp2 = "The relevant functions are already in the testing phase, and there are still some functional points left on the server side." Its word segmentation result is ['relevant', 'function', 'already', 'in', 'testing', 'phase', 'already', ', ', ', 'then', 'still', 'remaining', 'server side', 'of', 'some', 'function', 'points', '.']. Table 7 shows the comparative results of different performance evaluation indicators of Example 3.

[0171] Table 7 Comparison results of different performance evaluation indicators of Example 3

[0172]

[0173] (ref, hyp1) represents an example of editing the first predicted text hyp1 into the standard text ref, which is referred to as the first case for the convenience of subsequent description; (ref, hyp2) represents an example of editing the second predicted text hyp2 into the standard text ref, which is referred to as the second case for the convenience of subsequent description.

[0174] In Example 3, when a user says the standard text "ref," the speech recognition model may transcribe it into two possible errors: the first predicted text "hyp1" and the second predicted text "hyp2." In the first case, the conjunction "then" is missing, which has a negligible impact on the understanding of the entire sentence. In the second case, the noun "return" is missing, which has a more significant impact on the understanding of the sentence. Furthermore, when the user interprets the recognition results in the first and second cases, the errors in the first and second cases do not affect the user's emotions.

[0175] When the WER of related technologies is used, it cannot fully measure the performance of the speech recognition model. For example, the WER of the related technologies determined using Formula 1 in Example 3 is 0.058, which cannot differentiate the different levels of semantic impact that errors in different content have on users. As shown in Table 7, in Example 3, the fused semantic offset WER of the first case determined using Formula 7 is 0.0625, and the fused semantic offset and sentiment offset WER determined using Formula 11 is 0.0625; the fused semantic offset WER of the second case determined using Formula 7 is 0.735, and the fused semantic offset WER of the second case determined using Formula 11 is 0.0735. It can be seen that the WER of the first case of Example 3 is the same regardless of whether the sentiment offset is integrated, and the WER of the second case of Example 3 is the same regardless of whether the sentiment offset is integrated. When the semantic offset is fused in the first and second cases, the fused semantic offset WER is the same. When the semantic offset and sentiment offset are fused in the first and second cases, the fused semantic offset and sentiment offset WER is equal to the fused semantic offset WER. It can be seen that when the predicted text of Example 3 is transcribed into standard text, the error forms in the two cases do not have sentiment offset and will not bring negative emotions to the user's understanding, which is the same as the actual judgment.

[0176] In an optional manner, the method of correcting the edited quantization data by using semantic difference parameters in the exemplary embodiment of the present disclosure may be to introduce semantic difference parameters in a weighted manner, the semantic difference parameters may be used to correct the edited quantization data in a weighted manner, and the first parameter may be determined using the corrected edited quantization data, the semantic difference parameters may be used to correct the edited quantization data and identify the correct quantization data in a weighted manner, and the second parameter may be determined based on the corrected edited quantization data and the corrected identified correct quantization data.

[0177] The semantic difference parameter of the exemplary embodiment of the present disclosure includes a statistical weighting parameter. If the predicted text contains a target word segment, the recognition quantization parameter of the statistical object belonging to the target word segment is greater than 1.

[0178] In practical applications, a keyword table can be established to define the mapping relationship between target keywords and statistical weighting parameters, obtaining a second mapping relationship. After segmenting the predicted text and the standard text, it is possible to query from the second mapping relationship whether the segment to which the statistical object belongs is a target segment. If it is a target segment, the statistical weighting parameter corresponding to the target segment can be determined from the second mapping relationship. On this basis, if the statistical object belonging to this segment is correctly recognized, when counting the number of correctly recognized words in the predicted text, when counting the statistical object belonging to this segment, not only add the count of the statistical object of this segment (the value of the count of the statistical object of this segment is 1) to the already counted number of correctly recognized words, but also weight the count of the statistical object belonging to this segment using the statistical weighting parameter.

[0179] For example, for the predicted text hyp = "你在学习人工只能吗", its segmentation result is ['你', '在', '学习', '人工只能', '吗'], and the standard text ref = "你在学习人工智能吗?", its segmentation result is ['你', '在', '学习', '人工智能', '吗'].

[0180] Taking phrases as the smallest unit, the number of correctly recognized segments is 4, the number of substitution operations is 1, and the number of deletion and insertion operations are both 0. That is, S = 1, D = I = 0, C = 4. When there is no statistical weighting parameter, the performance evaluation index WER is calculated using the formula of Equation 1 as WER = (1 + 0 + 0) / (1 + 0 + 4) = 1 / 5 = 0.2.

[0181] When transcribing the predicted text hyp into the standard text ref, if the object after substitution "人工智能" of the substitution operation exists in the keyword table and its corresponding statistical weighting parameter is 20, the performance evaluation index WER is calculated using the formula of Equation 1 as WER = (1 * 20 + 0 + 0) / (1 * 20 + 0 + 4) = 20 / 24 = 5 / 6 = 0.83. If the correctly recognized segment "在" exists in the target keywords and its corresponding statistical weighting parameter is 20, the performance evaluation index WER is calculated using the formula of Equation 1 as WER = (1 + 0 + 0) / (1 + 0 + 1 + 1 + 1 + 1 * 20) = 1 / 24 = 0.042. By comparison, it can be found that when transcribing the predicted text hyp into the standard text ref, a keyword table can be designed to assign statistical weighting parameters to keywords with substantial meanings, forming a second mapping relationship. If the statistical keyword table contains segments of the standard sample or the predicted sample and these segments have substantial meanings, then through the form of statistical weighting parameters, this situation can be reflected in the performance evaluation index.

[0182] It should be noted that when the semantic difference parameters of the exemplary embodiment of the present disclosure include statistical weighting parameters, the editing importance of the editing operation and the emotional difference data of the editing operation, in the process of calculating the editing importance of the editing operation and the emotional difference data of the editing operation, the S, D and I involved do not need to be weighted by statistical weighting parameters.

[0183] As can be seen from the above, in one or more technical solutions provided in the exemplary embodiments of the present disclosure, when the speech recognition model is in the training stage, when the statistical object recognition error occurs, the recognition quantification parameter includes the editing quantification data of the editing operation, and the semantic difference parameter is used to correct the editing quantification data. Therefore, when determining the performance evaluation index based on the editing quantification data and the recognition correct quantification data, the semantic difference data can be used to correct the editing quantification data, so that the corrected editing quantification data can not only reflect the objective difference between the predicted text and the standard text, but also reflect the subjective difference between the predicted text and the standard text from the perspective of semantic understanding. Based on this, after determining the performance evaluation index based on the corrected editing quantification data and the recognition correct quantification data, when using the performance evaluation index to evaluate the error rate of the predicted text of the speech sample relative to the standard text, the performance evaluation index can distinguish the impact of semantic deviation on the recognition result. Therefore, the performance evaluation index used by the method of the exemplary embodiment of the present disclosure can differentially measure the impact of the predicted text on the user's understanding, and avoid the recognition text with a large semantic deviation from being mistaken for the correct text output.

[0184] The above mainly introduces the solution provided by the embodiment of the present disclosure from the perspective of an electronic device. It is understandable that, in order to realize the above functions, the electronic device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0185] The embodiments of the present disclosure can divide the functional units of the electronic device according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present disclosure is schematic and is only a logical function division. In actual implementation, there may be other division methods.

[0186] In the case of dividing each functional module according to each function, an exemplary embodiment of the present disclosure provides a voice recognition method, which can be an electronic device or a chip applied to the electronic device. Figure 5 FIG. 1 shows a schematic block diagram of the functional modules of a speech recognition device according to an exemplary embodiment of the present disclosure. Figure 5 As shown, the speech recognition device 500 includes:

[0187] Acquisition module 501, used to acquire target speech;

[0188] The recognition module 502 is used to determine the recognized text of the target speech based on the speech recognition model, wherein the speech recognition model is trained in the following manner:

[0189] The speech recognition model is used to predict the predicted text of the speech sample. In response to the performance evaluation index of the speech recognition model satisfying the iteration condition, the model parameters of the speech recognition model are updated. The performance evaluation index is used to measure the error rate of the predicted text of the speech sample relative to the standard text. The parameters of the performance evaluation index include the recognition quantization parameter and semantic difference parameter of the statistical object contained in the predicted text. When the statistical object is correctly recognized, the recognition quantization parameter includes the correct recognition quantization data. When the statistical object is incorrectly recognized, the recognition quantization parameter includes the editing quantization data of the editing operation. The semantic difference parameter is used to correct the editing quantization data.

[0190] In one possible implementation, the performance evaluation index is determined by a first parameter and a second parameter, the performance evaluation index is positively correlated with the first parameter, the performance evaluation index is negatively correlated with the second parameter, and the semantic difference parameter is positively correlated with both the first parameter and the second parameter.

[0191] In a possible implementation, the semantic difference parameter includes the editing importance of the editing operation, and the editing importance of the editing operation is positively correlated with the semantic limitation level of the word to which the editing object of the editing operation belongs.

[0192] In a possible implementation, the semantic limitation level is determined by the part of speech of the segmented word to which the statistical object of the recognition error belongs.

[0193] In a possible implementation, the editing operation is a replacement operation, and the editing importance P of the i-th replacement operation is Si satisfy:

[0194] lev(w Si ) represents the semantic qualification level quantization value of the word to which the object belongs before the replacement operation of the i-th replacement operation, lev(w Si)' represents the quantized value of the semantic qualification level of the replaced object in the i-th replacement operation, and n represents the total number of semantic qualification levels.

[0195] In a possible implementation, the editing operation is a deletion operation, and the editing importance P of the jth deletion operation is Dj satisfy:

[0196] lev(w Dj ) represents the quantified value of the semantic qualification level of the word to which the deleted object of the j-th deletion operation belongs, and n represents the total number of semantic qualification levels.

[0197] In a possible implementation, the editing operation is an insert operation, and the editing importance P of the kth insert operation is Ik satisfy:

[0198] lev(w Ik ) represents the quantified value of the semantic qualification level of the word to which the insertion object of the k-th insertion operation belongs, and n represents the total number of semantic qualification levels.

[0199] In a possible implementation, the semantic difference parameter includes emotion difference data of the editing operation, and the emotion difference data of the editing operation includes a difference between emotion quantization data of the object before editing and emotion quantization data of the object after editing.

[0200] In a possible implementation, the semantic difference parameter includes a statistical weighting parameter of the editing operation. If the predicted text contains a target segmentation, the statistical weighting parameter is greater than 1 for the recognition quantization parameter of the statistical object belonging to the target segmentation.

[0201] In a possible implementation, the statistical object is a statistical object with segmentation as the smallest unit, and the semantic difference parameters of the editing objects belonging to different segmentations are independent of each other; or,

[0202] The statistical object is a statistical object with a character as the smallest unit. If the editing objects belong to the same word segmentation, the semantic difference parameter is shared.

[0203] In the case of dividing each functional module according to each function, an exemplary embodiment of the present disclosure provides a training method, which can be used for a server or a chip applied to the server. Figure 6 FIG. 1 shows a schematic block diagram of functional modules of a training device according to an exemplary embodiment of the present disclosure. Figure 6 As shown, the training device 600 includes:

[0204] Prediction module 601, used to predict the predicted text of the speech sample using the speech recognition model;

[0205] An updating module 602 is configured to update the model parameters of the speech recognition model in response to a performance evaluation indicator of the speech recognition model satisfying an iteration condition. The performance evaluation indicator is used to measure the error rate of the predicted text of the speech sample relative to the standard text. The parameters of the performance evaluation indicator include recognition quantification parameters and semantic difference parameters of the statistical objects contained in the predicted text. When the statistical object is correctly identified, the recognition quantification parameters include correct recognition quantification data. When the statistical object is incorrectly identified, the recognition quantification parameters include editing quantification data of the editing operation. The semantic difference parameters are used to correct the editing quantification data.

[0206] Figure 7 FIG. 1 shows a schematic block diagram of a chip according to an exemplary embodiment of the present disclosure. Figure 7 As shown, the chip 700 includes one or more (including two) processors 701 and a communication interface 702. The communication interface 702 can perform the acquisition step in the above method, and the processor 701 can perform the processing step in the above method.

[0207] Optional, such as Figure 7 As shown, the chip 700 also includes a memory 703, which may include a read-only memory and a random access memory, and provides operating instructions and parameters to the processor. A portion of the memory may also include a non-volatile random access memory (NVRAM).

[0208] In some embodiments, as Figure 7 As shown, the processor 701 performs corresponding operations by calling the operation instructions stored in the memory (the operation instructions may be stored in the operating system). The processor 701 controls the processing operations of any one of the terminal devices, and the processor may also be called a central processing unit (CPU). The memory 703 may include a read-only memory and a random access memory, and provides instructions and parameters to the processor 701. A portion of the memory 703 may also include NVRAM. For example, in an application, the memory, the communication interface, and the memory are coupled together through a bus system, wherein the bus system may include a power bus, a control bus, and a status signal bus in addition to a parameter bus. However, for the sake of clarity, in Figure 7 Various buses are labeled as bus system 704 .

[0209] The methods disclosed in the above embodiments of the present disclosure can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0210] The exemplary embodiments of the present disclosure further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being configured to cause the electronic device to perform a method according to an exemplary embodiment of the present disclosure when executed by the at least one processor.

[0211] Exemplary embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present disclosure.

[0212] Exemplary embodiments of the present disclosure further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to perform the method according to the embodiment of the present disclosure.

[0213] refer to Figure 8, a block diagram of an electronic device that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0214] like Figure 8 As shown, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and parameters required for the operation of device 800 can also be stored in RAM 803. Computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0215] like Figure 8 As shown, multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, output unit 807, storage unit 808, and communication unit 809. Input unit 806 can be any type of device capable of inputting information into electronic device 800. Input unit 806 can receive input digital or character information and generate key signal input related to user settings and / or function control of the electronic device. Output unit 807 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 808 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 809 allows electronic device 800 to exchange information with other devices via computer networks such as the Internet and / or various telecommunication networks and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth device, a Wireless Fidelity (WiFi) device, a Worldwide Interoperability for Microwave Access (WiMax) device, a cellular communication device, and / or the like.

[0216] like Figure 8As shown, the computing unit 801 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (Central Processing Unit / Processor, abbreviated as CPU), a graphics processing unit (Graphics Processing Unit, GPU), various dedicated artificial intelligence (Artificial Intelligence, AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (Digital Signal Processing, DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above. For example, in some embodiments, the method of the exemplary embodiments of the present disclosure can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 can be configured to perform the method of the exemplary embodiments of the present disclosure by any other appropriate means (e.g., by means of firmware).

[0217] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0218] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0219] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube or liquid crystal display monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0220] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0221] A computer system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0222] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present disclosure are performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user device, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).

[0223] Although the present disclosure has been described with reference to specific features and embodiments thereof, it will be apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present disclosure. Accordingly, this specification and the drawings are merely illustrative of the present disclosure as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present disclosure. Obviously, those skilled in the art may make various modifications and variations to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, the present disclosure is intended to include such modifications and variations if they fall within the scope of the claims of the present disclosure and their equivalents.

Claims

1. A speech recognition method, characterized in that: include: Obtain a target speech, and determine the recognized text of the target speech based on a speech recognition model, wherein the speech recognition model is trained in the following manner: Predicting a predicted text of a speech sample using the speech recognition model, and updating model parameters of the speech recognition model in response to a performance evaluation indicator of the speech recognition model satisfying an iteration condition, wherein the performance evaluation indicator is used to measure an error rate of the predicted text of the speech sample relative to a standard text, and the parameters of the performance evaluation indicator include recognition quantization parameters and semantic difference parameters of statistical objects contained in the predicted text, wherein when the statistical object is correctly recognized, the recognition quantization parameters include recognition correctness quantization data, and when the statistical object is incorrectly recognized, the recognition quantization parameters include edited quantization data of an edit operation, and the semantic difference parameters are used to correct the edited quantization data; The semantic difference parameter includes the editing importance of the editing operation, which is positively correlated with the semantic qualification level of the word segment to which the editing object of the editing operation belongs. The editing importance of the editing operation is determined by the quantified value of the semantic qualification level, and the semantic qualification level is determined by the part of speech of the word segment to which the statistical object of the recognition error belongs.

2. The method according to claim 1, characterized in that The performance evaluation index is determined by a first parameter and a second parameter, the performance evaluation index is positively correlated with the first parameter, the performance evaluation index is negatively correlated with the second parameter, and the semantic difference parameter is positively correlated with both the first parameter and the second parameter.

3. The method according to claim 1, characterized in that The semantic difference parameter further includes emotion difference data of the editing operation, and the emotion difference data of the editing operation includes a difference between the emotion quantization data of the object before editing and the emotion quantization data of the object after editing.

4. The method according to any one of claims 1 to 3, characterized in that The semantic difference parameter also includes a statistical weighting parameter of the editing operation. If the predicted text contains a target segmentation, the statistical weighting parameter is greater than 1 for the recognition quantization parameter of the statistical object belonging to the target segmentation.

5. The method according to any one of claims 1 to 3, characterized in that The statistical object is a statistical object with segmentation as the smallest unit, and the semantic difference parameters of the editing objects belonging to different segmentations are independent of each other; or, The statistical object is a statistical object with a character as the smallest unit. If the editing objects belong to the same word segmentation, the semantic difference parameter is shared.

6. A training method, characterized in that: include: Use speech recognition models to predict text from speech samples; In response to the performance evaluation index of the speech recognition model satisfying an iteration condition, updating a model parameter of the speech recognition model; The performance evaluation index is used to measure the error rate of the predicted text of the speech sample relative to the standard text. The parameters of the performance evaluation index include recognition quantization parameters and semantic difference parameters of the statistical objects contained in the predicted text. When the statistical objects are correctly recognized, the recognition quantization parameters include recognition correct quantization data. When the statistical objects are incorrectly recognized, the recognition quantization parameters include edit quantization data of the editing operation, and the semantic difference parameters are used to correct the edit quantization data. The semantic difference parameter includes the editing importance of the editing operation, which is positively correlated with the semantic qualification level of the word segment to which the editing object of the editing operation belongs. The editing importance of the editing operation is determined by the quantified value of the semantic qualification level, and the semantic qualification level is determined by the part of speech of the word segment to which the statistical object of the recognition error belongs.

7. A speech recognition device, characterized in that: include: An acquisition module is used to acquire the target speech; The recognition module is used to determine the recognized text of the target speech based on the speech recognition model, wherein the speech recognition model is trained in the following manner: Predicting a predicted text of a speech sample using the speech recognition model, and updating model parameters of the speech recognition model in response to a performance evaluation indicator of the speech recognition model satisfying an iteration condition, wherein the performance evaluation indicator is used to measure an error rate of the predicted text of the speech sample relative to a standard text, and the parameters of the performance evaluation indicator include recognition quantization parameters and semantic difference parameters of statistical objects contained in the predicted text, wherein when the statistical object is correctly recognized, the recognition quantization parameters include recognition correctness quantization data, and when the statistical object is incorrectly recognized, the recognition quantization parameters include edited quantization data of an edit operation, and the semantic difference parameters are used to correct the edited quantization data; The semantic difference parameter includes the editing importance of the editing operation, which is positively correlated with the semantic qualification level of the word segment to which the editing object of the editing operation belongs. The editing importance of the editing operation is determined by the quantified value of the semantic qualification level, and the semantic qualification level is determined by the part of speech of the word segment to which the statistical object of the recognition error belongs.

8. A training device, characterized in that: include: A prediction module, used to predict the predicted text of the speech sample using the speech recognition model; an updating module, configured to update model parameters of the speech recognition model in response to a performance evaluation indicator of the speech recognition model satisfying an iteration condition, wherein the performance evaluation indicator is used to measure an error rate of a predicted text of a speech sample relative to a standard text, and the parameters of the performance evaluation indicator include a recognition quantization parameter and a semantic difference parameter of a statistical object contained in the predicted text; when the statistical object is correctly recognized, the recognition quantization parameter includes correct recognition quantization data; when the statistical object is incorrectly recognized, the recognition quantization parameter includes edit quantization data of an edit operation, and the semantic difference parameter is used to correct the edit quantization data; The semantic difference parameter includes the editing importance of the editing operation, which is positively correlated with the semantic qualification level of the word segment to which the editing object of the editing operation belongs. The editing importance of the editing operation is determined by the quantified value of the semantic qualification level, and the semantic qualification level is determined by the part of speech of the word segment to which the statistical object of the recognition error belongs.

9. An electronic device, characterized in that: include: processor; as well as Memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech recognition method and device, equipment and storage medium

    CN114520001A