Model loss value determination method and device, equipment and medium
By preprocessing and dynamically adjusting the model loss value in the speech recognition performance evaluation of deep learning models, the problem of low accuracy of model loss value in the prior art is solved, and more accurate performance evaluation is achieved.
Patent Information
- Application Number
- CN202510274361.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-27
AI Technical Summary
When determining the speech recognition performance of deep learning models, the prior art directly processes the predicted text sequence, real text sequence and related numerical values of speech, resulting in cumulative errors in model loss values and low accuracy.
By inputting the input speech from the speech dataset to the target prediction model, a sequence of predicted texts is obtained, and preprocessed according to their length and adjusted to the same length. At the same time, dynamically adjust the value related to the model loss value to ensure that it is within the available value range. Then, using the block calculation unit, based on the preprocessed text sequence and the dynamically adjusted numerical values, the model loss value is calculated and the model loss value is finally determined.
It effectively avoids the cumulative error of model loss value, improves the accuracy of model loss value, and makes the performance evaluation of deep learning models more accurate.
Smart Images

Figure CN120048266A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a method, apparatus, device, and medium for determining a model loss value. Background Art
[0002] With the development of artificial intelligence technologies, more and more enterprises have begun to perform speech recognition through deep learning models. The deep learning model trained based on a large amount of data can recognize the input speech, convert the speech into a text sequence for representing the speech content of the speech, and output the text sequence obtained by converting the speech.
[0003] In the process of training and using a deep learning model, it is necessary to evaluate the performance of the deep learning model according to a specified number of speeches, and determine the model loss value of each speech. The predicted text sequence of a speech is the text sequence output by the deep learning model obtained by inputting the speech into the deep learning model for representing the speech content of the speech. The true text sequence of a speech is the correct text sequence for representing the speech content of the speech. The model loss value is a numerical value determined according to the predicted text sequence and the true text sequence of the speech for representing the difference between the predicted text sequence of the speech output by the deep learning model and the true text sequence of the speech. The smaller the model loss value, the smaller the difference between the predicted text sequence of the speech output by the deep learning model and the true text sequence of the speech, the better the performance of the deep learning model, and the ability to accurately convert the input speech into a text sequence for representing the speech content of the speech. The larger the model loss value, the larger the difference between the predicted text sequence of the speech output by the deep learning model and the true text sequence of the speech, the poorer the performance of the deep learning model, and the inability to accurately convert the input speech into a text sequence for representing the speech content of the speech, and further adjustment is required.
[0004] In the related art, the commonly used solution for determining the model loss value is: input each speech into the deep learning model to obtain the predicted text sequence of each speech, and then directly determine the model loss value calculation parameter of each speech according to the predicted text sequence, true text sequence, and related numerical values of each speech, and further determine the model loss value of each speech according to the model loss value calculation parameter of each speech. There may be large differences in the lengths of the predicted text sequences and true text sequences of each speech and the numerical sizes of the related numerical values of each speech. Directly processing the predicted text sequences, true text sequences, and related numerical values of each speech with large differences will result in cumulative errors in the determined model loss value of the speech. The solution for determining the model loss value in the related art directly processes the predicted text sequences, true text sequences, and related numerical values of each speech, and cannot avoid the cumulative errors in the determined model loss value of the speech, resulting in low accuracy of the determined model loss value of the speech. Summary of the Invention
[0005] The present invention provides a method, apparatus, device and medium for determining a model loss value, so as to solve the problem that the existing method for determining the model loss value directly processes the predicted text sequences, true text sequences and related numerical values of each voice, and cannot avoid the cumulative error of the determined model loss value of the voice, resulting in low accuracy of the determined model loss value of the voice.
[0006] According to one aspect of the present invention, there is provided a method for determining a model loss value, including:
[0007] Inputting each input voice in the voice dataset corresponding to the target prediction model into the target prediction model to obtain the predicted text sequence of each input voice;
[0008] Preprocessing the predicted text sequence and true text sequence of each input voice according to the lengths of the predicted text sequence and true text sequence of each input voice;
[0009] Dynamically adjusting the numerical values related to the model loss value of each input voice to adjust the numerical values related to the model loss value of each input voice to the available numerical value range;
[0010] For each input voice, through each block calculation unit, calculating according to the preprocessed predicted text sequence, preprocessed true text sequence and dynamically adjusted numerical values related to the model loss value of the input voice to obtain the model loss value calculation parameters of the input voice;
[0011] Determining the model loss value of each input voice according to the model loss value calculation parameters of each input voice.
[0012] According to another aspect of the present invention, there is provided a device for determining a model loss value, including:
[0013] A sequence determination module, configured to input each input voice in the voice dataset corresponding to the target prediction model into the target prediction model to obtain the predicted text sequence of each input voice;
[0014] A sequence preprocessing module, configured to preprocess the predicted text sequence and true text sequence of each input voice according to the lengths of the predicted text sequence and true text sequence of each input voice;
[0015] A numerical value dynamic adjustment module, configured to dynamically adjust the numerical values related to the model loss value of each input voice to adjust the numerical values related to the model loss value of each input voice to the available numerical value range;
[0016] A parameter calculation module, configured to calculate, for each input speech, through each block calculation unit, based on the preprocessed predicted text sequence of the input speech, the preprocessed true text sequence, and the values related to the dynamically adjusted model loss value, to obtain the model loss value calculation parameters of the input speech;
[0017] A loss value calculation module, configured to determine the model loss values of the respective input speeches according to the model loss value calculation parameters of the respective input speeches.
[0018] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0019] At least one processor;
[0020] And a memory communicatively connected to the at least one processor;
[0021] Wherein, the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for determining the model loss value according to any embodiment of the present invention.
[0022] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the method for determining the model loss value according to any embodiment of the present invention when executed.
[0023] In the technical solution of the embodiment of the present invention, each input voice in the voice dataset corresponding to the target prediction model is input into the target prediction model to obtain the predicted text sequence of each input voice; then, according to the lengths of the predicted text sequence and the true text sequence of each input voice, preprocessing is performed on the predicted text sequence and the true text sequence of each input voice; the numerical values related to the model loss value of each input voice are dynamically adjusted to adjust the numerical values related to the model loss value of each input voice to the available numerical range; for each input voice, through each block calculation unit, calculations are performed according to the preprocessed predicted text sequence of the input voice, the preprocessed true text sequence, and the dynamically adjusted numerical values related to the model loss value to obtain the calculation parameters of the model loss value of the input voice; according to the calculation parameters of the model loss value of each input voice, the model loss value of each input voice is determined, solving the problem that the determination scheme of the model loss value in the related technology directly processes the predicted text sequence, the true text sequence, and the related numerical values of each voice, and it is impossible to avoid the cumulative error of the model loss value of the determined voice, resulting in a low accuracy of the determined model loss value of the voice. It is possible to preprocess the predicted text sequence and the true text sequence of each voice according to the lengths of the predicted text sequence and the true text sequence of each voice, adjust the lengths of the predicted text sequence and the true text sequence of each voice to the same length, and dynamically adjust the numerical values related to the model loss value of each voice to adjust the numerical values related to the model loss value of each voice to the available numerical range. Furthermore, based on each block calculation unit and the preprocessed predicted text sequence, the preprocessed true text sequence, and the dynamically adjusted numerical values related to the model loss value of each voice, the calculation parameters of the model loss value of each voice are determined, and then according to the calculation parameters of the model loss value of each voice, the model loss value of each voice is determined, thereby effectively avoiding the cumulative error of the model loss value of the voice determined by processing the predicted text sequence, the true text sequence, and the numerical values related to the model loss value of each voice, and improving the accuracy of the determined model loss value of the input voice.
[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 Flowchart of a method for determining a model loss value provided in Embodiment 1 of the present invention.
[0027] Figure 2 Flowchart of a method for determining a model loss value provided in Embodiment 2 of the present invention.
[0028] Figure 3 Schematic structural diagram of a device for determining a model loss value provided in Embodiment 3 of the present invention.
[0029] Figure 4 Schematic structural diagram of an electronic device for implementing the method for determining a model loss value in an embodiment of the present invention. Detailed implementation manners
[0030] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] It should be noted that the terms "target", "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising", "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] It should be noted that the relevant information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards in the relevant regions.
[0033] Embodiment 1
[0034] Figure 1The flowchart of a method for determining the model loss value provided in the first embodiment of the present invention. This embodiment is applicable to the process of training and using a deep learning model for speech recognition, evaluating the performance of the deep learning model according to a specified number of speeches, and determining the model loss value of each speech. This method can be executed by a device for determining the model loss value, which can be implemented in the form of hardware and / or software, and the device for determining the model loss value can be configured in an electronic device. The electronic device can be an artificial intelligence (AI) chip used for training and using the deep learning model. As Figure 1 shown, the method includes:
[0035] Step 101, input each input speech in the speech dataset corresponding to the target prediction model into the target prediction model to obtain the predicted text sequence of each input speech.
[0036] Optionally, the target prediction model is a deep learning model for speech recognition that needs to be evaluated for performance. The target prediction model is used to recognize the input speech, convert the speech into a text sequence for characterizing the speech content of the speech, and output the text sequence obtained by converting the speech. The text sequence can be a sequence composed of one or more characters. The predicted text sequence of the speech is the text sequence output by the target prediction model obtained by inputting the speech into the target prediction model for characterizing the speech content of the speech. The input of the target prediction model is the speech, and the output of the target prediction model is the predicted text sequence of the speech. The true text sequence of the speech is the correct text sequence for characterizing the speech content of the speech.
[0037] Optionally, the speech dataset corresponding to the target prediction model consists of multiple input speeches. Each input speech is a pre-collected speech for evaluating the performance of the target prediction model. Each input speech in the speech dataset is a speech authorized by the user or fully authorized by all parties, and the collection, use, and processing of the input speech comply with the relevant laws, regulations, and standards of the relevant region. The predicted text sequence of the input speech is the text sequence output by the target prediction model obtained by inputting the input speech into the target prediction model for characterizing the speech content of the speech. The true text sequence of the input speech is the correct text sequence for characterizing the speech content of the input speech. The speech dataset corresponding to the target prediction model and the true text sequence of each input speech are stored in the electronic device. The speech dataset corresponding to the target prediction model and the true text sequence of each input speech can be stored in the electronic device by technicians.
[0038] Optionally, after receiving the evaluation indication information corresponding to the target prediction model, each input speech in the speech dataset corresponding to the target prediction model can be input into the target prediction model to obtain the predicted text sequence of each input speech. The evaluation indication information corresponding to the target prediction model can be information used to indicate evaluating the performance of the target prediction model according to each input speech in the speech dataset corresponding to the target prediction model and determining the model loss value of each input speech. The evaluation indication information corresponding to the target prediction model can be sent by a technician through a terminal device.
[0039] Optionally, for each input speech in the speech dataset corresponding to the target prediction model, the input speech can be input into the target prediction model. The target prediction model will recognize the input speech and convert the input speech into a text sequence for representing the speech content of the input speech, that is, the predicted text sequence of the input speech, and then output the predicted text sequence of the input speech. The predicted text sequence of the input speech output by the target prediction model can be obtained, so as to obtain the predicted text sequence of the input speech.
[0040] Step 102: Preprocess the predicted text sequences and real text sequences of each input speech according to the lengths of the predicted text sequences and real text sequences of each input speech.
[0041] Optionally, the length of the text sequence can refer to the total number of characters included in the text sequence.
[0042] Optionally, preprocessing the predicted text sequences and real text sequences of each input speech according to the lengths of the predicted text sequences and real text sequences of each input speech includes: determining the maximum value among the lengths of the predicted text sequences and real text sequences of each input speech; performing text insertion on the text sequences with lengths less than the maximum value among the predicted text sequences and real text sequences of each input speech, and adjusting the lengths of the text sequences with lengths less than the maximum value to the maximum value.
[0043] Optionally, the lengths of the predicted text sequences of each input speech and the lengths of the real text sequences of each input speech can be detected, and then the maximum value among the lengths can be counted, so as to determine the maximum value among the lengths of the predicted text sequences and real text sequences of each input speech.
[0044] Optionally, perform text insertion on the predicted text sequences and the true text sequences of each input speech for the text sequences with lengths less than the maximum value, and adjust the lengths of the text sequences with lengths less than the maximum value to the maximum value, including: performing the following operations on each text sequence with a length less than the maximum value in the predicted text sequences and the true text sequences of each input speech: inserting one or more characters '0' at the end of the text sequence to increase the length of the text sequence to the maximum value.
[0045] Thus, preprocessing can be performed on the predicted text sequences and the true text sequences of each input speech, inserting one or more characters '0' at the end of the text sequences with smaller lengths to increase the lengths of the text sequences to the maximum value of the lengths of each text sequence, so as to adjust the lengths of the predicted text sequences and the true text sequences of each input speech to the same length without affecting the speech content represented by each text sequence.
[0046] Step 103: Dynamically adjust the values related to the model loss value of each input speech, and adjust the values related to the model loss value of each input speech to the available value range.
[0047] Optionally, the embodiments of the present invention use the Connectionist Temporal Classification (CTC) algorithm to determine the model loss value of speech. The model loss value of each speech is calculated according to the calculation parameters of the model loss value of the speech. The calculation parameters of the predicted text sequence model loss value of the speech may include: the probability that the predicted text sequence of the speech is the true text sequence of the speech, the forward recurrence probability and the backward recurrence probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech. Where k = 1, 2,..., N. N is the total number of characters included in the predicted text sequence of the speech. The forward recurrence probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech may refer to the probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech calculated using the forward probability algorithm. The backward recurrence probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech may refer to the probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech calculated using the backward probability algorithm. The forward recurrence probability and the backward recurrence probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech are the forward recurrence probability and the backward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech. The numerical values related to the model loss value of the speech are the numerical values required for calculating the calculation parameters of the model loss value of the speech. The numerical values related to the model loss value of the speech include, but are not limited to, the length of the predicted text sequence of the speech and other numerical values required for calculating the calculation parameters of the model loss value of the speech.
[0048] Optionally, the available numerical range may be a preset numerical range. The available numerical range includes an upper limit value and a lower limit value. The numerical values within the available numerical range refer to the numerical values that are greater than the lower limit value of the available numerical range and less than the upper limit value of the available numerical range. When the numerical values related to the model loss value of each input speech are within the available numerical range, the numerical differences between the numerical values related to the model loss value of each input speech are relatively small. When the lengths of the predicted text sequences and the true text sequences of each speech are adjusted to the same length and the numerical differences between the numerical values related to the model loss value of each speech are relatively small, the cumulative error of the model loss value of the speech determined by processing the predicted text sequences, the true text sequences, and the numerical values related to the model loss value of each speech can be effectively avoided, and the accuracy of the determined model loss value of the speech can be improved.
[0049] Optionally, the numerical values related to the model loss value of each input voice can be detected to obtain the numerical values related to the model loss value of each input voice, and then the numerical values related to the model loss value of each input voice can be dynamically adjusted to adjust the numerical values related to the model loss value of each input voice within the available numerical value range.
[0050] Optionally, dynamically adjusting the numerical values related to the model loss value of each input voice to adjust the numerical values related to the model loss value of each input voice within the available numerical value range includes: performing the following operations on the numerical values related to the model loss value of each input voice: if the numerical value related to the model loss value of the input voice is less than the lower limit value of the available numerical value range, then amplify the numerical value related to the model loss value of the input voice so that the numerical value related to the model loss value of the input voice is greater than the lower limit value of the available numerical value range and less than the upper limit value of the available numerical value range. If the numerical value related to the model loss value of the input voice is less than the lower limit value of the available numerical value range, the numerical value related to the model loss value of the input voice can be amplified so that the numerical value related to the model loss value of the input voice is greater than the lower limit value of the available numerical value range and less than the upper limit value of the available numerical value range, thereby adjusting the numerical value related to the model loss value of the input voice within the available numerical value range.
[0051] Optionally, dynamically adjusting the numerical values related to the model loss value of each input voice to adjust the numerical values related to the model loss value of each input voice within the available numerical value range further includes: performing the following operations on the numerical values related to the model loss value of each input voice: if the numerical value related to the model loss value of the input voice is greater than the upper limit value of the available numerical value range, then reduce the numerical value related to the model loss value of the input voice so that the numerical value related to the model loss value of the input voice is greater than the lower limit value of the available numerical value range and less than the upper limit value of the available numerical value range. If the numerical value related to the model loss value of the input voice is greater than the upper limit value of the available numerical value range, the numerical value related to the model loss value of the input voice can be reduced so that the numerical value related to the model loss value of the input voice is greater than the lower limit value of the available numerical value range and less than the upper limit value of the available numerical value range, thereby adjusting the numerical value related to the model loss value of the input voice within the available numerical value range.
[0052] Optionally, if the numerical value related to the model loss value of the input voice is greater than the lower limit value of the available numerical value range and less than the upper limit value of the available numerical value range, it is determined that the numerical value related to the model loss value of the input voice is already within the available numerical value range and no dynamic adjustment is required.
[0053] Step 104: For each input voice, through each block calculation unit, calculate according to the preprocessed predicted text sequence, the preprocessed true text sequence, and the dynamically adjusted numerical values related to the model loss value of the input voice to obtain the model loss value calculation parameters of the input voice.
[0054] Optionally, each block calculation unit may be a software module or a hardware module that is pre-set to calculate specified parameters among the model loss value calculation parameters of the speech. Each block calculation unit may include: a block calculation unit for calculating the probability that the predicted text sequence of the speech is the true text sequence of the speech, a block calculation unit for calculating the forward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech, and a block calculation unit for calculating the backward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech.
[0055] Optionally, for each input speech, through each block calculation unit, calculations are performed according to the pre-processed predicted text sequence of the input speech, the pre-processed true text sequence, and the dynamically adjusted numerical values related to the model loss value, to obtain the model loss value calculation parameters of the input speech, including: performing the following operations for each input speech: inputting the pre-processed predicted text sequence of the input speech, the pre-processed true text sequence, and the dynamically adjusted numerical values related to the model loss value into each block calculation unit, and obtaining the model loss value calculation parameters of the input speech calculated by each block calculation unit according to the pre-processed predicted text sequence of the input speech, the pre-processed true text sequence, and the dynamically adjusted numerical values related to the model loss value.
[0056] Optionally, for each input speech, the pre-processed predicted text sequence of the input speech, the pre-processed true text sequence, and the dynamically adjusted numerical values related to the model loss value are input into the block calculation unit for calculating the probability that the predicted text sequence of the speech is the true text sequence of the speech. The block calculation unit for calculating the probability that the predicted text sequence of the speech is the true text sequence of the speech will perform calculations according to the pre-processed predicted text sequence of the input speech, the pre-processed true text sequence, and the dynamically adjusted numerical values related to the model loss value, to obtain the probability that the predicted text sequence of the input speech is the true text sequence of the input speech, and then output the probability that the predicted text sequence of the input speech is the true text sequence of the input speech. The probability that the predicted text sequence of the input speech is the true text sequence of the input speech output by the block calculation unit for calculating the probability that the predicted text sequence of the speech is the true text sequence of the speech can be obtained.
[0057] Optionally, for each input speech, the pre-processed predicted text sequence of the input speech, the pre-processed true text sequence, and the values related to the dynamically adjusted model loss value are input into a block calculation unit for calculating the forward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech. The block calculation unit for calculating the forward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech calculates based on the pre-processed predicted text sequence of the input speech, the pre-processed true text sequence, and the values related to the dynamically adjusted model loss value, to obtain the forward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech, and then outputs the forward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech. The forward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech output by the block calculation unit for calculating the forward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech can be obtained.
[0058] Optionally, for each input speech, the pre-processed predicted text sequence of the input speech, the pre-processed true text sequence, and the values related to the dynamically adjusted model loss value are input into a block calculation unit for calculating the backward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech. The block calculation unit for calculating the backward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech calculates based on the pre-processed predicted text sequence of the input speech, the pre-processed true text sequence, and the values related to the dynamically adjusted model loss value, to obtain the backward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech, and then outputs the backward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech. The backward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech output by the block calculation unit for calculating the backward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech can be obtained.
[0059] Step 105: Calculate parameters based on the model loss values of the respective input speeches to determine the model loss values of the respective input speeches.
[0060] Optionally, the model loss value calculation parameters of each input speech are used to determine the model loss value of each input speech, including: substituting the model loss value calculation parameters of each input speech into a preset loss value calculation formula to obtain the model loss value of each input speech. The preset loss value calculation formula can be a formula that is preset for calculating the model loss value of speech based on the model loss value calculation parameters of the speech.
[0061] Optionally, for each input speech, substitute the model loss value calculation parameters of the input speech into the following preset loss value calculation formula to obtain the model loss value of the input speech:
[0062]
[0063] Wherein, is the model loss value of the input speech, is the probability that the predicted text sequence of the input speech is the true text sequence of the input speech, and α t (l k ) is the forward recurrence probability that the k-th character in the predicted text sequence of the input speech is the k-th character in the true text sequence of the input speech, and β t (l k ) is the backward recurrence probability that the k-th character in the predicted text sequence of the input speech is the k-th character in the true text sequence of the input speech, k = 1, 2,..., N, and N is the total number of characters included in the predicted text sequence of the speech. α t (l k )β t (l k ) is the product of the forward recurrence probability and the backward recurrence probability corresponding to the k-th character in the predicted text sequence of the input speech, is the sum of the products of the forward recurrence probabilities and the backward recurrence probabilities corresponding to each character in the predicted text sequence of the input speech.
[0064] Optionally, after determining the model loss value of each input speech according to the model loss value calculation parameters of each input speech, it further includes: providing the model loss value of each input speech to the target user corresponding to the target prediction model. The target user corresponding to the target prediction model can be a technical person in charge of managing the target prediction model. Providing the model loss value of each input speech to the target user corresponding to the target prediction model includes: sending the model loss value of each input speech to the terminal device of the target user corresponding to the target prediction model. The terminal device of the target user can refer to the terminal device used by the target user.
[0065] In the technical solution of the embodiment of the present invention, each input voice in the voice dataset corresponding to the target prediction model is input into the target prediction model to obtain the predicted text sequence of each input voice; then, according to the lengths of the predicted text sequences and the true text sequences of each input voice, the predicted text sequences and the true text sequences of each input voice are preprocessed; the numerical values related to the model loss value of each input voice are dynamically adjusted to adjust the numerical values related to the model loss value of each input voice to the available numerical range; for each input voice, through each block calculation unit, according to the preprocessed predicted text sequence of the input voice, the preprocessed true text sequence, and the dynamically adjusted numerical values related to the model loss value, calculations are performed to obtain the model loss value calculation parameters of the input voice; according to the model loss value calculation parameters of each input voice, the model loss value of each input voice is determined, solving the problem that the determination scheme of the model loss value in the related technology directly processes the predicted text sequences, true text sequences, and related numerical values of each voice, and it is impossible to avoid the cumulative error of the model loss value of the determined voice, resulting in a low accuracy of the determined model loss value of the voice. It is possible to preprocess the predicted text sequences and true text sequences of each voice according to the lengths of the predicted text sequences and true text sequences of each voice, adjust the lengths of the predicted text sequences and true text sequences of each voice to the same length, and dynamically adjust the numerical values related to the model loss value of each voice to adjust the numerical values related to the model loss value of each voice to the available numerical range. Furthermore, based on each block calculation unit and the preprocessed predicted text sequence, preprocessed true text sequence, and dynamically adjusted numerical values related to the model loss value of each voice, the model loss value calculation parameters of each voice are determined, and then according to the model loss value calculation parameters of each voice, the model loss value of each voice is determined, thereby effectively avoiding the cumulative error of the model loss value of the voice determined by processing the predicted text sequences, true text sequences, and numerical values related to the model loss value of each voice, and improving the accuracy of the determined model loss value of the voice.
[0066] Embodiment 2
[0067] Figure 2 The flowchart of a method for determining a model loss value provided by Embodiment 2 of the present invention. The embodiments of the present invention can be combined with each optional solution in one or more of the above embodiments. As Figure 2 shown, the method includes:
[0068] Step 201: Input each input voice in the voice dataset corresponding to the target prediction model into the target prediction model to obtain the predicted text sequence of each input voice.
[0069] Step 202: Determine the maximum value among the lengths of the predicted text sequences and the true text sequences of each input voice.
[0070] Step 203: Perform text insertion on the text sequences with lengths less than the maximum value in the predicted text sequences and the true text sequences of each input voice, and adjust the lengths of the text sequences with lengths less than the maximum value to the maximum value.
[0071] Step 204: Dynamically adjust the values related to the model loss value of each input voice, and adjust the values related to the model loss value of each input voice to the available value range.
[0072] Step 205: For each input voice, through each block calculation unit, calculate according to the preprocessed predicted text sequence, the preprocessed true text sequence, and the dynamically adjusted values related to the model loss value of the input voice, to obtain the model loss value calculation parameters of the input voice.
[0073] Step 206: Substitute the model loss value calculation parameters of each input voice into the preset loss value calculation formula to obtain the model loss values of each input voice.
[0074] Step 207: Provide the model loss values of each input voice to the target user corresponding to the target prediction model.
[0075] The technical solution of the embodiment of the present invention can perform text insertion according to the lengths of the predicted text sequences and the true text sequences of each voice, adjust the lengths of the predicted text sequences and the true text sequences of each voice to the same length, can dynamically adjust the values related to the model loss value of each voice, and adjust the values related to the model loss value of each voice to the available value range. Furthermore, based on each block calculation unit and the preprocessed predicted text sequence, the preprocessed true text sequence, and the dynamically adjusted values related to the model loss value of each voice, determine the model loss value calculation parameters of each voice. Then, substitute the model loss value calculation parameters of each voice into the preset loss value calculation formula to obtain the model loss values of each voice, thereby effectively avoiding the cumulative error of the model loss value of the voice determined by processing the predicted text sequence, the true text sequence, and the values related to the model loss value of each voice, and improving the accuracy of the determined model loss value of the voice.
[0076] Embodiment III
[0077] Figure 3 It is a schematic structural diagram of a device for determining a model loss value provided by Embodiment III of the present invention. The device can be configured in an electronic device. As Figure 3As shown in the figure, the device includes: a sequence determination module 301, a sequence preprocessing module 302, a numerical value dynamic adjustment module 303, a parameter calculation module 304, and a loss value calculation module 305.
[0078] Among them, the sequence determination module 301 is configured to input each input voice in the voice dataset corresponding to the target prediction model into the target prediction model to obtain a predicted text sequence of each input voice; the sequence preprocessing module 302 is configured to preprocess the predicted text sequence and the true text sequence of each input voice according to the lengths of the predicted text sequence and the true text sequence of each input voice; the numerical value dynamic adjustment module 303 is configured to dynamically adjust the numerical values related to the model loss value of each input voice and adjust the numerical values related to the model loss value of each input voice to the available numerical value range; the parameter calculation module 304 is configured to calculate, for each input voice, through each block calculation unit, based on the preprocessed predicted text sequence of the input voice, the preprocessed true text sequence, and the numerically adjusted numerical values related to the model loss value, to obtain the model loss value calculation parameters of the input voice; the loss value calculation module 305 is configured to determine the model loss value of each input voice according to the model loss value calculation parameters of each input voice.
[0079] In the technical solution of the embodiment of the present invention, by inputting each input voice in the voice dataset corresponding to the target prediction model into the target prediction model, a predicted text sequence of each input voice is obtained; then, according to the lengths of the predicted text sequence and the true text sequence of each input voice, preprocessing is performed on the predicted text sequence and the true text sequence of each input voice; the values related to the model loss value of each input voice are dynamically adjusted, and the values related to the model loss value of each input voice are adjusted to the available value range; for each input voice, through each block calculation unit, calculations are performed according to the preprocessed predicted text sequence, the preprocessed true text sequence, and the dynamically adjusted values related to the model loss value of the input voice, to obtain the model loss value calculation parameters of the input voice; according to the model loss value calculation parameters of each input voice, the model loss value of each input voice is determined, solving the problem that the determination scheme of the model loss value in the related technology directly processes the predicted text sequence, the true text sequence, and the related values of each voice, and it is impossible to avoid the cumulative error of the determined model loss value of the voice, resulting in a low accuracy of the determined model loss value of the voice. It is possible to preprocess the predicted text sequence and the true text sequence of each voice according to the lengths of the predicted text sequence and the true text sequence of each voice, adjust the lengths of the predicted text sequence and the true text sequence of each voice to the same length, and dynamically adjust the values related to the model loss value of each voice, and adjust the values related to the model loss value of each voice to the available value range. Furthermore, based on each block calculation unit and the preprocessed predicted text sequence, the preprocessed true text sequence, and the dynamically adjusted values related to the model loss value of each voice, the model loss value calculation parameters of each voice are determined, and then according to the model loss value calculation parameters of each voice, the model loss value of each voice is determined, thereby effectively avoiding the cumulative error of the determined model loss value of the voice after processing the predicted text sequence, the true text sequence, and the values related to the model loss value of each voice, and improving the accuracy of the determined model loss value of the voice.
[0080] In an alternative implementation manner of the embodiment of the present invention, optionally, the sequence preprocessing module 302 is specifically configured to: determine the maximum value among the lengths of the predicted text sequence and the true text sequence of each input voice; perform text insertion on the text sequences with lengths less than the maximum value among the predicted text sequence and the true text sequence of each input voice, and adjust the lengths of the text sequences with lengths less than the maximum value to the maximum value.
[0081] In an alternative embodiment of the embodiment of the present invention, optionally, the numerical value dynamic adjustment module 303 is specifically configured to perform the following operations on the numerical value related to the model loss value of each input voice: if the numerical value related to the model loss value of the input voice is less than the lower limit value of the available numerical value range, then amplify the numerical value related to the model loss value of the input voice so that the numerical value related to the model loss value of the input voice is greater than the lower limit value of the available numerical value range and less than the upper limit value of the available numerical value range.
[0082] In an alternative embodiment of the embodiment of the present invention, optionally, the numerical value dynamic adjustment module 303 is further configured to perform the following operations on the numerical value related to the model loss value of each input voice: if the numerical value related to the model loss value of the input voice is greater than the upper limit value of the available numerical value range, then reduce the numerical value related to the model loss value of the input voice so that the numerical value related to the model loss value of the input voice is greater than the lower limit value of the available numerical value range and less than the upper limit value of the available numerical value range.
[0083] In an alternative embodiment of the embodiment of the present invention, optionally, the parameter calculation module 304 is specifically configured to perform the following operations on each input voice: input the preprocessed predicted text sequence, the preprocessed true text sequence, and the numerically value related to the dynamically adjusted model loss value of the input voice into each block calculation unit, and obtain the model loss value calculation parameters of the input voice calculated by each block calculation unit according to the preprocessed predicted text sequence, the preprocessed true text sequence, and the numerically value related to the dynamically adjusted model loss value of the input voice.
[0084] In an alternative embodiment of the embodiment of the present invention, optionally, the loss value calculation module 305 is specifically configured to substitute the model loss value calculation parameters of each input voice into a preset loss value calculation formula to obtain the model loss values of each input voice.
[0085] In an alternative embodiment of the embodiment of the present invention, optionally, it further includes: a loss value providing module, configured to provide the model loss values of each input voice to the target user corresponding to the target prediction model.
[0086] The device for determining the model loss value provided by the embodiment of the present invention can execute the method for determining the model loss value provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0087] Embodiment 4
[0088] Figure 4FIG. 0 shows a schematic structural diagram of an electronic device 10 that can be used to implement the method for determining the model loss value according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, electronic devices, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0089] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0090] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0091] The processor 11 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for determining the model loss value.
[0092] In some embodiments, the method for determining the model loss value may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed onto the heterogeneous hardware accelerator via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the processor, one or more steps of the method for determining the model loss value described above may be performed. Alternatively, in other embodiments, the processor may be configured to execute the method for determining the model loss value by any other suitable means (e.g., by means of firmware).
[0093] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0094] The computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or electronic device.
[0095] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0096] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a heterogeneous hardware accelerator that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the heterogeneous hardware accelerator. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0097] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data electronic device), or a computing system that includes middleware components (e.g., an application electronic device), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0098] A computing system may include a client and an electronic device. The client and the electronic device are generally far from each other and usually interact via a communication network. The relationship between the client and the electronic device is generated by computer programs running on respective computers and having a client-electronic device relationship with each other. The electronic device may be a cloud electronic device, also known as a cloud computing electronic device or a cloud host, which is a host product in a cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0099] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0100] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining a model loss value, characterized in that: include: Input each input speech in the speech data set corresponding to the target prediction model into the target prediction model to obtain a predicted text sequence of each input speech; Preprocessing the predicted text sequence and the real text sequence of each input speech according to the length of the predicted text sequence and the real text sequence of each input speech; Dynamically adjust the model loss value related values of each input speech, and adjust the model loss value related values of each input speech to within the available value range; For each input speech, each block calculation unit calculates according to the preprocessed predicted text sequence of the input speech, the preprocessed real text sequence and the dynamically adjusted model loss value related values to obtain the model loss value calculation parameters of the input speech; The model loss value of each input speech is determined based on the model loss value calculation parameters of each input speech.
2. The method for determining the model loss value according to claim 1, characterized in that: According to the lengths of the predicted text sequence and the real text sequence of each input speech, the predicted text sequence and the real text sequence of each input speech are preprocessed, including: Determine the maximum value of the length of the predicted text sequence and the true text sequence of each input speech; Text insertion is performed on the text sequences whose lengths are less than the maximum value in the predicted text sequences and the real text sequences of each input speech, and the lengths of the text sequences whose lengths are less than the maximum value are adjusted to the maximum value.
3. The method for determining the model loss value according to claim 1, characterized in that: Dynamically adjust the model loss value related values of each input speech to adjust the model loss value related values of each input speech to within the available value range, including: For each input speech, perform the following operations on the model loss value: If the model loss value related value of the input speech is smaller than the lower limit value of the available numerical range, the model loss value related value of the input speech is amplified so that the model loss value related value of the input speech is greater than the lower limit value of the available numerical range and smaller than the upper limit value of the available numerical range.
4. The method for determining the model loss value according to claim 3, characterized in that: Dynamically adjusting the model loss value related values of each input speech to adjust the model loss value related values of each input speech to within the available value range, and also including: For each input speech, perform the following operations on the model loss value: If the model loss value related value of the input speech is greater than the upper limit value of the available numerical range, the model loss value related value of the input speech is reduced so that the model loss value related value of the input speech is greater than the lower limit value of the available numerical range and less than the upper limit value of the available numerical range.
5. The method for determining the model loss value according to claim 1, characterized in that: For each input speech, each block calculation unit calculates according to the preprocessed predicted text sequence of the input speech, the preprocessed real text sequence and the dynamically adjusted model loss value related values to obtain the model loss value calculation parameters of the input speech, including: For each input voice, perform the following operations: The preprocessed predicted text sequence, the preprocessed real text sequence and the dynamically adjusted model loss value related values of the input speech are input into each block calculation unit, and the model loss value calculation parameters of the input speech calculated by each block calculation unit according to the preprocessed predicted text sequence, the preprocessed real text sequence and the dynamically adjusted model loss value related values of the input speech are obtained.
6. The method for determining the model loss value according to claim 1, characterized in that: Calculate the parameters according to the model loss value of each input speech, and determine the model loss value of each input speech. include: Substitute the model loss value calculation parameters of each input speech into the preset loss value calculation formula to obtain the model loss value of each input speech.
7. The method for determining the model loss value according to claim 1, characterized in that: After calculating the parameters according to the model loss values of the input speech and determining the model loss values of the input speech, the method further includes: The model loss value of each input speech is provided to the target user corresponding to the target prediction model.
8. A device for determining a model loss value, characterized in that: include: A sequence determination module, used to input each input speech in the speech data set corresponding to the target prediction model into the target prediction model to obtain a predicted text sequence of each input speech; A sequence preprocessing module, used for preprocessing the predicted text sequence and the real text sequence of each input speech according to the length of the predicted text sequence and the real text sequence of each input speech; A numerical dynamic adjustment module is used to dynamically adjust the numerical value related to the model loss value of each input speech, and adjust the numerical value related to the model loss value of each input speech to within the available numerical range; A parameter calculation module is used to calculate, for each input speech, through each block calculation unit, according to the preprocessed predicted text sequence of the input speech, the preprocessed real text sequence and the dynamically adjusted model loss value related values, to obtain the model loss value calculation parameters of the input speech; The loss value calculation module is used to determine the model loss value of each input speech according to the model loss value calculation parameters of each input speech.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein, the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for determining the model loss value according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for determining a model loss value according to any one of claims 1 to 7 when executed.