Model loss value determination method and device, equipment and medium
By replacing the logarithmic multiplication function as a logarithmic addition function in the loss value parameter calculation unit, and packaging the speech of similar lengths in the speech data set into batches, the problem of low efficiency in the calculation of model loss value in the prior art is solved, and a fast and efficient calculation of model loss value is achieved.
Patent Information
- Application Number
- CN202510274360.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-27
AI Technical Summary
The determination scheme for model loss value in the prior art directly inputs each speech to the deep learning model, resulting in low efficiency and the inability to quickly obtain the predicted text sequence of each speech, which in turn affects the calculation efficiency of the model loss value.
By replacing the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with the logarithmic addition function, and performing feature analysis on the speech data set, input speech of similar length is packaged into an input speech batch and input it batch by batch to the deep learning model for calculation.
The calculation efficiency of the deep learning model is improved, and the predicted text sequences of each speech are quickly obtained, thereby improving the efficiency and performance of the process of determining the model loss value.
Smart Images

Figure CN120048265A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a method, apparatus, device, and medium for determining a model loss value. Background Art
[0002] With the development of artificial intelligence technologies, more and more enterprises have begun to perform speech recognition through deep learning models. The deep learning model trained based on a large amount of data can recognize the input speech, convert the speech into a text sequence for representing the speech content of the speech, and output the text sequence obtained by converting the speech.
[0003] During the training and use of the deep learning model, it is necessary to evaluate the performance of the deep learning model according to a specified number of speeches and determine the model loss value of each speech. The predicted text sequence of a speech is the text sequence output by the deep learning model obtained by inputting the speech into the deep learning model for representing the speech content of the speech. The true text sequence of a speech is the correct text sequence for representing the speech content of the speech. The model loss value is a numerical value determined according to the predicted text sequence and the true text sequence of the speech for representing the difference between the predicted text sequence of the speech output by the deep learning model and the true text sequence of the speech. The smaller the model loss value, the smaller the difference between the predicted text sequence of the speech output by the deep learning model and the true text sequence of the speech, the better the performance of the deep learning model, and the more accurately it can convert the input speech into a text sequence for representing the speech content of the speech. The larger the model loss value, the larger the difference between the predicted text sequence of the speech output by the deep learning model and the true text sequence of the speech, the worse the performance of the deep learning model, and it cannot accurately convert the input speech into a text sequence for representing the speech content of the speech and needs to be further adjusted.
[0004] In the related art, the common solution for determining the model loss value is as follows: Input each voice into a deep learning model to obtain the predicted text sequence of each voice. Then, based on the predicted text sequence, true text sequence, and related values of each voice, determine the calculation parameters of the model loss value for each voice. Subsequently, based on the calculation parameters of the model loss value for each voice, determine the model loss value for each voice. The lengths of each voice are usually different. Directly inputting each voice into the deep learning model will result in the lengths of the voices input into the deep learning model being constantly changing. The constantly changing length of the input voice will increase the dynamicity during the calculation of the deep learning model, severely affecting the calculation performance of the deep learning model, reducing the efficiency, and making it impossible to quickly obtain the predicted text sequences of each voice. The solution for determining the model loss value in the related art directly inputs each voice into the deep learning model to obtain the predicted text sequence of each voice, with low efficiency and unable to quickly obtain the predicted text sequences of each voice. Subsequently, based on the predicted text sequence, true text sequence, and related values of each voice, determine the calculation parameters of the model loss value and the model loss value for each voice. Summary of the Invention
[0005] The present invention provides a method, apparatus, device, and medium for determining a model loss value to solve the problem that the solution for determining the model loss value in the related art directly inputs each voice into the deep learning model to obtain the predicted text sequence of each voice, with low efficiency and unable to quickly obtain the predicted text sequences of each voice. Subsequently, based on the predicted text sequence, true text sequence, and related values of each voice, determine the calculation parameters of the model loss value and the model loss value for each voice.
[0006] According to one aspect of the present invention, there is provided a method for determining a model loss value, including:
[0007] Replace the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with a logarithmic addition function;
[0008] Perform feature analysis on the voice data set corresponding to the target prediction model, and pack the input voices with similar lengths into one input voice batch;
[0009] Input the input voices in each input voice batch into the target prediction model to obtain the predicted text sequences of the input voices in each input voice batch;
[0010] Preprocess the predicted text sequences and true text sequences of the input voices in each input voice batch according to the lengths of the predicted text sequences and true text sequences of the input voices in each input voice batch;
[0011] Through the loss value parameter calculation unit, calculate according to the preprocessed predicted text sequence, preprocessed true text sequence and model loss value related values of the input speech in each input speech batch, and obtain the model loss value calculation parameters of the input speech in each input speech batch;
[0012] Determine the model loss value of the input speech in each input speech batch according to the model loss value calculation parameters of the input speech in each input speech batch.
[0013] According to another aspect of the present invention, there is provided a device for determining a model loss value, including:
[0014] A function replacement module, configured to replace the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with a logarithmic addition function;
[0015] A dataset analysis module, configured to perform feature analysis on the speech dataset corresponding to the target prediction model, and pack input speeches of similar lengths into one input speech batch;
[0016] A sequence determination module, configured to input the input speech in each input speech batch into the target prediction model to obtain the predicted text sequence of the input speech in each input speech batch;
[0017] A sequence preprocessing module, configured to preprocess the predicted text sequence and true text sequence of the input speech in each input speech batch according to the lengths of the predicted text sequence and true text sequence of the input speech in each input speech batch;
[0018] A parameter calculation module, configured to calculate through the loss value parameter calculation unit according to the preprocessed predicted text sequence, preprocessed true text sequence and model loss value related values of the input speech in each input speech batch, and obtain the model loss value calculation parameters of the input speech in each input speech batch;
[0019] A loss value determination module, configured to determine the model loss value of the input speech in each input speech batch according to the model loss value calculation parameters of the input speech in each input speech batch.
[0020] According to another aspect of the present invention, there is provided an electronic device, and the electronic device includes:
[0021] At least one processor;
[0022] And a memory communicatively connected to the at least one processor;
[0023] Wherein, the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for determining the model loss value according to any embodiment of the present invention.
[0024] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for implementing the method for determining the model loss value according to any embodiment of the present invention when executed by a processor.
[0025] In the technical solution of the embodiment of the present invention, the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model is replaced with a logarithmic addition function; then, the speech data set corresponding to the target prediction model is subjected to feature analysis, and the input speeches with similar lengths are packed into an input speech batch; the input speeches in each input speech batch are input into the target prediction model to obtain the predicted text sequences of the input speeches in each input speech batch; according to the lengths of the predicted text sequences and the true text sequences of the input speeches in each input speech batch, the predicted text sequences and the true text sequences of the input speeches in each input speech batch are preprocessed; through the loss value parameter calculation unit, according to the preprocessed predicted text sequences, the preprocessed true text sequences and the model loss value related values of the input speeches in each input speech batch, the model loss value calculation parameters of the input speeches in each input speech batch are calculated; finally, according to the model loss value calculation parameters of the input speeches in each input speech batch, the model loss values of the input speeches in each input speech batch are determined, solving the problem that the determination scheme of the model loss value in the related technology directly inputs each speech into the deep learning model to obtain the predicted text sequences of each speech, with low efficiency and unable to quickly obtain the predicted text sequences of each speech, and then determining the model loss value calculation parameters and the model loss value of each speech according to the predicted text sequences, the true text sequences and the related values of each speech. The logarithmic multiplication function in the loss value parameter calculation unit can be replaced with a logarithmic addition function, so as to improve the calculation speed of the loss value parameter calculation unit without affecting the calculation function of the loss value parameter calculation unit. The speeches with similar lengths in the speech data set can be packed into multiple batches, and the lengths of the speeches in each batch are close. Then, the speeches in each batch are sequentially input into the deep learning model to be evaluated, which can reduce the amplitude of the length change of the speeches input into the deep learning model and avoid the increase in the dynamicity when the deep learning model performs calculations due to the continuous change of the length of the input speeches, thereby improving the efficiency of the deep learning model, quickly obtaining the predicted text sequences of each speech, and then determining the model loss value calculation parameters and the model loss value of each speech according to the predicted text sequences, the true text sequences and the related values of each speech, and improving the efficiency and performance of the determination process of the model loss value.
[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0028] Figure 1 It is a flowchart of a method for determining the model loss value provided in the first embodiment of the present invention.
[0029] Figure 2 It is a flowchart of a method for determining the model loss value provided in the second embodiment of the present invention.
[0030] Figure 3 It is a schematic structural diagram of a device for determining the model loss value provided in the third embodiment of the present invention.
[0031] Figure 4 It is a schematic structural diagram of an electronic device for implementing the method for determining the model loss value in the embodiments of the present invention. Detailed implementation manners
[0032] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0033] It should be noted that the terms "target", "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising", "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0034] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data comply with the relevant laws, regulations, and standards of the relevant regions.
[0035] Embodiment 1
[0036] Figure 1 The figure is a flowchart of a method for determining a model loss value provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of evaluating the performance of a deep learning model according to a specified number of voices during the training and use of a deep learning model for speech recognition, and determining the model loss value of each voice. This method can be executed by a device for determining a model loss value, and the device for determining a model loss value can be implemented in the form of hardware and / or software, and the device for determining a model loss value can be configured in an electronic device. The electronic device can be an artificial intelligence (AI) chip for training and using a deep learning model. As Figure 1 shown, the method includes:
[0037] Step 101, replace the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with a logarithmic addition function.
[0038] Optionally, the target prediction model is a deep learning model for speech recognition that needs to be evaluated for performance. The target prediction model is used to recognize the input voice, convert the voice into a text sequence for characterizing the voice content of the voice, and output the text sequence obtained by converting the voice. The text sequence can be a sequence composed of one or more characters. The predicted text sequence of the voice is the text sequence output by the target prediction model obtained by inputting the voice into the target prediction model for characterizing the voice content of the voice. The input of the target prediction model is the voice, and the output of the target prediction model is the predicted text sequence of the voice. The true text sequence of the voice is the correct text sequence for characterizing the voice content of the voice.
[0039] Optionally, the embodiments of the present invention use the Connectionist Temporal Classification (CTC) algorithm to determine the model loss value of speech. The model loss value of each speech is calculated according to the calculation parameters of the model loss value of the speech. The calculation parameters of the model loss value of the speech may include: the probability that the predicted text sequence of the speech is the true text sequence of the speech, the forward recurrence probability and the backward recurrence probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech. Wherein, k = 1, 2,..., N. N is the total number of characters included in the predicted text sequence of the speech. The forward recurrence probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech may refer to the probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech calculated using the forward probability algorithm. The backward recurrence probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech may refer to the probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech calculated using the backward probability algorithm. The forward recurrence probability and the backward recurrence probability that the k-th character in the predicted text sequence of the speech is the k-th character in the true text sequence of the speech are the forward recurrence probability and the backward recurrence probability that each character in the predicted text sequence of the speech is the character at the same position in the true text sequence of the speech. The numerical values related to the model loss value of the speech are the numerical values required for calculating the calculation parameters of the model loss value of the speech. The numerical values related to the model loss value of the speech include, but are not limited to, the length of the predicted text sequence of the speech and other numerical values required for calculating the calculation parameters of the model loss value of the speech.
[0040] Optionally, the loss value parameter calculation unit corresponding to the target prediction model may be a software module preset for calculating the calculation parameters of the model loss value of the speech. A logarithmic multiplication function is set in the loss value parameter calculation unit. The logarithmic multiplication function may be a function for calculating the result of multiplying two logarithms. The logarithmic multiplication function can be expressed as log a M * N. The logarithmic addition function may be a function for calculating the result of adding two logarithms. The logarithmic addition function can be expressed as log a M + log a N. The logarithmic multiplication function and the logarithmic addition function can be equivalently replaced mathematically. The calculation speed is slower when using the logarithmic multiplication function for calculation. The calculation speed is faster when using the logarithmic addition function for calculation. The logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model can be replaced with the logarithmic addition function, so as to improve the calculation speed of the loss value parameter calculation unit without affecting the calculation function of the loss value parameter calculation unit.
[0041] Step 102: Perform feature analysis on the speech dataset corresponding to the target prediction model, and pack input speeches of similar lengths into an input speech batch.
[0042] Optionally, the speech dataset corresponding to the target prediction model is composed of multiple input speeches. Each input speech is a pre-collected speech for evaluating the performance of the target prediction model. Each input speech in the speech dataset is a speech authorized by the user or fully authorized by all parties, and the collection, use, and processing of the input speech comply with the relevant laws, regulations, and standards in the relevant region. The lengths of the individual input speeches are usually different. The length of a speech can refer to the duration of the speech. The predicted text sequence of an input speech is the text sequence output by the target prediction model after inputting the input speech, which is used to represent the speech content of the speech. The true text sequence of an input speech is the correct text sequence used to represent the speech content of the input speech. The speech dataset corresponding to the target prediction model and the true text sequences of the individual input speeches are stored in the electronic device. The speech dataset corresponding to the target prediction model and the true text sequences of the individual input speeches can be stored in the electronic device by a technician.
[0043] Optionally, performing feature analysis on the speech dataset corresponding to the target prediction model and packing input speeches of similar lengths into an input speech batch includes: using a preset clustering algorithm to cluster the individual input speeches in the speech dataset corresponding to the target prediction model according to the lengths of the individual input speeches, to obtain at least two clustering results; wherein each clustering result is composed of input speeches of similar lengths; and determining each clustering result as an input speech batch.
[0044] Optionally, the preset clustering algorithm can be a pre-set clustering algorithm for clustering multiple speeches according to the lengths of the multiple speeches and clustering speeches of similar lengths into one category. Speeches of similar lengths can refer to multiple speeches with close lengths. Exemplarily, the length difference between multiple speeches with close lengths is less than 10 seconds, and the length difference between multiple speeches with non-close lengths is greater than or equal to 10 seconds. The preset clustering algorithm can be used to cluster multiple speeches according to the lengths of the multiple speeches to obtain at least two clustering results. Each clustering result is a group of multiple speeches clustered into one category. Each clustering result is composed of speeches of similar lengths. The lengths of the individual speeches in each clustering result are close.
[0045] Optionally, the input speech batch is a group of speeches of similar lengths in the speech dataset corresponding to the target prediction model. A preset clustering algorithm can be used to cluster each input speech in the speech dataset corresponding to the target prediction model according to the length of each input speech, resulting in at least two clustering results. Each clustering result is a group of multiple input speeches clustered into one category. Each clustering result is composed of input speeches of similar lengths. The lengths of the input speeches in each clustering result are close. Each clustering result is determined as an input speech batch, obtaining at least two input speech batches.
[0046] Thus, the input speeches of similar lengths in the speech dataset corresponding to the target prediction model are packed into an input speech batch, obtaining multiple input speech batches. Each input speech batch is a group of speeches of similar lengths in the speech dataset corresponding to the target prediction model.
[0047] Step 103: Input the input speeches in each input speech batch into the target prediction model to obtain the predicted text sequences of the input speeches in each input speech batch.
[0048] Optionally, a unique corresponding digital number can be set for each input speech batch, and then each input speech batch is sorted in ascending order according to the digital number to form a sequence. Then, starting from the input speech batch ranked first to the input speech batch ranked last, for each input speech batch in turn, the input speeches in the input speech batch are input into the target prediction model to obtain the predicted text sequences of the input speeches in the input speech batch.
[0049] Optionally, for each input speech in the input speech batch, the input speech can be input into the target prediction model. The target prediction model will recognize the input speech and convert the input speech into a text sequence representing the speech content of the input speech, that is, the predicted text sequence of the input speech, and then output the predicted text sequence of the input speech. The predicted text sequence of the input speech output by the target prediction model can be obtained, thereby obtaining the predicted text sequence of the input speech. Thus, the input speeches in each input speech batch are respectively input into the target prediction model to obtain the predicted text sequences of each input speech.
[0050] The lengths of the input speeches in each input speech batch are close. By sequentially inputting the input speeches in each input speech batch into the target prediction model, the variation range of the lengths of the speeches input into the target prediction model can be reduced, avoiding the increase in the dynamic nature during the calculation of the target prediction model due to the continuously changing lengths of the input speeches, thereby improving the efficiency of the deep learning model and quickly obtaining the predicted text sequences of each speech.
[0051] Step 104: Preprocess the predicted text sequences and the true text sequences of the input voices in each input voice batch according to the lengths of the predicted text sequences and the true text sequences of the input voices in each input voice batch.
[0052] Optionally, preprocessing the predicted text sequences and the true text sequences of the input voices in each input voice batch according to the lengths of the predicted text sequences and the true text sequences of the input voices in each input voice batch includes: performing the following operations for each input voice batch: determining the maximum value among the lengths of the predicted text sequences and the true text sequences of the input voices in the input voice batch; performing text insertion on the sequences in the predicted text sequences and the true text sequences of the input voices in the input voice batch whose lengths are less than the maximum value, and adjusting the lengths of the sequences whose lengths are less than the maximum value to the maximum value.
[0053] Optionally, the length of the text sequence may refer to the total number of characters included in the text sequence.
[0054] Optionally, starting from the first input voice batch to the last input voice batch, for each input voice batch in turn, the lengths of the predicted text sequences of the input voices in the input voice batch and the lengths of the true text sequences of the input voices can be detected, and then the maximum value among the lengths can be counted, so as to determine the maximum value among the lengths of the predicted text sequences and the true text sequences of the input voices in the input voice batch. Performing text insertion on the sequences in the predicted text sequences and the true text sequences of the input voices in the input voice batch whose lengths are less than the maximum value and adjusting the lengths of the sequences whose lengths are less than the maximum value to the maximum value includes: performing the following operations for each text sequence in the predicted text sequences and the true text sequences of the input voices whose length is less than the maximum value: inserting one or more characters '0' at the end of the text sequence so that the length of the text sequence increases to the maximum value. Thus, the predicted text sequences and the true text sequences of the input voices in each input voice batch can be preprocessed, and one or more characters '0' are inserted at the end of the text sequences with smaller lengths so that the lengths of the text sequences increase to the maximum value of the lengths of each text sequence, thereby adjusting the lengths of the predicted text sequences and the true text sequences of the input voices in each input voice batch to the same length without affecting the speech content represented by each text sequence.
[0055] Step 105: Through the loss value parameter calculation unit, calculate according to the preprocessed predicted text sequence, the preprocessed true text sequence, and the model loss value related values of the input speech in each input speech batch, to obtain the model loss value calculation parameters of the input speech in each input speech batch.
[0056] Optionally, the model loss value related values of the input speech are the values required for calculating the model loss value calculation parameters of the input speech. The model loss value related values of the input speech include, but are not limited to, the length of the predicted text sequence of the input speech and other values required for calculating the model loss value calculation parameters of the input speech. The model loss value related values of the input speech in each input speech batch can be detected to obtain the model loss value related values of the input speech in each input speech batch.
[0057] Optionally, through the loss value parameter calculation unit, calculate according to the preprocessed predicted text sequence, the preprocessed true text sequence, and the model loss value related values of the input speech in each input speech batch, to obtain the model loss value calculation parameters of the input speech in each input speech batch, including: for each input speech batch, perform the following operations: respectively input the preprocessed predicted text sequence, the preprocessed true text sequence, and the model loss value related values of each input speech in the input speech batch into the loss value parameter calculation unit, and obtain the model loss value calculation parameters of each input speech calculated by the loss value parameter calculation unit according to the preprocessed predicted text sequence, the preprocessed true text sequence, and the model loss value related values of each input speech.
[0058] Optionally, the model loss value calculation parameters of the input speech may include: the probability that the predicted text sequence of the input speech is the true text sequence of the input speech, the forward recurrence probability and the backward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech. Starting from the first input speech batch to the last input speech batch, for each input speech batch in turn, the preprocessed predicted text sequence, the preprocessed true text sequence, and the numerical values related to the model loss value of each input speech in the input speech batch are respectively input into the loss value parameter calculation unit. The loss value parameter calculation unit calculates based on the preprocessed predicted text sequence, the preprocessed true text sequence, and the numerical values related to the model loss value of the input speech, to obtain the probability that the predicted text sequence of the input speech is the true text sequence of the input speech, the forward recurrence probability and the backward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech, and then outputs the probability that the predicted text sequence of the input speech is the true text sequence of the input speech, the forward recurrence probability and the backward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech. The probability that the predicted text sequence of the input speech is the true text sequence of the input speech, the forward recurrence probability and the backward recurrence probability that each character in the predicted text sequence of the input speech is the character at the same position in the true text sequence of the input speech output by the loss value parameter calculation unit can be obtained.
[0059] Step 106: Determine the model loss values of the input speeches in each input speech batch according to the model loss value calculation parameters of the input speeches in each input speech batch.
[0060] Optionally, determining the model loss values of the input speeches in each input speech batch according to the model loss value calculation parameters of the input speeches in each input speech batch includes: performing the following operations for each input speech batch: respectively substituting the model loss value calculation parameters of each input speech in the input speech batch into a preset loss value calculation formula to obtain the model loss values of each input speech. The preset loss value calculation formula may be a formula preset for calculating the model loss value of speech according to the model loss value calculation parameters of speech. Starting from the first input speech batch to the last input speech batch, for each input speech batch in turn, the model loss value calculation parameters of each input speech in the input speech batch are respectively substituted into the preset loss value calculation formula to obtain the model loss values of each input speech.
[0061] Optionally, for each input speech, substitute the model loss value calculation parameters of the input speech into the following preset loss value calculation formula to obtain the model loss value of the input speech:
[0062]
[0063] Wherein, is the model loss value of the input speech, is the probability that the predicted text sequence of the input speech is the true text sequence of the input speech, and α t (l k ) is the forward recurrence probability that the k-th character in the predicted text sequence of the input speech is the k-th character in the true text sequence of the input speech, and β t (l k ) is the backward recurrence probability that the k-th character in the predicted text sequence of the input speech is the k-th character in the true text sequence of the input speech, k = 1, 2,..., N, and N is the total number of characters included in the predicted text sequence of the speech. α t (l k )β t (l k ) is the product of the forward recurrence probability and the backward recurrence probability corresponding to the k-th character in the predicted text sequence of the input speech, is the sum of the products of the forward recurrence probabilities and the backward recurrence probabilities corresponding to each character in the predicted text sequence of the input speech.
[0064] Optionally, it further includes: quantizing the predicted text sequence and the true text sequence of the input speech in each input speech batch according to the service scenario of the speech dataset.
[0065] Optionally, quantizing the predicted text sequence and the true text sequence of the input speech in each input speech batch may refer to compressing the data types of the predicted text sequence and the true text sequence of the input speech in each input speech batch from a data type with higher precision to a data type with lower precision. The data type with higher precision may refer to 32-bit floating-point number (fp32) or 64-bit floating-point number (fp64). The data type with lower precision may refer to 8-bit floating-point number (fp8) or 16-bit floating-point number (fp16).
[0066] Optionally, the business scenario of the speech dataset corresponding to the target prediction model may be information used to characterize the precision requirement for the model loss value of the input speech in the speech dataset. The business scenario is normal or low precision. When the business scenario of the speech dataset corresponding to the target prediction model is normal, it indicates that the precision requirement for the model loss value of the input speech in the speech dataset is the general standard, and there is no need to perform quantization processing on the predicted text sequence and the true text sequence of the input speech in the speech dataset. When the business scenario of the speech dataset corresponding to the target prediction model is low precision, it indicates that the precision requirement for the model loss value of the input speech in the speech dataset is relatively low, and the processing efficiency of the predicted text sequence and the true text sequence of the input speech can be improved by performing quantization processing on the predicted text sequence and the true text sequence of the input speech in the speech dataset.
[0067] Optionally, according to the business scenario of the speech dataset, performing quantization processing on the predicted text sequence and the true text sequence of the input speech in each input speech batch includes: if the business scenario of the speech dataset is low precision, performing quantization processing on the predicted text sequence and the true text sequence of the input speech in each input speech batch; if the business scenario of the speech dataset is normal, determining that there is no need to perform quantization processing on the predicted text sequence and the true text sequence of the input speech in each input speech batch.
[0068] Optionally, after calculating the parameters based on the model loss value of the input speech in each input speech batch and determining the model loss value of the input speech in each input speech batch, it further includes: providing the model loss value of the input speech in each input speech batch to the target user corresponding to the target prediction model. The target user corresponding to the target prediction model may be a technical person in charge of managing the target prediction model. Providing the model loss value of the input speech in each input speech batch to the target user corresponding to the target prediction model includes: sending the model loss value of the input speech in each input speech batch to the terminal device of the target user corresponding to the target prediction model. The terminal device of the target user may refer to the terminal device used by the target user.
[0069] The technical solution of the embodiment of the present invention replaces the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with a logarithmic addition function; then performs feature analysis on the speech data set corresponding to the target prediction model, and packs the input speeches with similar lengths into an input speech batch; inputs the input speeches in each input speech batch into the target prediction model to obtain the predicted text sequences of the input speeches in each input speech batch; preprocesses the predicted text sequences and the true text sequences of the input speeches in each input speech batch according to the lengths of the predicted text sequences and the true text sequences of the input speeches in each input speech batch; through the loss value parameter calculation unit, calculates according to the preprocessed predicted text sequences, the preprocessed true text sequences and the model loss value related values of the input speeches in each input speech batch to obtain the model loss value calculation parameters of the input speeches in each input speech batch; finally, determines the model loss values of the input speeches in each input speech batch according to the model loss value calculation parameters of the input speeches in each input speech batch, solving the problem that the determination scheme of the model loss value in the related technology directly inputs each speech into the deep learning model to obtain the predicted text sequences of each speech, with low efficiency and unable to quickly obtain the predicted text sequences of each speech, and then determining the model loss value calculation parameters and the model loss value of each speech according to the predicted text sequences, the true text sequences and the related values of each speech. The logarithmic multiplication function in the loss value parameter calculation unit can be replaced with a logarithmic addition function, so as to improve the calculation speed of the loss value parameter calculation unit without affecting the calculation function of the loss value parameter calculation unit. The speeches with similar lengths in the speech data set can be packed into multiple batches, and the lengths of the speeches in each batch are close. Then, the speeches in each batch are input into the deep learning model to be evaluated in turn, which can reduce the length change range of the speeches input into the deep learning model and avoid the increase in the dynamicity of the deep learning model during calculation due to the continuous change of the length of the input speeches, thereby improving the efficiency of the deep learning model, quickly obtaining the predicted text sequences of each speech, and then determining the model loss value calculation parameters and the model loss value of each speech according to the predicted text sequences, the true text sequences and the related values of each speech, improving the efficiency and performance of the determination process of the model loss value.
[0070] Embodiment 2
[0071] Figure 2 It is a flowchart of a method for determining a model loss value provided by Embodiment 2 of the present invention. The embodiment of the present invention can be combined with each optional solution in one or more of the above embodiments. As Figure 2 shown, the method includes:
[0072] Step 201: Replace the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with a logarithmic addition function.
[0073] Step 202: Through a preset clustering algorithm, cluster each input speech in the speech dataset corresponding to the target prediction model according to the length of each input speech, to obtain at least two clustering results.
[0074] Wherein, each clustering result is composed of input speeches with similar lengths.
[0075] Step 203: Determine each clustering result as an input speech batch.
[0076] Step 204: Input the input speeches in each input speech batch into the target prediction model, to obtain the predicted text sequences of the input speeches in each input speech batch.
[0077] Step 205: According to the lengths of the predicted text sequences and the true text sequences of the input speeches in each input speech batch, preprocess the predicted text sequences and the true text sequences of the input speeches in each input speech batch.
[0078] Step 206: Through the loss value parameter calculation unit, calculate according to the preprocessed predicted text sequences, the preprocessed true text sequences, and the model loss value related values of the input speeches in each input speech batch, to obtain the model loss value calculation parameters of the input speeches in each input speech batch.
[0079] Step 207: Determine the model loss values of the input speeches in each input speech batch according to the model loss value calculation parameters of the input speeches in each input speech batch.
[0080] The technical solution of the embodiment of the present invention can, based on the clustering algorithm, pack the speeches with similar lengths in the speech dataset into multiple batches, the lengths of the speeches in each batch are close, and then input the speeches in each batch into the deep learning model to be evaluated in turn, which can reduce the amplitude of the length change of the speeches input into the deep learning model, avoid the increase in the dynamicity during the calculation of the deep learning model due to the continuous change of the length of the input speeches, thereby improving the efficiency of the deep learning model, quickly obtaining the predicted text sequences of each speech, and then determining the model loss value calculation parameters and the model loss values of each speech according to the predicted text sequences, the true text sequences, and the related values of each speech, improving the efficiency and performance of the determination process of the model loss value.
[0081] Embodiment III
[0082] Figure 3The figure is a schematic structural diagram of an apparatus for determining a model loss value provided in Embodiment 3 of the present invention. The apparatus may be configured in an electronic device. As Figure 3 shown, the apparatus includes: a function replacement module 301, a data set analysis module 302, a sequence determination module 303, a sequence preprocessing module 304, a parameter calculation module 305, and a loss value determination module 306.
[0083] Among them, the function replacement module 301 is configured to replace the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with a logarithmic addition function; the data set analysis module 302 is configured to perform feature analysis on the speech data set corresponding to the target prediction model, and pack the input speeches with similar lengths into an input speech batch; the sequence determination module 303 is configured to input the input speeches in each input speech batch into the target prediction model to obtain the predicted text sequences of the input speeches in each input speech batch; the sequence preprocessing module 304 is configured to preprocess the predicted text sequences and the true text sequences of the input speeches in each input speech batch according to the lengths of the predicted text sequences and the true text sequences of the input speeches in each input speech batch; the parameter calculation module 305 is configured to calculate, through the loss value parameter calculation unit, the model loss value calculation parameters of the input speeches in each input speech batch according to the preprocessed predicted text sequences, the preprocessed true text sequences, and the values related to the model loss value of the input speeches in each input speech batch; the loss value determination module 306 is configured to determine the model loss values of the input speeches in each input speech batch according to the model loss value calculation parameters of the input speeches in each input speech batch.
[0084] In the technical solution of the embodiment of the present invention, by replacing the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with a logarithmic addition function; then performing feature analysis on the speech data set corresponding to the target prediction model, and packing the input speeches with similar lengths into an input speech batch; inputting the input speeches in each input speech batch into the target prediction model to obtain the predicted text sequences of the input speeches in each input speech batch; preprocessing the predicted text sequences and the true text sequences of the input speeches in each input speech batch according to the lengths of the predicted text sequences and the true text sequences of the input speeches in each input speech batch; calculating, by the loss value parameter calculation unit, the model loss value calculation parameters of the input speeches in each input speech batch according to the preprocessed predicted text sequences, the preprocessed true text sequences and the values related to the model loss value of the input speeches in each input speech batch; and finally determining the model loss values of the input speeches in each input speech batch according to the model loss value calculation parameters of the input speeches in each input speech batch, the problem in the related art that the determination scheme of the model loss value directly inputs each speech into the deep learning model to obtain the predicted text sequences of each speech, with low efficiency and unable to quickly obtain the predicted text sequences of each speech, and then determining the model loss value calculation parameters and the model loss value of each speech according to the predicted text sequences, the true text sequences and the related values of each speech, is solved. The logarithmic multiplication function in the loss value parameter calculation unit can be replaced with a logarithmic addition function, so as to improve the calculation speed of the loss value parameter calculation unit without affecting the calculation function of the loss value parameter calculation unit. The speeches with similar lengths in the speech data set can be packed into multiple batches, and the lengths of the speeches in each batch are close. Then, the speeches in each batch are sequentially input into the deep learning model to be evaluated, which can reduce the amplitude of the length change of the speeches input into the deep learning model and avoid the increase in the dynamicity during the calculation of the deep learning model due to the continuous change of the lengths of the input speeches, thereby improving the efficiency of the deep learning model, quickly obtaining the predicted text sequences of each speech, and then determining the model loss value calculation parameters and the model loss value of each speech according to the predicted text sequences, the true text sequences and the related values of each speech, and improving the efficiency and performance of the determination process of the model loss value.
[0085] In an alternative embodiment of the embodiment of the present invention, optionally, the data set analysis module 302 is specifically configured to: cluster each input speech in the speech data set corresponding to the target prediction model according to the length of each input speech in the speech data set through a preset clustering algorithm to obtain at least two clustering results; wherein each clustering result is composed of input speeches with similar lengths; and determine each clustering result as an input speech batch.
[0086] In an alternative embodiment of the embodiment of the present invention, optionally, the sequence preprocessing module 304 is specifically configured to perform the following operations for each input speech batch: determine the maximum value among the lengths of the predicted text sequences and the true text sequences of the respective input speeches in the input speech batch; perform text insertion on the predicted text sequences and the true text sequences of the respective input speeches in the input speech batch whose lengths are less than the maximum value, and adjust the lengths of the sequences with lengths less than the maximum value to the maximum value.
[0087] In an alternative embodiment of the embodiment of the present invention, optionally, the parameter calculation module 305 is specifically configured to perform the following operations for each input speech batch: respectively input the preprocessed predicted text sequences, the preprocessed true text sequences, and the values related to the model loss value of the respective input speeches in the input speech batch into the loss value parameter calculation unit, and obtain the model loss value calculation parameters of the respective input speeches calculated by the loss value parameter calculation unit according to the preprocessed predicted text sequences, the preprocessed true text sequences, and the values related to the model loss value of the respective input speeches.
[0088] In an alternative embodiment of the embodiment of the present invention, optionally, the loss value determination module 306 is specifically configured to perform the following operations for each input speech batch: respectively substitute the model loss value calculation parameters of the respective input speeches in the input speech batch into a preset loss value calculation formula to obtain the model loss values of the respective input speeches.
[0089] In an alternative embodiment of the embodiment of the present invention, optionally, it further includes: a quantization module, configured to perform quantization processing on the predicted text sequences and the true text sequences of the input speeches in each input speech batch according to the service scenario of the speech data set.
[0090] In an alternative embodiment of the embodiment of the present invention, optionally, it further includes: a loss value providing module, configured to provide the model loss values of the input speeches in each input speech batch to the target user corresponding to the target prediction model.
[0091] The device for determining the model loss value provided by the embodiment of the present invention can execute the method for determining the model loss value provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0092] Embodiment 4
[0093] Figure 4FIG. 0 shows a schematic structural diagram of an electronic device 10 that can be used to implement the method for determining the model loss value according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, electronic devices, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0094] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0095] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0096] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for determining the model loss value.
[0097] In some embodiments, the method for determining the model loss value may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed onto the heterogeneous hardware accelerator via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the processor, one or more steps of the method for determining the model loss value described above may be performed. Alternatively, in other embodiments, the processor may be configured to execute the method for determining the model loss value by any other suitable means (e.g., by means of firmware).
[0098] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0099] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or electronic device.
[0100] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0101] To provide for interaction with a user, the systems and techniques described herein can be implemented on a heterogeneous hardware accelerator that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the heterogeneous hardware accelerator. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0102] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data electronic device), or a computing system that includes middleware components (e.g., an application electronic device), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0103] A computing system may include a client and an electronic device. The client and the electronic device are generally far from each other and usually interact via a communication network. The relationship between the client and the electronic device is generated by computer programs running on respective computers and having a client-electronic device relationship with each other. The electronic device may be a cloud electronic device, also known as a cloud computing electronic device or a cloud host, which is a host product in a cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0104] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0105] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining a model loss value, characterized in that: include: Replace the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with a logarithmic addition function; Performing feature analysis on a speech data set corresponding to the target prediction model, and packaging input speech of similar length into an input speech batch; Inputting the input speech in each input speech batch into the target prediction model to obtain a predicted text sequence of the input speech in each input speech batch; Preprocessing the predicted text sequence and the real text sequence of the input speech in each input speech batch according to the length of the predicted text sequence and the real text sequence of the input speech in each input speech batch; The loss value parameter calculation unit calculates the model loss value calculation parameters of the input speech in each input speech batch according to the preprocessed predicted text sequence, the preprocessed real text sequence and the model loss value related values of the input speech in each input speech batch; The model loss value of the input speech in each input speech batch is determined according to the model loss value calculation parameter of the input speech in each input speech batch.
2. The method for determining the model loss value according to claim 1, characterized in that: Perform feature analysis on the speech data set corresponding to the target prediction model, and pack input speech of similar length into an input speech batch, including: Clustering each input speech in the speech data set according to the length of each input speech in the speech data set corresponding to the target prediction model by a preset clustering algorithm to obtain at least two clustering results; wherein each clustering result is composed of input speech of similar length; Each clustering result is determined as an input speech batch.
3. The method for determining the model loss value according to claim 1, characterized in that: Preprocessing the predicted text sequence and the real text sequence of the input speech in each input speech batch according to the length of the predicted text sequence and the real text sequence of the input speech in each input speech batch includes: For each batch of input speech, perform the following operations: Determine the maximum value of the length of the predicted text sequence and the true text sequence of each input speech in the input speech batch; Perform text insertion on the predicted text sequence and the real text sequence of each input speech in the input speech batch whose length is less than the maximum value, and adjust the length of the sequence whose length is less than the maximum value to the maximum value.
4. The method for determining the model loss value according to claim 1, characterized in that: The loss value parameter calculation unit calculates the model loss value calculation parameters of the input speech in each input speech batch according to the preprocessed predicted text sequence, the preprocessed real text sequence and the model loss value related values of the input speech in each input speech batch, including: For each batch of input speech, perform the following operations: The preprocessed predicted text sequence, the preprocessed real text sequence and the model loss value related numerical values of each input speech in the input speech batch are respectively input into the loss value parameter calculation unit, and the model loss value calculation parameters of each input speech calculated by the loss value parameter calculation unit according to the preprocessed predicted text sequence, the preprocessed real text sequence and the model loss value related numerical values of each input speech are obtained.
5. The method for determining the model loss value according to claim 1, characterized in that: The method of calculating the model loss value of the input speech in each input speech batch according to the model loss value of the input speech in each input speech batch comprises: For each batch of input speech, perform the following operations: The model loss value calculation parameters of each input speech in the input speech batch are respectively substituted into the preset loss value calculation formula to obtain the model loss value of each input speech.
6. The method for determining the model loss value according to claim 1, characterized in that: Also includes: According to the business scenario of the speech data set, the predicted text sequence and the real text sequence of the input speech in each input speech batch are quantized.
7. The method for determining the model loss value according to claim 1, characterized in that: After calculating the parameters according to the model loss values of the input speech in each input speech batch and determining the model loss values of the input speech in each input speech batch, the method further includes: The model loss value of the input speech in each input speech batch is provided to the target user corresponding to the target prediction model.
8. A device for determining a model loss value, characterized in that: include: A function replacement module, used to replace the logarithmic multiplication function in the loss value parameter calculation unit corresponding to the target prediction model with a logarithmic addition function; A data set analysis module, used to perform feature analysis on a speech data set corresponding to the target prediction model, and pack input speech of similar length into an input speech batch; A sequence determination module, used for inputting the input speech in each input speech batch into the target prediction model to obtain a predicted text sequence of the input speech in each input speech batch; A sequence preprocessing module, used for preprocessing the predicted text sequence and the real text sequence of the input speech in each input speech batch according to the length of the predicted text sequence and the real text sequence of the input speech in each input speech batch; A parameter calculation module, configured to obtain the model loss value calculation parameters of the input speech in each input speech batch by performing calculation according to the preprocessed predicted text sequence, the preprocessed real text sequence and the model loss value related values of the input speech in each input speech batch through the loss value parameter calculation unit; The loss value determination module is used to determine the model loss value of the input speech in each input speech batch according to the model loss value calculation parameters of the input speech in each input speech batch.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein, the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for determining the model loss value according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for determining a model loss value according to any one of claims 1 to 7 when executed.