Cross-Domain Voice Recognition Model Optimization Method
By extracting and analyzing the time-frequency characteristics of sounds in different domains, determining the differential characteristics, and prioritizing storage based on the priority storage coefficient, the problems of low recognition accuracy and excessive storage burden in cross-domain sound recognition are solved, and higher recognition accuracy and storage optimization are achieved.
Patent Information
- Application Number
- CN202510257266.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-05
AI Technical Summary
When traditional sound recognition models process sound data in different domains, their recognition accuracy is low, and because they cannot effectively process a large number of differential features, the storage burden is too heavy, which affects the operational efficiency and recognition accuracy.
By extracting the time-frequency characteristics of sounds in different domains, performing feature differences analysis, obtaining different features, and prioritizing the different features through priority storage coefficients to reduce the storage burden.
The recognition accuracy of the cross-domain sound recognition model is improved, storage is optimized, and the storage burden of the model is reduced, thereby improving the recognition accuracy of sounds in different domains.
Smart Images

Figure CN119785823B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent voice recognition, and specifically relates to a method for optimizing a cross-domain voice recognition model. Background Art
[0002] In the field of voice recognition, with the continuous expansion of application scenarios, the sources and types of voices to be recognized are becoming increasingly complex and diverse. Cross-domain voice recognition faces many challenges. When traditional voice recognition models process voice data in different domains, there are often problems of low recognition accuracy. At the same time, in terms of model storage and performance optimization, the problem of storing a large number of different features cannot be effectively handled, resulting in an excessive storage burden on the model, which affects its operation efficiency and recognition accuracy.
[0003] In the prior art, the model cannot effectively handle the problem of storing a large number of different features, resulting in an excessive storage burden on the model, which affects its operation efficiency and recognition accuracy. Therefore, in this application, the characteristic difference value reflects the difference in the variation characteristics of the time-frequency characteristics of voices in different domains in the time domain, as well as the difference in the intensity of the time-frequency characteristics of voices in different domains in the frequency domain, so as to complete the recognition of different voices, determine the different characteristics of voices, improve the recognition accuracy of the cross-domain voice recognition model, and reflect the storage ratio of the cross-domain voice recognition model to different features during multiple historical storage cycles through the storage occupancy ratio, which is beneficial to evaluating the proportion of the storage content of different features by the cross-domain voice recognition model. At the same time, obtain the priority storage coefficient, and preferentially store different features according to the priority storage coefficient. When the priority storage memory is equal to the storage occupancy threshold, stop the work of storing different features, so as to solve the problem that when the cross-domain voice recognition model stores multiple different features, due to the large memory occupancy of multiple different features, they cannot be completely stored, and then reasonably optimize the storage to reduce the storage burden of the cross-domain voice recognition model and improve the recognition accuracy of the cross-domain voice recognition model for voices in different domains.
[0004] Therefore, the present invention provides a method for optimizing a cross-domain voice recognition model. Summary of the Invention
[0005] In order to make up for the deficiencies of the prior art and solve at least one of the technical problems proposed in the background art.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0007] A method for optimizing a cross-domain voice recognition model, including the following modules:
[0008] Step 1: During the recognition capture period, import the features of voices in different domains into the cross-domain voice recognition model, extract the time-frequency features, and obtain multiple groups of time-frequency feature sequences;
[0009] Step 2: Conduct feature difference analysis on the time-frequency feature sequences corresponding to different domain sounds to obtain the zero-crossing rate difference value and formant difference value. Perform quantization processing to obtain difference features, and conduct induction to obtain the feature difference degree sequence;
[0010] Step 3: In multiple historical storage cycles, obtain the feature byte ratio and feature quantity ratio, perform quantization processing, and compare them with the memory corresponding to all difference features in the feature difference degree sequence to generate a storage optimization signal;
[0011] Step 4: Based on the storage optimization signal, obtain the priority storage coefficient, and complete the priority storage of the difference features according to the priority storage coefficient.
[0012] As a further solution of the present invention: The extraction process of the time-frequency features is as follows:
[0013] Divide the recognition capture period into several capture time periods in the way of equal time intervals;
[0014] Process the time-frequency features through short-time Fourier transform to obtain time-frequency feature values;
[0015] Sort the time-frequency feature values corresponding to the time-frequency features of different sounds in all capture time periods within the recognition capture period according to the time series corresponding to the capture time periods to obtain multiple groups of time-frequency feature sequences.
[0016] As a further solution of the present invention: The acquisition process of the zero-crossing rate difference value is as follows:
[0017] Arbitrarily extract a time-frequency feature sequence as the feature comparison sequence, and form comparison analysis groups one by one with the remaining time-frequency feature sequences after extraction and the feature comparison sequence to obtain multiple groups of comparison analysis groups;
[0018] Arbitrarily extract a comparison analysis group, substitute the time-frequency feature values corresponding to all time-frequency features into the zero-crossing rate formula to obtain the zero-crossing rate corresponding to the time-frequency features at each capture node;
[0019] In the way of obtaining the zero-crossing rates corresponding to all time-frequency features in the feature comparison sequence, obtain the zero-crossing rates corresponding to all time-frequency features in the remaining time-frequency feature sequences after extraction, and obtain the zero-crossing rate difference value through the formula.
[0020] As a further solution of the present invention: The acquisition method of the formant difference value is as follows:
[0021] Construct a feature comparison change curve, extract the time-frequency feature values corresponding to all wave peaks, and integrate them according to the time series to obtain a comparison wave peak set;
[0022] Construct the time-frequency feature change curve, extract the time-frequency feature values corresponding to all wave peaks, and integrate them according to the time series to obtain the time-frequency feature wave peak set;
[0023] Substitute the comparison wave peak set and the time-frequency feature wave peak set into the Euclidean distance formula to obtain the resonance wave peak difference value ;
[0024] Extract the time-frequency feature values corresponding to all wave valleys, and integrate them according to the time series to obtain the comparison wave valley set;
[0025] Extract the time-frequency feature values corresponding to all wave valleys, and integrate them according to the time series to obtain the time-frequency feature wave valley set;
[0026] Substitute the comparison wave valley set and the time-frequency feature wave valley set into the Euclidean distance formula to obtain the resonance wave valley difference value ;
[0027] Sum the resonance wave peak difference value and the resonance wave valley difference value to obtain the resonance peak difference value.
[0028] As a further solution of the present invention: the method for obtaining the feature difference degree sequence is as follows:
[0029] Sum the zero-crossing rate difference value and the resonance peak difference value to obtain the feature difference value;
[0030] If the feature difference value is greater than the feature difference threshold, generate a difference feature signal, and mark the time-frequency feature of the generated difference feature signal as a difference feature;
[0031] Sort all the difference features from largest to smallest according to the corresponding feature difference values to obtain the feature difference degree sequence.
[0032] As a further solution of the present invention: the method for obtaining the feature quantity ratio is as follows:
[0033] Randomly extract a historical storage period;
[0034] Divide the historical storage period into several historical storage time periods, obtain the number of stored difference features within the historical storage time period, and calculate the ratio with the number of all features corresponding to the sound to obtain the unit storage quantity ratio;
[0035] Sum and average the unit storage quantity ratios corresponding to all historical storage time periods within the historical storage period to obtain the period storage quantity ratio;
[0036] Compare the ratio of the number of cycles stored corresponding to all historical storage cycles, select the maximum and minimum ratios of the number of cycles stored, and calculate the sum average to obtain the feature quantity ratio.
[0037] As a further solution of the present invention: The method for obtaining the feature byte ratio is as follows:
[0038] Obtain the bytes corresponding to the storage difference features within the historical storage period, and calculate the ratio with the total bytes of all features corresponding to the sound to obtain the unit storage byte ratio;
[0039] Calculate the sum average of the unit storage byte ratios corresponding to all historical storage periods within the historical storage cycle to obtain the cycle storage byte ratio;
[0040] Compare the cycle storage byte ratios corresponding to all historical storage cycles, select the maximum and minimum cycle storage byte ratios, and calculate the sum average to obtain the feature byte ratio.
[0041] As a further solution of the present invention: The generation process of the storage optimization signal is as follows:
[0042] Calculate the product of the feature byte ratio and the feature quantity ratio to obtain the storage occupancy ratio;
[0043] If the storage occupancy ratio is less than or equal to the storage occupancy threshold, generate a storage optimization signal.
[0044] As a further solution of the present invention: The method for obtaining the priority storage coefficient is as follows:
[0045] Extract the bytes corresponding to the difference features, calculate the ratio with the total bytes corresponding to all difference features to obtain the feature byte value, and sort all difference features from largest to smallest according to the feature byte value corresponding to the difference features to obtain the feature byte storage sequence;
[0046] Obtain the bytes of any difference feature within the feature difference degree sequence, calculate the ratio with the total bytes of all features corresponding to the sound, and calculate the sum average to obtain the feature difference byte ratio;
[0047] Overlap and pair the feature byte storage sequence with the feature difference degree sequence, extract the overlapping and paired difference features, and obtain the corresponding feature difference value and feature difference byte ratio. Calculate the ratio of the feature difference value and the feature difference byte ratio to obtain the priority storage coefficient.
[0048] As a further solution of the present invention: According to the priority storage coefficient, prioritize the storage of difference features, and the process is as follows:
[0049] Compare the priority storage coefficients corresponding to the differential features within the feature difference degree sequence, and preferentially store the differential features in descending order;
[0050] Dynamically count the feature difference byte ratios corresponding to the preferentially stored differential features to obtain the preferential storage memory until the preferential storage memory is equal to the storage ratio threshold, and stop the differential feature storage work.
[0051] The beneficial effects of the present invention are as follows:
[0052] During the recognition capture period of the present invention, different domain voice features are imported into the cross-domain voice recognition model to obtain multiple groups of time-frequency feature sequences, and feature difference analysis is performed based on the multiple groups of time-frequency feature sequences to obtain feature difference values. The feature difference values reflect the differences in the variation characteristics of the time-frequency features of different domain voices in the time domain and the difference degrees of the intensities of the time-frequency features of different domain voices in the frequency domain, thereby completing the recognition of different voices, determining the differential features of different voices, and improving the recognition accuracy of the cross-domain voice recognition model;
[0053] Based on the differential feature signal, the present invention obtains the memory information of the differential features within multiple historical storage periods through PyTorch, and performs storage ratio analysis to obtain the storage ratio value. The storage ratio value reflects the storage ratio of the cross-domain voice recognition model to the differential features within multiple historical storage periods, which is beneficial to evaluating the proportion of the storage content of the differential features by the cross-domain voice recognition model, thereby storing the differential features of different voices specifically and providing data support for improving the recognition accuracy of the cross-domain voice recognition model for voices;
[0054] Based on the storage optimization signal, the present invention obtains the priority storage coefficient, preferentially stores the differential features according to the priority storage coefficient, and at the same time, dynamically obtains the preferential storage memory. When the preferential storage memory is equal to the storage ratio threshold, the differential feature storage work is stopped, thereby solving the problem that when the cross-domain voice recognition model stores multiple differential features, due to the large memory occupation of multiple differential features, they cannot be completely stored, and further reasonably optimizing the storage to reduce the storage burden of the cross-domain voice recognition model and improving the recognition accuracy of the cross-domain voice recognition model for different domain voices. Description of the Drawings
[0055] The present invention will be further described below with reference to the drawings.
[0056] Figure 1 It is a flowchart of the cross-domain voice recognition model optimization method of the present invention. Detailed Embodiments
[0057] In order to make the technical means, creative features, achieved objectives and effects realized by the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.
[0058] Embodiment 1
[0059] As Figure 1 shown, the cross-domain voice recognition model optimization method described in the embodiment of the present invention includes:
[0060] Step 1: During the recognition capture period, import the features of voices in different domains into the cross-domain voice recognition model for time-frequency feature extraction to obtain multiple groups of time-frequency feature sequences;
[0061] In some embodiments, the recognition capture period is divided into several capture time periods in an equal time interval manner;
[0062] Extract the time-frequency features of the voice through short-time Fourier transform, and the process is as follows:
[0063] A1. Select the window function ;
[0064] It should be noted that the role of the window function is to only consider the voice features within the capture time period covered by the window function when analyzing the time-frequency features of the voice;
[0065] A2. Divide the capture time period into n capture nodes in an equal time interval manner, and the time interval duration between adjacent capture nodes is equal;
[0066] Mark the voice feature corresponding to each capture node as , where T is the frame shift, that is, the time interval duration between adjacent capture nodes, represents the original voice feature value;
[0067] A3. Perform discrete Fourier transform on the voice feature corresponding to each capture node to transform it from the time domain to the frequency domain. The formula is: , calculate the time-frequency feature value corresponding to the nth capture node, where =0, 1,..., n-1, is an imaginary number;
[0068] Sort the time-frequency feature values corresponding to all capture time periods within the recognition capture period of different voices according to the time series corresponding to the capture time periods to obtain multiple groups of time-frequency feature sequences;
[0069] Step 2: Conduct feature difference analysis on the time-frequency feature sequences corresponding to different domain sounds to obtain feature difference data. The feature difference data includes the zero-crossing rate difference value and the formant difference value. Quantify the zero-crossing rate difference value and the formant difference value to obtain a feature difference value, compare it with the feature difference threshold to obtain a difference feature, and summarize the difference features to obtain a feature difference degree sequence;
[0070] In some embodiments, randomly extract one time-frequency feature sequence as a feature comparison sequence;
[0071] Form a group of comparison analysis groups by combining the remaining time-frequency feature sequences after extraction with the feature comparison sequence one by one to obtain multiple groups of comparison analysis groups;
[0072] Exemplarily, randomly extract one group of comparison analysis groups;
[0073] Substitute the time-frequency feature values corresponding to all time-frequency features in the feature comparison sequence into the formula:
[0074]
[0075] Calculate the zero-crossing rate corresponding to all time-frequency features in the feature comparison sequence , where T represents the time interval duration between adjacent capture nodes, is the time-frequency feature value corresponding to the time-frequency feature of the th capture node, represents the time-frequency feature value corresponding to the time-frequency feature of the nth capture node, () is a compliance function;
[0076] When X>0, () = 1, when X = 0, () = 0, when X<0, () = -1; It should be noted that the zero-crossing rate corresponding to all time-frequency features in the feature comparison sequence is calculated by counting the number of times the adjacent capture nodes comply with the change , divided by is to normalize the zero-crossing rate to the number of zero-crossings per unit time; Similarly, in the same way as obtaining the zero-crossing rate corresponding to all time-frequency features in the feature comparison sequence , obtain the zero-crossing rate corresponding to all time-frequency features in the remaining time-frequency feature sequences after extraction ;
[0077] Calculate the difference between the zero-crossing rates within the comparison analysis group through the formula to obtain the zero-crossing rate difference value. The process is as follows:
[0078] A1. Align the corresponding time-frequency features of two sounds within the comparison analysis group in terms of time and frequency to ensure the same specific time and frequency, as well as the same time interval duration between the same adjacent capture nodes;
[0079] A2. Through the formula: , calculate the zero-crossing rate difference value ;
[0080] Exemplarily, within the comparison analysis group, substitute all time-frequency features within the feature comparison sequence into a two-dimensional coordinate system to obtain a feature comparison change curve, where the X-axis is time and the Y-axis is the time-frequency feature;
[0081] Extract the time-frequency feature values corresponding to all the peaks within the feature comparison change curve and integrate them according to the time sequence to obtain a comparison peak set , where represents the time-frequency feature value corresponding to the m-th peak within the feature comparison change curve, and m represents the total number of time-frequency feature values corresponding to the peaks within the feature comparison change curve;
[0082] Similarly, substitute all time-frequency features within the remaining time-frequency feature sequence after extraction into a two-dimensional coordinate system to obtain a time-frequency feature change curve, where the X-axis is time and the Y-axis is the time-frequency feature;
[0083] Extract the time-frequency feature values corresponding to all the peaks within the time-frequency feature change curve and integrate them according to the time sequence to obtain a time-frequency feature peak set , where represents the time-frequency feature value corresponding to the m-th peak within the time-frequency feature change curve, and m represents the total number of time-frequency feature values corresponding to the peaks within the time-frequency feature change curve;
[0084] It should be noted that the total number of time-frequency feature values corresponding to the peaks within the time-frequency feature change curve is the same as the total number of time-frequency feature values corresponding to the peaks within the feature comparison change curve;
[0085] Through the Euclidean distance formula: , calculate the resonance peak difference value , where represents the time-frequency feature value corresponding to the m-th peak within the feature comparison change curve, represents the time-frequency feature value corresponding to the m-th peak within the time-frequency feature change curve, and m represents both the total number of time-frequency feature values corresponding to the peaks within the feature comparison change curve and the total number of time-frequency feature values corresponding to the peaks within the time-frequency feature change curve;
[0086] Similarly, extract the time-frequency feature values corresponding to all the valleys within the feature comparison change curve and integrate them according to the time sequence to obtain a comparison valley set , where represents the time-frequency eigenvalue corresponding to the r-th trough in the feature comparison change curve, and r represents the total number of time-frequency eigenvalues corresponding to the troughs in the feature comparison change curve;
[0087] Extract the time-frequency eigenvalues corresponding to all the troughs in the time-frequency feature change curve and integrate them according to the time series to obtain the time-frequency feature trough set , where represents the time-frequency eigenvalue corresponding to the r-th peak in the time-frequency feature change curve, and r represents the total number of time-frequency eigenvalues corresponding to the peaks in the time-frequency feature change curve;
[0088] It should be noted that the total number of time-frequency eigenvalues corresponding to the troughs in the time-frequency feature change curve is the same as the total number of time-frequency eigenvalues corresponding to the troughs in the feature comparison change curve;
[0089] Through the Euclidean distance formula: , the resonance trough difference value is calculated , where represents the time-frequency eigenvalue corresponding to the r-th peak in the feature comparison change curve, represents the time-frequency eigenvalue corresponding to the r-th peak in the time-frequency feature change curve, and r represents the total number of time-frequency eigenvalues corresponding to the peaks in the feature comparison change curve, as well as the total number of time-frequency eigenvalues corresponding to the peaks in the time-frequency feature change curve;
[0090] The resonance peak difference value and the resonance trough difference value are summed up to obtain the resonance peak difference value;
[0091] The zero-crossing rate difference value and the resonance peak difference value are summed up to obtain the feature difference value;
[0092] It can be understood that the meaning represented by the feature difference value is: integrating the feature differences of the sound in the time domain (zero-crossing rate) and the frequency domain (resonance peak), and used to measure the overall difference size between the time-frequency feature sequences of sounds in different domains. On the one hand, the zero-crossing rate difference value reflects the difference degree of the number of times the signal crosses the zero value per unit time in the time-frequency feature sequences of sounds in different domains, reflecting the difference in the change characteristics of the time-frequency features of sounds in different domains in the time domain. On the other hand, the resonance peak difference value reflects the difference degree of the intensity of the time-frequency features of sounds in different domains in the frequency domain;
[0093] Compare the feature difference value with the feature difference threshold, and the process is as follows:
[0094] If the feature difference value is greater than the feature difference threshold, it indicates that the variation characteristics of the time-frequency features of different-domain sounds in the time domain are quite different, and the difference degree of the intensities in the frequency domain is relatively large. Generate a difference feature signal, and mark the time-frequency features of the generated difference feature signal as difference features;
[0095] Sort all the difference features in descending order according to the corresponding feature difference values to obtain a feature difference degree sequence;
[0096] If the feature difference value is less than or equal to the feature difference threshold, it indicates that the variation characteristics of the time-frequency features of different-domain sounds in the time domain are relatively small, and the difference degree of the intensities in the frequency domain is relatively small. Generate a non-difference feature signal, and mark the time-frequency features of the generated non-difference feature signal as non-difference features;
[0097] The specific implementation manner of the embodiment of the present invention is as follows: within the recognition capture period, import the different-domain sound features into the cross-domain sound recognition model to obtain multiple groups of time-frequency feature sequences, and perform feature difference analysis based on the multiple groups of time-frequency feature sequences to obtain feature difference values. The feature difference values reflect the variation characteristics of the time-frequency features of different-domain sounds in the time domain and the difference degree of the intensities of the time-frequency features of different-domain sounds in the frequency domain, so as to complete the recognition of different sounds, determine the difference features of different sounds, and improve the recognition accuracy of the cross-domain sound recognition model.
[0098] Embodiment 2
[0099] As Figure 1 shown, on the basis of Embodiment 1, the cross-domain sound recognition model optimization method described in the embodiment of the present invention includes:
[0100] Step 3: Within multiple historical storage periods, obtain the memory information of the difference features through PyTorch. Among them, the memory information includes the feature byte ratio and the feature quantity ratio. Perform quantization processing on the feature byte ratio and the feature quantity ratio to obtain a storage occupancy ratio, and compare it with the storage occupancy threshold. If the storage occupancy ratio is greater than the storage occupancy threshold, generate a storage optimization signal;
[0101] In some embodiments, within multiple historical storage periods, randomly select one historical storage period;
[0102] Divide the historical storage period into several historical storage time segments, obtain the number of stored difference features within the historical storage time segment, and perform a ratio calculation with all the feature quantities corresponding to the sound to obtain a unit storage quantity ratio;
[0103] Perform a sum mean calculation on the unit storage quantity ratios corresponding to all the historical storage time segments within the historical storage period to obtain a period storage quantity ratio;
[0104] Compare the ratio of the number of cycles stored corresponding to all historical storage cycles, select the maximum and minimum ratios of the number of cycles stored, and calculate the sum mean to obtain the feature quantity ratio;
[0105] Obtain the bytes corresponding to the storage difference features within the historical storage period, and calculate the ratio with the total bytes of all features corresponding to the sound to obtain the unit storage byte ratio;
[0106] Calculate the sum mean of the unit storage byte ratios corresponding to all historical storage periods within the historical storage cycle to obtain the cycle storage byte ratio;
[0107] Compare the cycle storage byte ratios corresponding to all historical storage cycles, select the maximum and minimum cycle storage byte ratios, and calculate the sum mean to obtain the feature byte ratio;
[0108] Calculate the product of the feature byte ratio and the feature quantity ratio to obtain the storage occupancy ratio;
[0109] It can be understood that the meaning represented by the storage occupancy ratio is: it reflects the storage ratio of the cross-domain voice recognition model to the difference features within multiple historical storage cycles. Specifically, the storage occupancy ratio is obtained by calculating the product of the feature byte ratio and the feature quantity ratio. Therefore, on the one hand, from the perspective of the feature byte ratio, it reflects the relative proportion of the size of the difference feature bytes in the total bytes of all features of the voice. On the other hand, from the perspective of the feature quantity ratio, it reflects the relative proportion of the number of difference features in the total number of all features of the voice, which is beneficial to evaluating the proportion of the storage content of the cross-domain voice recognition model to the difference features, so as to store the difference features of different voices targeted, providing data support for improving the accuracy of the cross-domain voice recognition model for voice recognition;
[0110] Compare the storage occupancy ratio with the storage occupancy threshold, and the process is as follows:
[0111] If the storage occupancy ratio is greater than the storage occupancy threshold, it means that the storage occupancy of the difference features is small, and a storage non-optimization signal is generated;
[0112] If the storage occupancy ratio is less than or equal to the storage occupancy threshold, it means that the storage occupancy of the difference features is large, and a storage optimization signal is generated;
[0113] It should be noted that the way to obtain the storage occupancy threshold is:
[0114] Calculate the ratio of the number of all difference features in the feature difference degree sequence to the total number of time-frequency features corresponding to all different-domain voices to obtain the feature difference quantity ratio;
[0115] Obtain the bytes of any differential feature within the sequence of differential feature degrees, calculate the ratio with the total bytes of all features corresponding to the sound, and perform a summation average calculation to obtain the differential feature byte ratio;
[0116] Perform a product calculation on the proportion of differential feature quantity and the differential feature byte ratio to obtain the storage proportion threshold;
[0117] Based on the storage non-optimization signal, compare the sizes of the differential feature values corresponding to the differential features, and store them in descending order;
[0118] The specific implementation manner of the embodiment of the present invention is: Based on the differential feature signal, use PyTorch to obtain the memory information of the differential features in multiple historical storage cycles, and perform storage proportion analysis to obtain the storage proportion value. The storage proportion value reflects the storage proportion of the cross-domain voice recognition model for differential features in multiple historical storage cycles, which is beneficial to evaluating the storage content proportion of the cross-domain voice recognition model for differential features, so as to specifically store the differential features of different voices, providing data support for improving the accuracy of the cross-domain voice recognition model for voice recognition.
[0119] Embodiment 3
[0120] As Figure 1 shown, on the basis of Embodiment 1 and Embodiment 2, the voice recognition model optimization method described in the embodiment of the present invention includes:
[0121] Step 4: Based on the storage optimization signal, obtain the priority storage coefficient, and preferentially store the differential features according to the priority storage coefficient to complete the storage optimization work of the features of the voice recognition model;
[0122] In some embodiments, extract the bytes corresponding to the differential features, calculate the ratio with the total amount of bytes corresponding to all differential features to obtain the feature byte value, and sort all differential features in descending order according to the feature byte value corresponding to the differential features to obtain the feature byte storage sequence;
[0123] Obtain the bytes of any differential feature within the sequence of differential feature degrees, calculate the ratio with the total bytes of all features corresponding to the sound, and perform a summation average calculation to obtain the differential feature byte ratio;
[0124] Perform an overlapping pairing of the feature byte storage sequence and the differential feature degree sequence, extract the overlapping paired differential features, and obtain the corresponding differential feature value and differential feature byte ratio. Perform a ratio calculation on the differential feature value and the differential feature byte ratio to obtain the priority storage coefficient;
[0125] Based on the priority storage coefficient, compare the priority storage coefficients corresponding to the differential features in the feature difference degree sequence in terms of size, and preferentially store the differential features in descending order;
[0126] Dynamically count the feature difference byte ratios corresponding to the preferentially stored differential features to obtain the preferential storage memory until the stored preferential storage memory is equal to the storage ratio threshold, and stop the differential feature storage work;
[0127] The specific implementation manner of the embodiment of the present invention is as follows: Based on the storage optimization signal, obtain the priority storage coefficient, and preferentially store the differential features according to the priority storage coefficient. At the same time, dynamically obtain the preferential storage memory. When the preferential storage memory is equal to the storage ratio threshold, stop the differential feature storage work, so as to solve the problem that when the cross-domain voice recognition model stores multiple differential features, due to the large memory occupation of multiple differential features, they cannot be completely stored, and then reasonably optimize the storage to reduce the storage burden of the cross-domain voice recognition model and improve the recognition accuracy of the cross-domain voice recognition model for voices in different domains.
[0128] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A cross-domain sound recognition model optimization method, characterized in that: include: Step 1: In the recognition and capture cycle, the features of sounds in different domains are imported into the cross-domain sound recognition model, and the time-frequency features are extracted to obtain multiple groups of time-frequency feature sequences; Step 2: Perform feature difference analysis on the time-frequency feature sequences corresponding to sounds in different domains, obtain zero-crossing rate difference values and formant difference values, perform quantization processing to obtain difference features, and summarize to obtain a feature difference degree sequence; Step 3: In multiple historical storage cycles, obtain the feature byte ratio and feature quantity ratio, perform quantification processing, and compare them with the memory corresponding to all difference features in the feature difference degree sequence to generate a storage optimization signal; Step 4: Based on the storage optimization signal, obtain the priority storage coefficient, and complete the priority storage of the difference features according to the priority storage coefficient.
2. The cross-domain sound recognition model optimization method according to claim 1, characterized in that: The extraction process of time-frequency features is as follows: Divide the recognition and capture cycle into a number of capture time periods at equal time intervals; The time-frequency features are processed by short-time Wavelet transform to obtain the time-frequency feature values; The time-frequency feature values corresponding to all capture time periods in the recognition capture period of different sounds are sorted according to the time series corresponding to the capture time periods to obtain multiple groups of time-frequency feature sequences.
3. The cross-domain sound recognition model optimization method according to claim 2, characterized in that: The process of obtaining the zero-crossing rate difference value is as follows: A time-frequency feature sequence is randomly extracted as a feature comparison sequence, and the remaining time-frequency feature sequences after the extraction are combined with the feature comparison sequence one by one to form a comparison analysis group, thereby obtaining multiple comparison analysis groups; Randomly select a comparison analysis group, substitute the time-frequency feature values corresponding to all time-frequency features into the zero-crossing rate formula, and obtain the zero-crossing rate corresponding to the time-frequency feature of each capture node; By obtaining the zero-crossing rates corresponding to all time-frequency features in the feature comparison sequence, the zero-crossing rates corresponding to all time-frequency features in the remaining time-frequency feature sequence after extraction are obtained, and the zero-crossing rate difference value is obtained through the formula.
4. The cross-domain sound recognition model optimization method according to claim 2, characterized in that: The resonance peak difference value is obtained as follows: Construct a feature comparison change curve, extract the time-frequency feature values corresponding to all peaks, and integrate them according to the time series to obtain a comparison peak set; Construct a time-frequency feature change curve, extract the time-frequency feature values corresponding to all peaks, and integrate them according to the time series to obtain a set of time-frequency feature peaks; Substitute the comparison peak set and the time-frequency feature peak set into the Euclidean distance formula to obtain the resonance peak difference value ; Extract the time-frequency feature values corresponding to all troughs and integrate them according to the time series to obtain the comparison trough set; Extract the time-frequency feature values corresponding to all troughs and integrate them according to the time series to obtain the time-frequency feature trough set; Substitute the comparison trough set and the time-frequency feature trough set into the Euclidean distance formula to obtain the resonance trough difference value ; The resonance peak difference Difference from resonance trough Perform sum calculation to obtain the resonance peak difference value.
5. The cross-domain sound recognition model optimization method according to claim 1, characterized in that: The feature difference degree sequence is obtained as follows: The zero-crossing rate difference value and the resonance peak difference value are summed up to obtain a characteristic difference value; If the feature difference value is greater than the feature difference threshold, a difference feature signal is generated, and the time-frequency feature of the generated difference feature signal is marked as a difference feature; Sort all difference features from large to small according to the corresponding feature difference values. Get the feature difference degree sequence.
6. The cross-domain sound recognition model optimization method according to claim 5, characterized in that: The feature quantity ratio is obtained as follows: Arbitrarily extract a historical storage period; Divide the historical storage cycle into several historical storage periods, obtain the number of stored difference features in the historical storage period, and calculate the ratio with the number of all features corresponding to the sound to obtain the unit storage quantity ratio; The unit storage quantity ratios corresponding to all historical storage periods in the historical storage cycle are summed and averaged to obtain the period storage quantity ratio; The period storage quantity ratios corresponding to all historical storage periods are compared, the maximum period storage quantity ratio and the minimum period storage quantity ratio are selected, and the sum and mean are calculated to obtain the characteristic quantity ratio.
7. The cross-domain sound recognition model optimization method according to claim 6, characterized in that: The characteristic byte ratio is obtained as follows: Obtain the bytes corresponding to the storage difference features during the historical storage period, and calculate the ratio with the total bytes of all features corresponding to the sound to obtain the unit storage byte ratio; The unit storage byte ratios corresponding to all historical storage periods in the historical storage cycle are summed and averaged to obtain the period storage byte ratio; Compare the cycle storage byte ratios corresponding to all historical storage cycles, select the maximum cycle storage byte ratio and the minimum cycle storage byte ratio, and calculate the sum and average. Get the characteristic byte ratio.
8. The cross-domain sound recognition model optimization method according to claim 1, characterized in that: The generation process of storage optimization signal is as follows: The feature byte ratio is multiplied by the feature quantity ratio to obtain the storage ratio value; If the storage ratio value is less than or equal to the storage ratio threshold, a storage optimization signal is generated.
9. The cross-domain sound recognition model optimization method according to claim 7, characterized in that: The priority storage coefficient is obtained as follows: Extract the bytes corresponding to the difference features, calculate the ratio with the total amount of bytes corresponding to all the difference features, obtain the feature byte value, and sort all the difference features from large to small according to the feature byte values corresponding to the difference features, and obtain the feature byte storage sequence; Obtain the byte of any difference feature in the feature difference degree sequence, calculate the ratio with the total bytes of all features corresponding to the sound, and perform summation and mean calculation to obtain the feature difference byte ratio; The feature byte storage sequence is overlapped and paired with the feature difference degree sequence, the difference features of the overlapped pairing are extracted, and the corresponding feature difference value and feature difference byte ratio are obtained, and the feature difference value and feature difference byte ratio are calculated to obtain the priority storage coefficient.
10. The cross-domain sound recognition model optimization method according to claim 9, characterized in that: The difference features are stored preferentially according to the priority storage coefficient. The process is as follows: Compare the priority storage coefficients corresponding to the difference features in the feature difference degree sequence, and store the difference features in a priority order from large to small; The feature difference byte ratios corresponding to the priority stored difference features are dynamically counted to obtain the priority storage memory, until the priority storage memory is equal to the storage ratio threshold, and the difference feature storage work is stopped.
Citation Information
Patent Citations
Health status monitoring system based on speech analysis
AU2020102516A4
Small sample voice recognition method and system based on cross-domain transfer learning
CN114299986A