Voice intelligent model test data generation method and system guided by multiple coverage rates

Through a variety of coverage-guided voice intelligent model test data generation methods, the problem of insufficient diversity of test data in the prior art is solved, more sufficient and effective testing is achieved, and potential security defects of the model can be better exposed.

CN120032628APending Publication Date: 2025-05-23BEIJING XUANYU INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510017065.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing technology is difficult to generate diverse and effective voice intelligent model test data, resulting in insufficient testing and lack of full exposure to potential security defects of the model.

Method used

A variety of coverage-guided speech intelligent model test data generation methods are used to analyze the activation status of the model to be tested under different neural network coverage indicators, randomly select the mutation strategy to mutate the seed test data, define the number of strong mutation constraints, and filter and select the variation data that improves the coverage of the model.

Benefits of technology

Improve the diversity and effectiveness of test data, ensure the adequacy of voice intelligent model testing, and expose as many potential security flaws of the model under test as possible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032628A_ABST
    Figure CN120032628A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-coverage-rate-guided voice intelligent model test data generation method and system. The method comprises the steps that a to-be-tested voice intelligent model and a seed test data set are input; the coverage rate of the voice intelligent model to be tested under different neural network coverage rate indexes is analyzed; randomly selecting one or more variation strategies to perform variation on the seed test data; defining a strong variation number constraint condition, preliminarily filtering variation data, calculating four neural network coverage rates of the voice intelligent model to be tested by using the filtered variation data, and selecting variation data of which the coverage rate of the model to be tested is improved compared with that before variation to form variation data sets corresponding to different neural network coverage rate criteria; and outputting a coverage rate calculation result of the to-be-tested voice intelligent model on the seed test data set and the selected variation test data set. The diversity of test data is improved, and the sufficiency of voice intelligent model test is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for generating speech intelligent model test data guided by multiple coverage rates, and belongs to the technical field of intelligent software testing. Background Art

[0002] Like traditional software, voice intelligence models also need to be tested before deployment. Testing voice intelligence models in the early stages helps to discover potential defects and failures. However, due to the significant differences between voice intelligence models and traditional software in terms of technical architecture and development process, traditional software testing methods are difficult to directly apply to voice intelligence model testing. The performance of voice intelligence models is highly dependent on training data and model training parameters. During the inference stage, voice intelligence models exhibit certain uncertainty and volatility. Therefore, voice intelligence models need to be fully tested.

[0003] Sufficient testing is one of the key testing methods to ensure the credibility of voice intelligence models. Test data is the basis of testing. Diversified test data (different accents, speaking speeds, background noise) can not only comprehensively evaluate the performance and robustness of the model, but also ensure the fairness and generalization ability of the model, which is of great significance for revealing model defects. Manually collecting test data is inefficient and costly, and the diversity of test data is difficult to ensure. There is an urgent need to study the generation of voice test data.

[0004] The current status of research on speech data generation is as follows:

[0005] SpecAugment, proposed by the Google Brain team, randomly enhances the logarithmic Mel spectrum of audio data. Its enhancement strategies include feature distortion, frequency domain shielding, time domain shielding, etc. SpecAugment can synthesize a large amount of audio data, but due to the randomness of the process and lack of guidance, the diversity of the synthesized data is difficult to guarantee; DeepWalk, proposed by the Hong Kong University of Science and Technology, converts audio data into logarithmic Mel spectrum, constructs its image manifold, and then mutates the image manifold based on the generative adversarial network. DeepWalk can generate diverse and effective mutated data, but its training process requires a lot of computing resources and long training time; CrossASR and CrossASR++ proposed by the Singapore Management University generate speech data from text. Although they can generate a large amount of speech data, these methods are based on the text-to-speech conversion engine (TTS), require a large amount of text data set training, and the quality of the generated audio data depends on the training effect of TTS, and the quality is unstable.

[0006] The "Test Case Generation Method, Device, Equipment, Storage Medium and Program Product" (CN117932348A) proposed by Beijing New Energy Automobile Co., Ltd. describes a method for generating speech test data. It mutates the original data from three aspects: speech properties, noise properties and multi-tone properties of the speech data, but does not evaluate the effectiveness and diversity of the generated data.

[0007] "An Evaluation System for Deep Learning Models" proposed by Shanghai Anban Information Technology Co., Ltd.

[0008] (CN117493140A) The speech test set is mutated in terms of volume, speaking speed and pitch, and (five) neuron coverage index tests are performed on the mutated data. Euclidean distance similarity is used as the judgment criterion for data selection. However, the mutated data generated by this invention is intended to test the robustness of the model and is not suitable for testing the conventional performance indicators of the model.

[0009] In summary, the current test data generation methods for speech intelligence models have problems such as insufficient diversity, poor quality, and lack of effectiveness evaluation. Therefore, in order to generate diverse and effective speech test data, it is necessary to explore more efficient test data that is suitable for the characteristics of speech intelligence models to improve the adequacy and credibility of speech intelligence model testing. Summary of the invention

[0010] The technical problem solved by the present invention is: to overcome the shortcomings of the prior art and provide a method and system for generating speech intelligence model test data guided by multiple coverage rates, aiming to solve the problem of insufficient testing caused by insufficient diversity of speech intelligence model test data, improve the diversity of test data, ensure the adequacy of speech intelligence model testing, and expose as many potential security defects of the tested model as possible.

[0011] The technical solution of the present invention is: in the first aspect, a method for generating speech intelligence model test data guided by multiple coverage rates is provided, comprising:

[0012] Input data, including the voice intelligence model to be tested and the seed test data set;

[0013] Analyze the activation state of neurons in the neural network under a given seed test data set under different neural network coverage indicators of the voice intelligence model to be tested, and obtain the coverage under different neural network coverage indicators; neural network coverage indicators include: neuron coverage, Top-k neuron coverage, neuron strong activation coverage, and multi-node neuron coverage;

[0014] Randomly select one or more mutation strategies to mutate the seed test data; the mutation strategies include random silence, random gain, Gaussian noise, speech rate change, pitch shift, time shift, dynamic compression, and hyperbolic tangent distortion;

[0015] Define strong mutation number constraints, preliminarily filter mutation data, use the filtered mutation data to calculate the four neural network coverage rates of the speech intelligence model to be tested, select mutation data that improves the coverage rate of the model to be tested compared to before mutation, and form mutation data sets corresponding to different neural network coverage criteria;

[0016] Output the coverage calculation results of the speech intelligence model to be tested on the seed test dataset and the selected variant test dataset.

[0017] Preferably, the set of all neurons in the speech intelligence model to be tested is N={n 1 ,n 2 ,…,n m};

[0018] All seed test datasets are Among them, x i represents audio sample i.

[0019] Preferably, the activation state of neurons in the neural network under a given seed test data set is specifically:

[0020] Let O j (x i ) indicates that given a test input x i The number of neurons in the speech intelligence model to be tested is n j The output value of

[0021] A j (x i )∈{0,1} means the input is x i When neuron n j The activation state of , where:

[0022] When O j (x i )>O min When A j (x i ) = 1, indicating that the neuron is activated;

[0023] When O j (x i )≤O min When A j (x i )=0, indicating that the neuron is not activated.

[0024] Preferably, the neural network coverage criteria of neuron coverage, neuron strong activation coverage, Top-k neuron coverage, and multi-node neuron coverage are as follows:

[0025] Neuron coverage: The neuron coverage of a set of test inputs is defined as the ratio of the number of activated neurons in all test inputs to the total number of neurons in the model under test, that is, the neuron coverage is the ratio of the number of activated neurons to the total number of neurons; the definition is as follows:

[0026]

[0027] Neuron strong activation coverage: When training the model, each neuron n j The upper boundary value can be calculated based on the training set data analysis If the output value of a neuron is higher than its upper boundary value, the neuron is considered to be strongly activated; the strong activation coverage is the ratio of the number of strongly activated neurons to the total number of neurons; it is defined as follows:

[0028]

[0029] Top-k neuron coverage: Assuming a neural network has l layers, for the rth layer, use top k (x i ,r) means when the input is x i When , the neurons corresponding to the top K values ​​with the largest neuron output value in the rth layer; under all test inputs, the Top-k neuron coverage is the proportion of the top K most active neurons in all layers of the neural network to the total number of neurons after taking the union; it is defined as follows:

[0030]

[0031] Multi-section neuron coverage: Determine the activation range of each neuron based on the training data and divide this range into K equal segments. During testing, statistically test the input x i ∈X is the output value O of the j∈Nth neuron j (x i ) falls in the K segments of the jth neuron as a proportion of the total number of neuron segments; it is defined as follows:

[0032]

[0033] Among them, Fall k (O j (x i )) represents the test input x i The output value O of the jth neuron j (x i) falls in the kth segment of the jth neuron.

[0034] Preferably, the number of mutations m for the seed test data is defined by the user, and the seed data x is defined as i The result after the jth mutation is

[0035] The random silence, random gain, Gaussian noise, speech rate change, pitch shift, time shift, dynamic compression, and hyperbolic tangent distortion mutation strategies are:

[0036] Random silence: Add silence to a randomly selected part of the audio signal to simulate a brief signal loss or blockage; the audio data after adding random silence is represented as t represents time;

[0037] Random gain: Multiply the audio signal by a random amplitude factor to adjust the volume by increasing or decreasing the overall amplitude of the audio signal. The loudness of the audio signal will change, but the pitch and rhythm remain unchanged. The audio data after random gain is represented as x′ ij =α·x i , α>1 means increasing the volume, α<1 means decreasing the volume;

[0038] Gaussian Noise: Adds Gaussian noise to the audio signal The audio data after adding Gaussian noise is represented as x′ ij =x i +g;

[0039] Speech rate change: change the speed or duration of the audio signal without changing the pitch; the audio data after speech rate change is represented as Where α is the time stretching factor. α>1 means the audio is stretched and the speed is slowed down; α<1 means the audio is compressed and the speed is accelerated.

[0040] Pitch shift: Raise or lower the pitch of an audio signal without changing the audio rhythm or duration. The audio data after pitch shift is represented as p represents the shift operation, α represents the number of semitones of pitch shift, α>0 represents raising the pitch, and α<0 represents lowering the pitch;

[0041] Time Shift: Shift the audio signal forward or backward; the audio data after applying the shift is represented as Δt represents the amount of time for translation, positive values ​​represent forward translation, negative values ​​represent backward translation, and L represents the total duration of the input audio signal;

[0042] Dynamic compression: limits the maximum amplitude of the audio signal to prevent the signal from exceeding a preset threshold. The audio data after dynamic compression is represented as Where T represents the amplitude threshold, sgn(x i ) represents the signal x i The symbolic function of

[0043] Hyperbolic tangent distortion: By i Apply the hyperbolic tangent function to add distortion effect; the audio data after applying the hyperbolic tangent distortion is represented as x′ ij =tanh(k·x i ), k represents an adjustable gain factor, which is used to control the intensity of the distortion effect.

[0044] Preferably, the strong mutation number constraint condition is defined to initially filter the mutation data through the strong mutation number constraint, specifically:

[0045] Calculate the changes in the data values ​​before and after the mutation Seed data x i The result after the jth mutation is

[0046] Calculate the mean amplitude of the data changes before and after the mutation The number of values ​​that have changed statistically and exceed the mean amplitude (i.e., strong variation)

[0047] Determine x′ ij Whether the number of strong mutations does not exceed 50% of the data dimension, that is, If the constraint condition is satisfied, the next step is selected; otherwise, the mutation strategy is adopted again to mutate the seed test data.

[0048] Preferably, the filtered variant data is used to calculate the four neural network coverages of the speech intelligence model to be tested, and the variant data that improves the coverage of the model to be tested compared to before the mutation is selected to form a variant data set corresponding to different neural network coverage criteria, specifically:

[0049] The mutation data x′ that satisfies the strong mutation number constraint ij As the input of the speech intelligence model to be tested, the coverage of the tested model is calculated according to each neural network coverage criterion. If the coverage is improved compared with the previous result, the mutated data is retained, otherwise it is discarded and mutated again until all samples reach the set number of mutations of the seed test data; a set of corresponding mutated data sets is formed for each coverage criterion.

[0050] In the second aspect, a system for generating test data for a speech intelligence model guided by multiple coverage rates is provided, including: a data input module, a neuron activation state analysis module, a speech data variation module, a variation data selection module, and a result visualization module; specifically:

[0051] The data input module stores and inputs the speech intelligence model to be tested and the seed test data set to other modules;

[0052] The neuron activation state analysis module analyzes the activation state of each neuron in the neural network under a given seed test data set, and outputs the coverage under different neural network coverage indicators; the neural network coverage criteria include neuron coverage, Top-k neuron coverage, neuron strong activation coverage, and multi-section neuron coverage.

[0053] The speech data mutation module automatically mutates the seed test data set. The mutation strategies include random silence, random gain, Gaussian noise, speech rate change, pitch shift, time shift, dynamic compression, and hyperbolic tangent distortion.

[0054] The variant data selection module defines strong variant number constraints, automatically filters variant data that do not meet the constraints, and automatically calculates the neural network coverage of the speech intelligence model to be tested using the filtered variant data according to each neural network coverage criterion, and selects and retains variant data that improves model coverage;

[0055] The result visualization module presents the coverage calculation results of the speech intelligence model to be tested on the seed test dataset and the filtered variant test dataset in the form of charts, and displays the statistical results of the selected variant data.

[0056] Compared with the prior art, the present invention has the following advantages:

[0057] The present invention discloses a speech intelligence model test data generation system guided by multiple neural network coverage. The system first defines 8 speech data mutation strategies (random silence, random gain, Gaussian noise, speech rate change, pitch offset, time offset, dynamic compression, hyperbolic tangent distortion), automatically mutates the initial seed test case, and then introduces 4 neural network coverage criteria (neuron coverage, neuron strong activation coverage, Top-k neuron coverage and multi-network area neuron coverage), automatically selects and retains the mutated data, and calculates the coverage of the model to be tested on different neural network coverage criteria and its retained variant data. The higher the coverage of the model to be tested, the more diverse the test input data. The system guides the generation of variant data through the neural network coverage criterion, and provides targeted guidance for test input generation. The speech intelligence model test data generation system of the present invention can improve the diversity of test data, ensure the adequacy of speech intelligence model testing, and expose as many potential security defects of the model to be tested as possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1A framework diagram of a speech intelligence model test data generation system based on multiple neural network coverage guidance according to the present invention;

[0059] Figure 2 It is a flow chart of the speech intelligence model test data generation system guided by multiple neural network coverage in the present invention. DETAILED DESCRIPTION

[0060] The current speech test data generation methods have problems such as incomplete coverage of boundary conditions or extreme conditions, difficulty in ensuring data quality, and unstable results. It is difficult to ensure the effectiveness and diversity of speech test data, which may lead to insufficient testing and unreliable test results. The neural network coverage criterion can reflect the diversity of test inputs, guide the generation of test inputs, improve the quality and diversity of test inputs, and make up for the defects of poor data quality and insufficient diversity generated by traditional methods. The automated mutation and data selection process ensures the efficiency of data generation, thereby ensuring the adequacy and efficiency of testing.

[0061] The present invention proposes a multi-coverage guided speech intelligence model test data generation system, which includes five modules: data input module, neuron activation state analysis module, speech data variation module, variation data selection module and result visualization module. Specifically:

[0062] (1) Data input module: The input of this module includes the speech intelligence model to be tested and the seed test data set.

[0063] (2) Neuron activation state analysis module. This module analyzes the activation state of each neuron in the neural network under a given seed test data set, providing a basis for the calculation of neuron coverage, Top-k neuron coverage, neuron strong activation coverage, and multi-section neuron coverage.

[0064] (3) Voice data mutation module. This module automatically mutates the seed test data set, including eight audio data mutation strategies: random silence, random gain, Gaussian noise, speech rate change, pitch shift, time shift, dynamic compression, and hyperbolic tangent distortion.

[0065] (4) Mutation data selection module. This module first defines strong mutation number constraints, automatically filters mutation samples that do not meet the constraints, and then automatically calculates the coverage of the model under test under the mutated test data according to each neural network coverage criterion, and selects and retains mutation data that improves the model coverage.

[0066] (5) Result visualization module. This module presents the four coverage calculation results of the speech intelligence model to be tested on the seed test dataset and the selected variant test dataset in the form of charts, and displays the statistical results of the selected variant data.

[0067] The following is a further detailed description of the method for generating test data for a speech intelligence model guided by multiple coverage rates provided by an embodiment of the present invention in conjunction with the accompanying drawings. The system structure is as follows: Figure 1 As shown, the method flow of using this system to test data is as follows Figure 2 shown.

[0068] Step 1: Use the data input module to input the speech intelligence model to be tested and the seed test data set.

[0069] Specifically, all neurons of the speech intelligence model to be tested are represented by the set N = {n 1 ,n 2 ,…,n m} means that all seed test inputs are Indicates that, where x i represents audio sample i.

[0070] Step 2: Use the neuron activation state analysis module to analyze the activation state of each neuron in the neural network under the same seed test input and different neural network coverage criteria. This module includes four neural network coverage criteria: neuron coverage, neuron strong activation coverage, Top-k neuron coverage, and multi-section neuron coverage.

[0071] Specifically, O j (x i ) indicates that given a test input x i The number of neurons in the speech intelligence model to be tested is n j The output value of A j (x i )∈{0,1} to indicate that the input is x i When neuron n j The activation state, when O j (x i )>O min When A j (x i )=1, indicating that the neuron is activated. j (x i )≤O min When A j (x i )=0, indicating that the neuron is not activated.

[0072] The four coverage criteria for this module are as follows:

[0073] Neuron coverage (NC): The neuron coverage of a set of test inputs is defined as the ratio of the number of activated neurons in all test inputs to the total number of neurons in the model under test. In other words, neuron coverage is the ratio of the number of activated neurons to the total number of neurons. Neuron coverage is defined as follows:

[0074]

[0075] Strong Neuron Activation Coverage (SNAC): When training the model, each neuron n j The upper boundary value can be calculated based on the training set data analysis If the output value of a neuron is higher than its upper boundary value, the neuron is considered to be strongly activated. The strong activation coverage is the ratio of the number of strongly activated neurons to the total number of neurons. The definition of neuron strong activation coverage is as follows:

[0076]

[0077] Top-k neuron coverage (TKNC): Assume that a neural network has l layers. For the rth layer, use top k (x i ,r) means when the input is x i When , the neurons corresponding to the top K values ​​with the largest neuron output value in the rth layer. Under all test inputs, the Top-k neuron coverage is the proportion of the top K most active neurons in all layers of the neural network to the total number of neurons after taking the union. The definition of Top-k neuron coverage is as follows:

[0078]

[0079] Multi-section neuron coverage (KMNC): Determine the activation range of each neuron based on the training data and divide this range into K equal segments. During testing, the test input x is statistically analyzed. i ∈X is the output value O of the j∈Nth neuron j (x i ) is the ratio of the number of neurons that fall in the K segments of the jth neuron to the total number of neuron segments. The definition of multi-segment neuron coverage is as follows:

[0080]

[0081] Among them, Fall k (O j (x i )) represents the test input x i The output value O of the jth neuron j (x i ) falls in the kth segment of the jth neuron.

[0082] Based on the above principles, this module counts the activation value distribution of each neuron in the seed test data set to obtain the coverage under different neural network coverage indicators.

[0083] Step 3: Use the voice data mutation module to randomly select mutation strategies for the seed test dataset Mutate each seed test data in turn, and perform a maximum of m (m is user-defined, the default is 20) mutations. The seed data x i The result after the jth mutation is

[0084] The 8 mutation strategies of this module are as follows:

[0085] Random silence (TimeMask): Add silence to a randomly selected portion of the audio signal to simulate a brief signal loss or occlusion. The audio data after adding random silence is represented as t represents time.

[0086] Random gain (Gain): Multiply the audio signal by a random amplitude factor to adjust the volume by increasing or decreasing the overall amplitude of the audio signal. The loudness of the audio signal will change, but the pitch and rhythm remain unchanged. The audio data after random gain is represented as x′ ij =α·x i , α>1 means increasing the volume, and α<1 means decreasing the volume.

[0087] Gaussian Noise: Add Gaussian noise to the audio signal The audio data after adding Gaussian noise is represented as x′ ij =x i +g.

[0088] TimeStretch: Change the speed or duration of an audio signal without changing the pitch. The audio data after the time-stretch is expressed as Where α is the time stretching factor. α>1 means the audio is stretched and the speed is slowed down; α<1 means the audio is compressed and the speed is accelerated.

[0089] Pitch Shift: Raise or lower the pitch of an audio signal without changing the rhythm or duration of the audio. The audio data after pitch shift is represented as Indicates the shift operation, α indicates the number of semitones of pitch shift, α>0 indicates raising the pitch, and α<0 indicates lowering the pitch.

[0090] Time Shift: Shift the audio signal forward or backward. The audio data after applying the shift is represented as Δt represents the amount of time to pan, which can be positive (panning forward) or negative (panning backward), and L represents the total duration of the input audio signal.

[0091] Dynamic compression (Limiter): Limits the maximum amplitude of the audio signal to prevent the signal from exceeding the preset threshold. The audio data after dynamic compression is represented as Where T represents the amplitude threshold, sgn(x i ) represents the signal x i The symbol function of .

[0092] Hyperbolic tangent distortion (TanhDistortion): By i Apply the hyperbolic tangent function (tanh) to add distortion. The audio data after applying the hyperbolic tangent distortion is represented as x′ ij =tanh(k·x i ), k represents an adjustable gain factor, which is used to control the intensity of the distortion effect.

[0093] Step 4: Use the variant data selection module to first define strong variant number constraints to perform preliminary filtering on the mutated data, then calculate the coverage of the model to be tested under the filtered variant data according to each neural network coverage criterion, and further select variant data that improves the neural network coverage.

[0094] Specifically, for the seed test data x i Speech data after the jth mutation And the voice intelligence model to be tested, the mutation data selection is carried out through the following steps:

[0095] (a) Define the strong mutation number constraint condition, perform preliminary filtering on the mutation data, and retain the mutation data whose strong mutation number does not exceed 50% of the data dimension. The strong mutation number constraint is as follows:

[0096] Assume that the mutated data is x′ ij , calculate the changes in the data values ​​before and after the mutation Calculate the mean amplitude of the data changes before and after the mutation The number of values ​​that have changed statistically and exceed the mean amplitude (i.e., strong variation) Check x′ ij Does the number of strong mutations not exceed 50% of the data dimension (i.e. ) of the mutation data, if the constraint conditions are met, the next step is selected, otherwise it goes to step three for the next mutation.

[0097] (b) The mutation data x′ that satisfies the strong mutation number constraint is ij As the input of the tested voice intelligence model, the coverage of the tested model is calculated according to each neural network coverage criterion (neuron coverage, neuron strong activation coverage, Top-k neuron coverage, multi-section neuron coverage). If the coverage is improved compared with the previous result, the variant data is retained, otherwise it is discarded and the next mutation is performed in step three. Steps three and four are repeated until all samples have been mutated m times, and a corresponding mutated data set is formed for each coverage criterion.

[0098] Step 5: Result visualization module: This module presents the four coverage calculation results of the speech intelligence model to be tested on the seed test data set and the selected variant test data set in the form of charts, and displays the statistical results of the selected variant data.

[0099] Specifically, for the coverage calculation results, this module uses a line chart to show the trend of each coverage on the seed test data and variant data, and uses a radar chart to show the differences between different coverage criteria.

[0100] The present invention proposes a method and system for generating test data for a speech intelligence model. In order to solve the problem of insufficient testing caused by insufficient diversity of test data for speech intelligence models, a speech test data generation method guided by four neural network coverage rates is proposed, wherein the speech data mutation method includes eight types, namely, random silence, random gain, Gaussian noise, speech rate change, pitch offset, time offset, dynamic compression, and hyperbolic tangent distortion, and the neural network coverage rate criteria include four types, namely, neuron coverage rate, neuron strong activation coverage rate, Top-k neuron coverage rate, and multi-section neuron coverage rate. The validity and diversity of speech data are guided and evaluated from multiple dimensions, and the generated data are selected. The mutation method of the present invention does not rely on complex models, does not have a complex training process, has diverse mutation strategies, and has diverse mutation effect evaluation methods, which ensures the sufficiency and validity of the mutation data. The generated data can not only test the robustness of the model, but also support the testing of the conventional performance indicators of the model. In addition, the data mutation module and the mutation data selection module of the present invention are automated processes, which ensure the efficiency of speech data generation.

[0101] The contents not described in detail in the specification of the present invention belong to the prior art known to the professional and technical personnel in this field.

Claims

1. A method for generating test data for a speech intelligence model guided by multiple coverage rates, characterized in that include: Input data, including the voice intelligence model to be tested and the seed test data set; Analyze the activation state of neurons in the neural network under a given seed test data set under different neural network coverage indicators of the voice intelligence model to be tested, and obtain the coverage under different neural network coverage indicators; Neural network coverage indicators include: neuron coverage, Top-k neuron coverage, neuron strong activation coverage, and multi-node neuron coverage; Randomly select one or more mutation strategies to mutate the seed test data; the mutation strategies include random silence, random gain, Gaussian noise, speech rate change, pitch shift, time shift, dynamic compression, and hyperbolic tangent distortion; Define strong mutation number constraints, preliminarily filter mutation data, use the filtered mutation data to calculate the four neural network coverage rates of the speech intelligence model to be tested, select mutation data that improves the coverage rate of the model to be tested compared to before mutation, and form mutation data sets corresponding to different neural network coverage criteria; Output the coverage calculation results of the speech intelligence model to be tested on the seed test dataset and the selected variant test dataset.

2. The method for generating speech intelligence model test data guided by multiple coverage rates according to claim 1, characterized in that: The set of all neurons in the speech intelligence model to be tested is N = {n1, n2, ..., n m }; All seed test datasets are Among them, x i represents audio sample i.

3. The method for generating speech intelligence model test data guided by multiple coverage rates according to claim 2, characterized in that: The activation state of the neurons in the neural network under a given seed test data set is: Let O j (x i ) indicates that given a test input x i The number of neurons in the speech intelligence model to be tested is n j The output value, A j (x i )∈{0,1} means the input is x i When neuron n j The activation state of , where: When O j (x i )>O min When A j (x i ) = 1, indicating that the neuron is activated; When O j (x i )≤O min When A j (x i )=0, indicating that the neuron is not activated.

4. The method for generating speech intelligence model test data guided by multiple coverage rates according to claim 2, characterized in that: The neural network coverage criteria for neuron coverage, neuron strong activation coverage, Top-k neuron coverage, and multi-section neuron coverage are as follows: Neuron coverage: The neuron coverage of a set of test inputs is defined as the ratio of the number of activated neurons in all test inputs to the total number of neurons in the model under test, that is, the neuron coverage is the ratio of the number of activated neurons to the total number of neurons; the definition is as follows: Neuron strong activation coverage: When training the model, each neuron n j The upper boundary value can be calculated based on the training set data analysis If the output value of a neuron is higher than its upper boundary value, the neuron is considered to be strongly activated; the strong activation coverage is the ratio of the number of strongly activated neurons to the total number of neurons; it is defined as follows: Top-k neuron coverage: Assuming a neural network has l layers, for the rth layer, use top k (x i ,r) means when the input is x i When , the neurons corresponding to the top K values ​​with the largest neuron output value in the rth layer; under all test inputs, the Top-k neuron coverage is the proportion of the top K most active neurons in all layers of the neural network to the total number of neurons after taking the union; it is defined as follows: Multi-section neuron coverage: Determine the activation range of each neuron based on the training data and divide this range into K equal segments. During testing, statistically test the input x i ∈X is the output value O of the j∈Nth neuron j (x i ) falls in the K segments of the jth neuron as a proportion of the total number of neuron segments; it is defined as follows: Among them, Fall k (O j (x i )) represents the test input x i The output value O of the jth neuron j (x i ) falls in the kth segment of the jth neuron.

5. The method for generating speech intelligence model test data guided by multiple coverage rates according to claim 2, characterized in that: The number of mutations m for the seed test data is user-defined, defining the seed data x i The result after the jth mutation is The random silence, random gain, Gaussian noise, speech rate change, pitch shift, time shift, dynamic compression, and hyperbolic tangent distortion mutation strategies are: Random silence: Add silence to a randomly selected part of the audio signal to simulate a brief signal loss or blockage; the audio data after adding random silence is represented as t represents time; Random gain: Multiply the audio signal by a random amplitude factor to adjust the volume by increasing or decreasing the overall amplitude of the audio signal. The loudness of the audio signal will change, but the pitch and rhythm remain unchanged. The audio data after random gain is represented as x′ ij =α·x i , α>1 means increasing the volume, α<1 means decreasing the volume; Gaussian Noise: Adds Gaussian noise to the audio signal The audio data after adding Gaussian noise is represented as x′ ij =x i +g; Speech rate variation: changing the speed or duration of an audio signal without changing the pitch; The audio data after the speech rate changes is expressed as Where α is the time stretching factor, α>1 means the audio is stretched and the speed is slowed down; α<1 means the audio is compressed and the speed is fast; Pitch shift: Raise or lower the pitch of an audio signal without changing the audio rhythm or duration. The audio data after pitch shift is represented as Indicates the shift operation, α indicates the number of semitones of pitch shift, α>0 indicates raising the pitch, and α<0 indicates lowering the pitch; Time Shift: Shift the audio signal forward or backward; the audio data after applying the shift is represented as Δt represents the amount of time for translation, positive values ​​represent forward translation, negative values ​​represent backward translation, and L represents the total duration of the input audio signal; Dynamic compression: limits the maximum amplitude of the audio signal to prevent the signal from exceeding a preset threshold. The audio data after dynamic compression is represented as Where T represents the amplitude threshold, sgn(x i ) represents the signal x i The symbolic function of Hyperbolic tangent distortion: By applying the input audio signal x i Apply the hyperbolic tangent function to add distortion effect; the audio data after applying the hyperbolic tangent distortion is represented as x′ ij =tanh(k·x i ), k represents an adjustable gain factor, which is used to control the intensity of the distortion effect.

6. The method for generating speech intelligence model test data guided by multiple coverage rates according to claim 2, characterized in that: Define strong mutation number constraints to preliminarily filter mutation data through strong mutation number constraints, specifically: Calculate the changes in the data values ​​before and after the mutation Seed data x i The result after the jth mutation is Calculate the mean amplitude of the data changes before and after the mutation The number of values ​​that have changed statistically and exceed the mean amplitude (i.e., strong variation) Determine x′ ij Whether the number of strong mutations does not exceed 50% of the data dimension, that is, If the constraint condition is satisfied, the next step is selected; otherwise, the mutation strategy is adopted again to mutate the seed test data.

7. The method for generating speech intelligence model test data guided by multiple coverage rates according to claim 1, characterized in that: Use the filtered variant data to calculate the coverage of the four neural networks of the voice intelligence model to be tested, select the variant data that improves the coverage of the model to be tested compared to before the mutation, and form a variant data set corresponding to different neural network coverage criteria, specifically: The mutation data x′ that satisfies the strong mutation number constraint ij As the input of the speech intelligence model to be tested, the coverage of the model to be tested is calculated according to each neural network coverage criterion. If the coverage is improved compared with the previous result, the mutated data is retained, otherwise it is discarded and mutated again until all samples reach the set number of mutations of the seed test data; a set of corresponding mutated data sets is formed for each coverage criterion.

8. A system for generating test data for speech intelligence models guided by multiple coverage rates, characterized in that include: Data input module, neuron activation state analysis module, speech data mutation module, mutation data selection module, and result visualization module; specifically: The data input module stores and inputs the speech intelligence model to be tested and the seed test data set to other modules; The neuron activation state analysis module analyzes the activation state of each neuron in the neural network under a given seed test data set, and outputs the coverage under different neural network coverage indicators; the neural network coverage criteria include neuron coverage, Top-k neuron coverage, neuron strong activation coverage, and multi-section neuron coverage. The speech data mutation module automatically mutates the seed test data set. The mutation strategies include random silence, random gain, Gaussian noise, speech rate change, pitch shift, time shift, dynamic compression, and hyperbolic tangent distortion. The variant data selection module defines strong variant number constraints, automatically filters variant data that do not meet the constraints, and automatically calculates the neural network coverage of the speech intelligence model to be tested using the filtered variant data according to each neural network coverage criterion, and selects and retains variant data that improves model coverage; The result visualization module presents the coverage calculation results of the speech intelligence model to be tested on the seed test dataset and the filtered variant test dataset in the form of charts, and displays the statistical results of the selected variant data.

Citation Information

Patent Citations

  • Evaluation system for deep learning model

    CN117493140A

  • Automatic training generation method and system for smart home interaction test case

    CN117932348A

  • Intelligent system test data generation method based on uncertainty

    CN113762335A

  • Neural network model test method based on weighted neuron coverage rate

    CN114492740A

  • Loop code test data generation method based on deep learning fuzzy test

    CN116680164A

Cited By

  • Off-line voice control intelligent loudspeaker test system and method

    CN121240025A