A method and device for generating a simulation test set, electronic equipment and a storage medium

By automating the processing and annotation of audio data in new energy vehicles to generate simulation test sets, the problem of low testing efficiency in existing technologies is solved, and rapid simulation testing of multiple vehicle models and scenarios is realized.

CN119763613BActive Publication Date: 2026-03-27BEIJING CO WHEELS TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing voice testing methods for new energy vehicles suffer from low testing efficiency and high manpower and material costs, especially in simulation testing across multiple models and scenarios, which is complex and time-consuming.

Method used

By playing a preset audio dataset in a preset vehicle scenario, a second audio dataset is automatically collected and labeled to generate a simulation test set, which is then processed automatically using a speech test set automatic conversion device.

Benefits of technology

It improves the efficiency of generating simulation test sets, reduces the time and effort required by staff, and enables rapid testing of multiple vehicle models and scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119763613B_ABST
    Figure CN119763613B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a simulation test set generation method and device, electronic equipment and storage medium; the method comprises: a voice test set automatic conversion device plays a preset first audio data set in a vehicle in a preset vehicle scene; from the vehicle, a second audio data set collected for the first audio data set is obtained, and log information generated by the vehicle in the process of playing the first audio data set; from the log information, the target recognition result of the vehicle for the second audio data set is obtained; according to the target recognition result, the second audio data set is labeled, and according to the labeled result and the second audio data set, a simulation test set is generated. Since the simulation test set is automatically generated throughout the process, compared with the existing technology through the staff to play, screen and label, the embodiments of the present application can improve the efficiency of simulation test set generation, reduce the time and energy input cost of the staff.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to a simulation test set generation method, a simulation test set generation device, an electronic device and a computer readable storage medium. BACKGROUND

[0002] In new energy vehicles, with the increasing popularity of intelligent voice technology, it is particularly important to ensure the performance of the voice system. For the performance of the voice system, large-scale and repetitive tests are often needed to ensure the performance. The existing test mode generally includes manual voice wake-up, speaking or playing pre-recorded audio in the car through a sound system, and real-time monitoring of errors in voice wake-up and recognition. The errors are selected, extracted, summarized and counted from a large number of logs to become test results. In the market environment of multiple vehicle models, long-term continuous testing has certain limitations.

[0003] Therefore, in the environment of rapid iteration of voice products, a solution is needed to quickly complete multi-vehicle and multi-scene testing. Currently, simulation testing is usually used. For the simulation of voice test set production, complex manual recording, playing and data extraction are often needed, and complex manual annotation of data and audio and table alignment operations are needed, which consumes a lot of time and effort. SUMMARY

[0004] In view of the above problems, a simulation test set generation method, a simulation test set generation device, an electronic device and a computer readable storage medium are provided to overcome the above problems or at least partially solve the above problems.

[0005] The embodiment of the present application provides a simulation test set generation method applied to a voice test set automatic conversion device, and the method comprises the following steps:

[0006] The voice test set automatic conversion device plays a preset first audio data set in a vehicle in a preset vehicle scene;

[0007] The second audio data set collected for the first audio data set and log information generated by the vehicle in the process of playing the first audio data set are obtained from the vehicle;

[0008] The target recognition result of the vehicle for the second audio data set is obtained from the log information;

[0009] The second audio data set is annotated according to the target recognition result, and a simulation test set is generated according to the annotated result and the second audio data set.

[0010] Optionally, the method further comprises:

[0011] obtaining preset first audio data and scene configuration information set for each first audio data;

[0012] storing the scene configuration information and the first audio data into a preset first data table to obtain the first audio data set;

[0013] playing the preset first audio data set in a vehicle in a preset vehicle scene, comprising:

[0014] constructing the preset vehicle scene in the vehicle according to the scene configuration information in the first data table, and playing first audio data in the first audio data set in the vehicle.

[0015] Optionally, the storing the scene configuration information and the first audio data into a preset first data table comprises:

[0016] storing the first audio data and the scene configuration information into the preset first data table according to a preset format requirement.

[0017] Optionally, the scene configuration information comprises vehicle control parameters and playing control parameters, and the constructing the preset vehicle scene in the vehicle according to the scene configuration information in the first data table, and playing first audio data in the first audio data set in the vehicle comprises:

[0018] controlling the vehicle according to the vehicle control parameters to construct the preset vehicle scene in the vehicle;

[0019] playing first audio data in the first audio data set in the vehicle according to the playing control parameters.

[0020] Optionally, the generating a simulation test set according to the labeling result and the second audio data set comprises:

[0021] determining, from the first data table, scene configuration information corresponding to first audio data corresponding to each second audio data in the second audio data set;

[0022] generating the simulation test set according to the labeling result, the scene configuration information corresponding to each second audio data in the second audio data set, and the second audio data set.

[0023] Optionally, the generating a simulation test set according to the labeling result and the second audio data set comprises:

[0024] According to the target recognition result, a time point of a target sentence in each second audio data in the second audio data set is determined, and the time point of the target sentence in each second audio data is taken as a result of labeling;

[0025] According to the result of labeling and the second audio data set, a simulation test set is generated.

[0026] Optionally, the method further comprises:

[0027] cleaning up a target storage space; the target storage space is used for storing the log information and / or the second audio data set.

[0028] Optionally, the method further comprises:

[0029] performing simulation test using the second audio data set in the simulation test set to obtain a first simulation test result;

[0030] According to the first simulation test result and the result of labeling in the simulation test set, the simulation test is evaluated.

[0031] Embodiments of the present application also provide a simulation test set generation device applied to a voice test set automatic conversion device, the device comprising:

[0032] a playing module configured to play a preset first audio data set in a vehicle in a preset vehicle scene by the voice test set automatic conversion device;

[0033] a collecting module configured to acquire a second audio data set collected for the first audio data set and log information generated by the vehicle in a process of playing the first audio data set from the vehicle;

[0034] an extracting module configured to acquire a target recognition result of the vehicle for the second audio data set from the log information;

[0035] a generating module configured to label the second audio data set according to the target recognition result, and generate a simulation test set according to a result of labeling and the second audio data set.

[0036] Optionally, the device further comprises:

[0037] a scene information acquiring module configured to acquire preset first audio data and scene configuration information set for each first audio data; store the scene configuration information and the first audio data into a preset first data table to obtain the first audio data set;

[0038] The playing module is configured to construct the preset vehicle scene in the vehicle according to the scene configuration information in the first data table, and play the first audio data in the first audio data set in the vehicle.

[0039] Optionally, the scene information acquisition module is configured to store the first audio data and the scene configuration information in a preset first data table according to preset format requirements.

[0040] Optionally, the scene configuration information comprises vehicle control parameters and playing control parameters, and the playing module is configured to control the vehicle according to the vehicle control parameters to construct the preset vehicle scene in the vehicle, and play the first audio data in the first audio data set in the vehicle according to the playing control parameters.

[0041] Optionally, the generating module is configured to determine, from the first data table, the scene configuration information corresponding to the first audio data corresponding to each second audio data in the second audio data set, and generate the simulation test set according to the labeled result, the scene configuration information corresponding to each second audio data in the second audio data set, and the second audio data set.

[0042] Optionally, the generating module is configured to determine, according to the target recognition result, the time point of the target word or phrase in each second audio data in the second audio data set, and take the time point of the target word or phrase in each second audio data as the labeled result, and generate the simulation test set according to the labeled result and the second audio data set.

[0043] Optionally, the apparatus further comprises:

[0044] The cleaning module is configured to clean a target storage space, and the target storage space is configured to store the log information and / or the second audio data set.

[0045] Optionally, the apparatus further comprises:

[0046] The evaluation module is configured to perform simulation testing using the second audio data set in the simulation test set to obtain a first simulation testing result, and evaluate the simulation testing according to the first simulation testing result and the labeled result in the simulation test set.

[0047] Embodiments of the present application also provide an electronic device comprising a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program is executed by the processor to implement the simulation test set generation method as above.

[0048] The embodiment of the present application also provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the method for generating the simulation test set.

[0049] The embodiment of the present application has the following advantages:

[0050] In the embodiment of the present application, the voice test set automatic conversion device plays a preset first audio data set in a vehicle in a preset vehicle scene, acquires a second audio data set collected for the first audio data set and log information generated by the vehicle in the process of playing the first audio data set from the vehicle, acquires a target recognition result of the vehicle for the second audio data set from the log information, labels the second audio data set according to the target recognition result, and generates a simulation test set according to the labeled result and the second audio data set. Since the simulation test set is automatically generated throughout the process, compared with the playing, screening and labeling mode by the staff in the prior art, the embodiment of the present application can improve the efficiency of generating the simulation test set and reduce the time and energy input cost of the staff. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the description of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0052] Figure 1 is a step flow chart of a simulation test set generation method of the embodiment of the present application;

[0053] Figure 2 is a step flow chart of another simulation test set generation method of the embodiment of the present application;

[0054] Figure 3a is a scene schematic diagram of the embodiment of the present application;

[0055] Figure 3b is another scene schematic diagram of the embodiment of the present application;

[0056] Figure 4a is a step flow chart of a simulation test set generation method of the embodiment of the present application;

[0057] Figure 4b is a step flow chart of another simulation test set generation method of the embodiment of the present application;

[0058] Figure 5 is a structure block diagram of a simulation test set generation device of the embodiment of the present application. DETAILED DESCRIPTION

[0059] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0060] The simulation test is a method for evaluating performance by simulating real environment or conditions; in actual application, a simulation test set needs to be obtained first, which can be used to simulate real environment or conditions. Specifically, for the scene of vehicle voice wake-up and voice recognition, a worker needs to drive the vehicle first so that the vehicle is in different vehicle scenes (for example, high-speed driving scene, low-speed driving scene, etc.); then, the worker can play preset audio data in the vehicle.

[0061] Then, the worker can obtain the audio data recorded by the vehicle for the played preset audio data from the vehicle, and screen out the corresponding information from the log information of the vehicle, and then generate a simulation test set; in this process, the worker needs to spend a lot of time and effort to play, screen, label, etc. In order to reduce the time consumed by the worker to produce the simulation test set, the embodiment of the present application provides a simulation test set generation method, which automatically generates a simulation test set through automatic playing, automatic screening and automatic labeling.

[0062] Specifically, reference can be made to Figure 1 , Figure 1 The step flowchart of the simulation test set generation method of the embodiment of the present application is shown, which can be applied to a voice test set automatic conversion device.

[0063] As shown in Figure 1 , the method can include the following steps:

[0064] Step 101, the voice test set automatic conversion device plays a preset first audio data set in a vehicle in a preset vehicle scene.

[0065] The first audio data set can include a plurality of audio data, and the plurality of audio data can include different contents; for example, the plurality of audio data can include audio data of voice wake-up of different age groups, different genders and different speech speeds, and audio data of recognition sentences of different language types, and the embodiment of the present application does not limit this.

[0066] The preset vehicle scene can refer to a scene in which the vehicle plays the audio data in the first audio data set, for example, a high-speed driving scene or a low-speed driving scene.

[0067] In actual applications, the vehicle can be controlled to be in the preset vehicle scene first, and then the vehicle or the playing device deployed in the vehicle can be controlled to play the audio data in the first audio data set.

[0068] For example, the method for generating the simulation test set mentioned in the present application can be implemented by a voice test set automatic conversion device. Specifically, when the preset first audio data set is played, the voice test set automatic conversion device can send the audio data in the first audio data set to the vehicle or the playing device deployed in the vehicle, so that the vehicle or the playing device deployed in the vehicle plays the audio data in the first audio data set.

[0069] Step 102: obtaining, from the vehicle, a second audio data set collected for the first audio data set and log information generated by the vehicle in the process of playing the first audio data set.

[0070] When the audio data in the first audio data set is played, the audio recording device (for example, a microphone; the microphone can include one or more, and the one or more microphones can be arranged in the vehicle according to actual conditions; when the microphone includes multiple, the multiple microphones can be distributed at different positions in the vehicle) in the vehicle can record the audio data, thereby obtaining the second audio data set.

[0071] When the simulation test set is generated, the voice test set automatic conversion device can automatically read, from the vehicle, the second audio data set collected by the vehicle for the played first audio data set; the second audio data set can also include multiple audio data, and the multiple audio data are obtained by recording the audio data in the played first audio data set.

[0072] It should be noted that the difference between the audio data in the first audio data set and the audio data in the second audio data set is that the audio data in the second audio data set includes not only the sound of the audio data in the first audio data set, but also the sound of the preset vehicle scene, for example, the sound generated by the vehicle driving or the sound of other people speaking, and the present application does not limit this.

[0073] In some possible embodiments, after receiving the audio data in the first audio data set, the vehicle can play the audio data; while playing the audio data, the infotainment system in the vehicle can collect the audio data, and identify the collected data (i.e., the second audio data) to generate corresponding log information, which can include the collection of the played audio data, and the recognition result corresponding to the identification of the collected second audio data, for example, the text content obtained after voice recognition, or the time corresponding to each word, phrase or sentence in the audio data after identifying the audio data in the second audio data set.

[0074] For example, when collecting the audio data in the played first audio data set, the vehicle can also collect the sound in the environment; then, the vehicle can identify the collected data, and generate corresponding log information based on the identification result.

[0075] In actual applications, when generating the simulation test set, the voice test set automatic conversion device can also automatically read the log information generated by the vehicle in the process of playing the first audio data set from the vehicle.

[0076] In some possible embodiments, the vehicle for outputting the log information and the second audio data set can be a vehicle with certain voice recognition capability, which can accurately identify and analyze the audio data.

[0077] Step 103, obtaining the target recognition result of the vehicle for the second audio data set from the log information.

[0078] After obtaining the log information, the target recognition result of the vehicle for the audio data in the second audio data set can be obtained; the target recognition result can be used to represent the information of each audio data in the second audio data set, for example, whether the audio data is a text sentence including a wake-up word or a text sentence not including a wake-up word; if it is a sentence including a wake-up word, the target recognition result can include the time period of the wake-up word in the audio data; if it is a text sentence not including a wake-up word, the target recognition result can also include the specific text content of the text sentence, and / or the time period of the word, phrase or sentence in the text content in the audio data (for example, the time period of a word in the audio data can be represented by the start time and the end time of the word in the audio data).

[0079] Step 104, labeling the second audio data set according to the target recognition result, and generating a simulation test set according to the labeling result and the second audio data set.

[0080] After obtaining the target recognition result, the voice test set automatic conversion device can automatically label the audio data in the second audio data set, so that the corresponding content of each audio data in the second audio data set and the time when the corresponding content is in the second audio data can be determined when subsequent simulation testing is performed.

[0081] After labeling the audio data in the second audio data set, the voice test set automatic conversion device can generate a simulation test set according to the labeling result and the second audio data set. The simulation test set can include the labeling result and the second audio data set. The second audio data set can include specific audio data files or storage addresses and names of specific audio data, which are not limited in the embodiments of the present application.

[0082] In some possible embodiments, after obtaining the simulation test set, it can be applied to different types of vehicles or different vehicle simulators to perform simulation testing on different types of vehicles. Since the simulation test set is automatically generated by the voice test set automatic conversion device, compared with the existing technology which plays, filters and labels by workers, the embodiments of the present application can improve the efficiency of generating the simulation test set and reduce the cost of time and energy input of workers.

[0083] For example, when performing simulation testing, the audio data in the second audio data set can be played to the vehicle or the vehicle simulator. Then, the corresponding recognition result can be obtained from the vehicle or the vehicle simulator, and then the performance of the vehicle or the vehicle simulator can be determined based on the recognition result and the labeling result in the simulation test set. For example, if the recognition result matches the labeling result in the simulation test set, it means that the performance of the vehicle or the vehicle simulator is passed. Otherwise, it means that the performance of the vehicle or the vehicle simulator needs to be adjusted, which is not limited in the embodiments of the present application.

[0084] In the embodiments of the present application, in the vehicle in the preset vehicle scene, the preset first audio data set is played. The second audio data set collected by the vehicle for the first audio data set and the log information generated by the vehicle in the process of playing the first audio data set are obtained from the vehicle. The target recognition result of the vehicle for the second audio data set is obtained from the log information. The second audio data set is labeled according to the target recognition result, and a simulation test set is generated according to the labeling result and the second audio data set. Since the simulation test set is automatically generated, compared with the existing technology which plays, filters and labels by workers, the embodiments of the present application can improve the efficiency of generating the simulation test set and reduce the cost of time and energy input of workers.

[0085] Reference Figure 2Fig. 2 shows a flow chart of steps of another method for generating a simulation test set according to an embodiment of the present application, which can include the following steps:

[0086] In step 201, preset first audio data and scenario configuration information set for each first audio data are obtained.

[0087] In actual applications, a plurality of first audio data can be recorded in advance, which can include audio data of voice wake-up of different age groups, different genders and different speech speeds, and audio data of recognition sentences of different language types, and the present application is not limited in this regard.

[0088] In addition, corresponding scenario configuration information can be set for different first audio data respectively, which can be used to configure a vehicle scenario in which the first audio data is played.

[0089] In some possible embodiments, the voice test set automatic conversion device can first obtain a plurality of preset first audio data and scenario configuration information set for each first audio data.

[0090] In step 202, the scenario configuration information and the first audio data are stored in a preset first data table to obtain a first audio data set.

[0091] After obtaining a plurality of preset first audio data and scenario configuration information set for each first audio data, the voice test set automatic conversion device can store the scenario configuration information and the first audio data in a preset first data table according to a corresponding relationship, thereby obtaining a first audio data set.

[0092] In some possible embodiments, the scenario configuration information and the first audio data can be sequentially stored in the preset first data table according to a playing order, so that the voice test set automatic conversion device can play the first audio data in the first audio data set according to the order in the future, and the present application is not limited in this regard.

[0093] In an embodiment of the present application, the above step 202 can be implemented in the following manner:

[0094] The first audio data and the scenario configuration information are stored in the preset first data table according to a preset format requirement.

[0095] In some possible embodiments, in order to ensure that the audio data can be correctly played and processed, the first audio data and the scene configuration information can be stored in a preset first data table according to preset format requirements; specifically, the first audio data of 16000Hz, single channel and wav format and the scene configuration information corresponding to the first audio data can be stored in the preset first data table, and the format requirements of the embodiments of the present application are not limited, which can be set according to actual conditions.

[0096] Step 203, constructing a preset vehicle scene in the vehicle according to the scene configuration information in the first data table, and playing the first audio data in the first audio data set in the vehicle.

[0097] In some possible embodiments, after obtaining the scene configuration information in the first data table, the voice test set automatic conversion device can first construct a preset vehicle scene in the vehicle according to the scene configuration information.

[0098] Then, the voice test set automatic conversion device can play the first audio data in the first audio data set in the vehicle in the preset vehicle scene.

[0099] In an embodiment of the present application, the scene configuration information can include vehicle control parameters and playing control parameters; wherein the vehicle control parameters can include parameters for controlling the vehicle to construct a preset vehicle scene, for example: vehicle speed, etc.; the playing control parameters can include parameters for controlling the playing to construct a preset vehicle scene, for example: the position of the playing and the volume of the playing, etc., which are not limited by the embodiments of the present application. In actual application, step 203 can be realized through the following sub-steps:

[0100] Sub-step 11, controlling the vehicle according to the vehicle control parameters to construct a preset vehicle scene in the vehicle.

[0101] After obtaining the vehicle control parameters, the voice test set automatic conversion device can first control the vehicle to make the vehicle in a preset vehicle scene.

[0102] For example, when the preset vehicle scene is a high-speed driving scene, the vehicle can be controlled to drive at a speed exceeding a preset speed value (for example: 80km / h); when the vehicle speed exceeds the preset speed value, it can be determined that the vehicle is in a high-speed driving scene, and at this time, the first audio data in the first audio data set can be played according to the playing control parameters.

[0103] In some possible embodiments, in-vehicle conversations and other scenes can also be added; for example, when the first audio data set is played in the vehicle, audio data of human-to-human conversations can be played through other audio playing devices to enrich the simulation test set.

[0104] Sub-step 12, playing the first audio data in the first audio data set in the vehicle according to the playing control parameter.

[0105] After the vehicle is caused to be in the preset vehicle scene, the first audio data in the first audio data set can be played in the vehicle according to the playing control parameter; for example, the first audio data played at the corresponding position, or the first audio data played at the set volume, which is not limited by the embodiments of the present application.

[0106] Step 204, obtaining, from the vehicle, a second audio data set collected by the vehicle for the first audio data set, and log information generated by the vehicle in the process of playing the first audio data set.

[0107] In the process of playing the audio data in the first audio data set, the vehicle can record the audio data, thereby obtaining the second audio data set.

[0108] In the process of generating the simulation test set, the voice test set automatic conversion device can automatically read, from the vehicle, the second audio data set collected by the vehicle for the played first audio data set.

[0109] After receiving the audio data in the first audio data set, the vehicle can play the audio data; in the process of playing the audio data, the vehicle can collect the audio data, and identify the collected data (i.e., the audio data in the second audio data set) to generate corresponding log information.

[0110] In actual application, in the process of generating the simulation test set, the voice test set automatic conversion device can also automatically read, from the vehicle, the log information generated by the vehicle in the process of playing the first audio data set.

[0111] Step 205, obtaining, from the log information, a target recognition result of the vehicle for the second audio data set.

[0112] After obtaining the log information, the voice test set automatic conversion device can obtain the target recognition result of the vehicle for the audio data in the second audio data set.

[0113] Step 206, determining, according to the target recognition result, a time point of the target sentence in each second audio data in the second audio data set, and taking the time point of the target sentence in each second audio data as an annotation result.

[0114] In some feasible embodiments, after obtaining the target recognition result, the voice test set automatic conversion device can determine the time point of the target sentence in each second audio data in the second audio data set based on the time information in the target recognition result, so as to determine the position of the target sentence in the second audio data based on the time point.

[0115] In actual application, the time point corresponding to the target sentence in each second audio data can be taken as a result of labeling the target sentence in the second audio data. Based on the labeling result, the time at which the target sentence in the second audio data is located can be determined. Further, in the simulation test, the time at which the same target sentence is located in the result output by the simulation test can be used to evaluate the simulation test.

[0116] The target sentence can be a wake-up word or a non-wake-up word, and the embodiments of the present application do not limit the target sentence.

[0117] In step 207, a simulation test set is generated based on the labeling result and the second audio data set.

[0118] After obtaining the time point of the target sentence in each second audio data, the voice test set automatic conversion device can generate a simulation test set based on the time point corresponding to the target sentence in each second audio data and the second audio data set.

[0119] In an embodiment of the present application, step 207 can be implemented in the following manner:

[0120] From the first data table, the scene configuration information corresponding to the first audio data corresponding to each second audio data in the second audio data set is determined. The simulation test set is generated based on the labeling result, the scene configuration information corresponding to each second audio data in the second audio data set, and the second audio data set.

[0121] In some feasible embodiments, when the simulation test set is generated, the voice test set automatic conversion device can also read the scene configuration information corresponding to the first audio data recorded by each second audio data from the first data table, and take the scene configuration information as the scene configuration information of the second audio data.

[0122] After labeling the second audio data and determining the scene configuration information corresponding to each second audio data, the voice test set automatic conversion device can generate a simulation test set based on the labeling result, the scene configuration information corresponding to each second audio data in the second audio data set, and the second audio data set.

[0123] Specifically, the labeling result, the scene configuration information corresponding to each second audio data in the second audio data set, and the name or address of the second audio data in the second audio data set can be written into the second data table according to the corresponding relationship, and the second data table can be taken as the simulation test set.

[0124] Based on the name or address of the second audio data in the simulation test set, the corresponding second audio data can be obtained; based on the obtained second audio data and the labeled results and scene configuration information in the simulation test set, simulation tests in different vehicle scenes can be performed.

[0125] In some possible embodiments, when performing the simulation test, the scene to which the simulation test is directed can be determined based on the scene configuration information; then, the second audio data in the second audio data set can be played, and the simulation test can be evaluated based on the results output by the simulation test with the labeled results in the simulation test set as a benchmark, so as to determine the performance of the real vehicle or the vehicle simulator that performs the simulation test.

[0126] In an embodiment of the present application, the method can further include the following steps:

[0127] The target storage space is cleaned; the target storage space is used to store the log information and / or the second audio data set.

[0128] In some possible embodiments, after the log information and / or the second audio data set are obtained, the log information and / or the second audio data set can be stored in the target storage space. Before storing the log information and / or the second audio data set, in order to avoid interference of other information, the target storage space can be cleaned first, and then the log information and / or the second audio data set obtained by currently processing the first audio data set can be stored in the target storage space.

[0129] In actual application, after the simulation data set is output, the target storage space can be cleaned to process the next audio data set for transcription to obtain another simulation data set, and the embodiment of the present application does not limit this.

[0130] In an embodiment of the present application, the method can further include the following steps:

[0131] The simulation test is performed using the second audio data set in the simulation test set to obtain a first simulation test result; and the simulation test is evaluated according to the first simulation test result and the labeled results in the simulation test set.

[0132] In some possible embodiments, the simulation test can be performed using the second audio data set in the simulation test set; for example, the second audio data set can be played in a real vehicle. After the simulation test is performed using the simulation test set, a first simulation test result can be obtained, which can represent the recognition of the audio data in the played second audio data set in the simulation test, for example, whether a wake-up word is woken up or specific content recognized.

[0133] After obtaining the first simulation test result, the simulation test can be evaluated according to the identified result and the first simulation test result in the simulation test set. Specifically, if the identified result matches the first simulation test result or the matching degree reaches a preset value, it can be indicated that the simulation test passes, that is, the real vehicle or vehicle simulator test passes. For example: whether the same text content corresponds to the same wake-up word or non-wake-up word at the same time.

[0134] In some possible embodiments, the simulation test can also be used to evaluate the pros and cons of the simulation test set. For example, if the first simulation test result deviates greatly from the identified result, it can also indicate that the simulation test set may have problems. In order to further verify the simulation test set, multiple simulation tests can be performed, and the pros and cons of the simulation test set can be evaluated based on the results of the multiple simulation tests.

[0135] In the embodiment of the application, the preset first audio data and the scene configuration information set for each first audio data are obtained; the scene configuration information and the first audio data are stored in the preset first data table to obtain a first audio data set; according to the scene configuration information in the first data table, a preset vehicle scene is constructed in the vehicle, and the first audio data in the first audio data set is played in the vehicle; from the vehicle, a second audio data set collected for the first audio data set and log information generated by the vehicle in the process of playing the first audio data set are obtained; from the log information, a target recognition result of the vehicle for the second audio data set is obtained; according to the target recognition result, the time point of the target sentence in each second audio data in the second audio data set is determined, and the time point of the target sentence in each second audio data is taken as the labeled result; and according to the labeled result and the second audio data set, a simulation test set is generated. Since the simulation test set is automatically generated throughout, compared with the existing technology which plays, filters and labels by the staff, the embodiment of the application can improve the efficiency of generating the simulation test set, and reduce the time and energy input cost of the staff.

[0136] In the following, the above method is further described through a specific example:

[0137] I. Automatic conversion preparation:

[0138] 1.1. According to the requirements of the simulation test set, a plurality of first audio data required in advance is prepared, and the plurality of first audio data contains audio of voice wake-up of different age groups, different genders and different speech speeds, and recognition sentences of different language types.

[0139] 1.2, the first audio data which needs to be converted by the voice test set automatic conversion device is stored as 16000Hz, single channel and wav format audio according to the format requirements of the program, and the first data table required by the program to read is written.

[0140] In addition, scene configuration information can also be written in the first data table, thereby obtaining the first audio data set.

[0141] For example, as shown in Tables 1 and 2 below, the table headers of the first audio data sets of two different audio types are shown, and the specific data can be set according to actual needs:

[0142] Table 1:

[0143] Audio data storage address Speaker number Speaker gender Speaker age group Environment Location Speech rate

[0144] Table 2:

[0145] Audio data storage address Speaker gender Text content ***-**-** Female Third **-****-** Male Close windows

[0146] 1.3, the voice test set automatic conversion device is arranged on the real vehicle, and a sound card and a sound playing device such as a sound are connected to play the first audio data set. The voice test set automatic conversion device can also be connected with the vehicle to obtain log information and the second audio data set from the vehicle.

[0147] According to the required position, high-fidelity sound equipment can be installed at different seat positions, and when playing the first audio data set, the volume is controlled at 75DB to 80DB to simulate the volume closest to the human voice in the car. The maximum degree restores the high frequency and low frequency of human voice, avoids the problem of distortion or does not meet the real sound.

[0148] II. Automatic conversion operation:

[0149] 2.1, on the PC end deployed with the voice test set automatic conversion device, according to the requirements, the high-fidelity sound playing audio through sound card switching or direct connection is played, the sound card is called through different card numbers to play the first audio data set, and the sound is directly called to play, realizing the distinction on the playing audio.

[0150] As shown in Figure 3a and Figure 3b , the sound can be placed in the vehicle 30, and the connection between the PC end 301 deployed with the voice test set automatic conversion device and the vehicle end 302 is established; the PC end 301 can be connected with the vehicle end 302 through USB (Universal Serial Bus, Universal Serial Bus).

[0151] The PC end 301 can also be connected with the first sound 304 through the sound card 303, or directly connected with the second sound 305; the microphone 306 is deployed in the car machine end 302, which can be used to collect the sound output by the first sound 304 and / or the second sound 305, so as to obtain the second audio data.

[0152] 2.2, connect the PC end deployed with the voice test set automatic conversion device to the car machine end of the vehicle, start the program, confirm the sound card number of the sound card, and use the corresponding sound card number to call the playing device to play the first audio data in the first audio data set.

[0153] 2.3, the voice test set automatic conversion device first preoperates the car machine end, such as obtaining root permission, clearing the logs and audio in the target storage space, automatically restarting the voice system of the car machine end, ensuring that the test and transcribed audio can be normally performed, the audio on the disk has no redundant interference, and ensuring that the data collected this time is the test this time.

[0154] 2.4, the voice test set automatic conversion device reads the first data table, obtains the content and path of the first audio data set to be played, reads the second data table, and records the position of the second data table. After writing data in the second data table, remember the position of the existing data to avoid the subsequent written data covering the existing data.

[0155] 2.5, the voice test set automatic conversion device plays the prepared audio data in sequence according to the first data table, records the time before playing the audio data, such as: “today the weather is really good”, then collect the time at the moment before playing the word “today”.

[0156] 2.6, after playing each audio, obtain real-time log information, and capture the target recognition result required, including but not limited to wake-up word, wake-up sound area or sentence recognition result, etc.

[0157] For example, as shown in Table 3, multiple sentence recognition results can be obtained, and the audio data storage address and speaker gender of the first audio data corresponding to each sentence recognition result.

[0158] Table 3:

[0159] Audio data storage address Speaker gender Text content ***-**-** Female Reduce air conditioning air volume **-****-** Male Turn on reading light

[0160] 2.7, write the key information captured in the second data table in real time.

[0161] For example, as shown in Table 4 below, the table header of the target recognition result corresponding to the first audio data of the wake-up type, wherein the specific data can be set according to actual needs, and as shown in Table 5, the table header of the target recognition result corresponding to the non-wake-up word sentence, wherein the specific data can be set according to actual needs:

[0162] Table 4:

[0163]

[0164] Table 5:

[0165]

[0166] 2.8, after completing the test and dubbing of an audio data set, automatically export the current log information and the converted Sse (Speech Signal Enhancement, speech signal enhancement) -input name with a time stamp, multi-channel long original audio in PCM (Pulse Code Modulation, pulse code modulation) format.

[0167] For example, the second audio data can be named as "03-20_14-04-03_977_-Sse-input.pcm", and the corresponding log information can be named as "SS3_Android-00000700001890000-2024_03_20_14_03_22+0800.zip", which is not limited by the embodiments of the application.

[0168] 2.9, delete the historical log and audio of the target storage space again, automatically restart the voice system, prepare to read the new first data table, and perform new audio test and conversion.

[0169] 2.10, the above process is executed in a loop until all audio data conversion of the current vehicle placement position is completed to multi-channel original audio Sse-input.pcm.

[0170] Three, automatic conversion operation:

[0171] 3.1, after ending the part of the real vehicle, the audio is processed on the PC end where the voice test set automatic conversion device is deployed. The second data table of the real vehicle test is read, and the saved multi-channel original audio Sse-input.pcm is read.

[0172] 3.2, configure the scene configuration information to which each audio belongs; ensure that the output data after running processing is the multi-channel original audio Sse-input.pcm matching each scene and type recorded, such as:

[0173] wavepath = 'Cartest / W01 / Mandarin_kws / Quiet_1 / ' path

[0174] Type_str = 'Mandarin_kws'

[0175] Background_srt = 'Quiet'

[0176] folder_path = r"D:\testset"

[0177] xlsx_name = r"D:\TestSet\OUT_KWS_Quiet.xlsx"

[0178] Wherein, wavepath represents the audio data storage address; Type_str represents the test type; Background_srt represents the test scenario; folder_path represents the test table folder address; and xlsx_name represents the test table (i.e., the second data table) file name.

[0179] 3.3. Begin timing calculations on the multi-channel raw audio file Sse-input.pcm. Calculate the time each sentence appears in the long raw audio file by recording the time in the actual vehicle. For example, the first sentence, "**Student's" sentence, appears between the 10th and 12th seconds of the long audio file, and the second sentence, "**Student's" sentence, appears between the 23rd and 25th seconds of the long audio file.

[0180] 3.4 Complete the information matching in step 3.2 and the time calculation in step 3.3 to generate a simulation test set, as shown in Table 6. The table header of the simulation test set is shown, and the specific contents are set according to the actual situation.

[0181] Table 6:

[0182]

[0183] The final generated simulation test set may include the path of the converted long original audio uploaded to the cloud for simulation, the name of the ripped short audio, the start time of the target words, the end time of the target words, the duration, the test type, the test scenario, and the seat position in the vehicle, etc. This embodiment of the invention does not impose any restrictions on these.

[0184] 3.5 Finally, the automated voice test set conversion device can directly read the converted original audio Sse-input.pcm stored in the cloud to generate a simulation test set for simulation testing of different vehicle models and scenarios.

[0185] 3.6 After completing the first version of the simulation test, compare it with the test results of the real vehicle under the actual recording environment to verify the simulation test set, verify the quality of the simulation test set, and ensure that the test error is within a certain reliable range during the conversion process from real vehicle test to simulation test.

[0186] like Figure 4a and Figure 4b As shown:

[0187] First, data preparation and scenario setup can be carried out; then, the PC can play audio data in the vehicle to conduct real-vehicle testing and collect data.

[0188] Then, the automated voice test set conversion device deployed on the PC can perform real-vehicle testing based on the first audio dataset and output the results of the real-vehicle testing, which may include log information such as wake-up rate.

[0189] In addition, the automated voice test set conversion device can also annotate audio and generate simulation test sets based on the annotation results and the results of real vehicle tests.

[0190] After obtaining the simulation test set, it can be uploaded to the cloud server. Then, the simulation test set can be used to conduct simulation tests, and the results of the first simulation test can be compared with the results of the real vehicle test (i.e., the results marked in the simulation test set) to determine the quality of the simulation test set.

[0191] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0192] Reference Figure 5 This diagram illustrates a structural schematic of a simulation test set generation device according to an embodiment of the present invention, applied to an automated speech test set conversion device. The simulation test set generation device may include the following modules:

[0193] The playback module 501 is used to play a preset first audio dataset in a vehicle in a preset vehicle scenario by the automated voice test set conversion device.

[0194] The collection module 502 is used to obtain from the vehicle a second audio dataset collected for the first audio dataset, as well as log information generated by the vehicle during the playback of the first audio dataset;

[0195] The extraction module 503 is configured to acquire, from the log information, a target recognition result of the vehicle for the second audio data set.

[0196] The generation module 504 is configured to label the second audio data set according to the target recognition result, and generate a simulation test set according to the labeled result and the second audio data set.

[0197] In an optional embodiment of the present application, the device further comprises:

[0198] The scene information acquisition module is configured to acquire preset first audio data and scene configuration information set for each first audio data, store the scene configuration information and the first audio data in a preset first data table to obtain a first audio data set.

[0199] The playing module 501 is configured to construct a preset vehicle scene in the vehicle according to the scene configuration information in the first data table, and play the first audio data in the first audio data set in the vehicle.

[0200] In an optional embodiment of the present application, the scene information acquisition module is configured to store the first audio data and the scene configuration information in the preset first data table according to preset format requirements.

[0201] In an optional embodiment of the present application, the scene configuration information comprises vehicle control parameters and playing control parameters, and the playing module 501 is configured to control the vehicle according to the vehicle control parameters to construct the preset vehicle scene in the vehicle, and play the first audio data in the first audio data set in the vehicle according to the playing control parameters.

[0202] In an optional embodiment of the present application, the generation module 504 is configured to determine, from the first data table, scene configuration information corresponding to the first audio data corresponding to each second audio data in the second audio data set, and generate the simulation test set according to the labeled result, the scene configuration information corresponding to each second audio data in the second audio data set, and the second audio data set.

[0203] In an optional embodiment of the present application, the generation module 504 is configured to determine, from the target recognition result, a time point of a target sentence in each second audio data in the second audio data set, and take the target sentence in each second audio data and the time point of the target sentence as the labeled result, and generate the simulation test set according to the labeled result and the second audio data set.

[0204] In an optional embodiment of the present application, the device further comprises:

[0205] The cleaning module is configured to clean a target storage space, and the target storage space is configured to store the log information and / or the second audio data set.

[0206] In an optional embodiment of the present application, the device further comprises:

[0207] The evaluation module is configured to perform simulation testing using the second audio data set in the simulation testing set to obtain a first simulation testing result, and evaluate the simulation testing according to the first simulation testing result and the identified result in the simulation testing set.

[0208] In the vehicle in the preset vehicle scene, the preset first audio data set is played, the second audio data set collected for the first audio data set and the log information generated by the vehicle in the process of playing the first audio data set are obtained from the vehicle, the target recognition result of the vehicle for the second audio data set is obtained from the log information, the second audio data set is labeled according to the target recognition result, and the simulation testing set is generated according to the labeled result and the second audio data set. Since the simulation testing set is automatically generated throughout, compared with the existing technology, the simulation testing set is generated by the staff to play, filter and label, the embodiment of the present application can improve the efficiency of generating the simulation testing set, and reduce the time and energy input cost of the staff.

[0209] The embodiment of the present application also provides an electronic device, which comprises a processor, a memory, and a computer program stored on the memory and capable of running on the processor, and the computer program is executed by the processor to realize the simulation testing set generation method as above.

[0210] The embodiment of the present application also provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the simulation testing set generation method as above.

[0211] For the device embodiment, since it is basically similar to the method embodiment, it is described more simply, and the related parts are referred to the part of the method embodiment.

[0212] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same and similar parts of each embodiment are referred to each other.

[0213] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present application can adopt a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer usable program code.

[0214] The embodiments of the present application are described with reference to the flowchart illustrations and / or block diagrams of the methods, terminal devices (systems) and computer program products according to the embodiments of the present application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing terminal devices to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal devices, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0215] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal devices to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0216] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal devices, such that a series of operational steps are carried out on the computer or other programmable terminal devices to produce a computer implemented process so that the instructions executed on the computer or other programmable terminal devices provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0217] Although preferred embodiments of the present application have been described, those skilled in the art will be able to make additional modifications and variations to these embodiments without departing from the scope of the present application. Accordingly, the appended claims are intended to encompass all such modifications and variations as falling within the scope of the present application.

[0218] Finally, it is to be understood that the phraseology or terminology such as "first" and "second" etc. used herein is merely intended to differentiate one entity or operation from another entity or operation, without necessarily requiring or implying any actual such relationship or order between such entities or operations. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0219] The above provides a detailed introduction to the method for generating a simulation test set, the device for generating a simulation test set, the electronic device and the computer readable storage medium. The principles and implementation manners of the present application are described by using specific examples. The above example is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges can be changed. In conclusion, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A method for generating a simulation test set, characterized in that, The method, applied to an automated conversion device for speech test sets, includes: The automated voice test set conversion device plays a preset first audio dataset in a vehicle in a preset vehicle scenario. From the vehicle, obtain a second audio dataset collected for the first audio dataset, and log information generated by the vehicle during the playback of the first audio dataset; From the log information, obtain the target recognition result of the vehicle for the second audio dataset; Based on the target recognition results, the second audio dataset is labeled, and a simulation test set is generated based on the labeling results and the second audio dataset. The method further includes: Acquire preset first audio data and scene configuration information set for each first audio data, wherein the scene configuration information includes vehicle control parameters and playback control parameters; The scene configuration information and the first audio data are stored in a preset first data table to obtain the first audio dataset; The step of generating a simulation test set based on the annotation results and the second audio dataset includes: From the first data table, determine the scene configuration information corresponding to the first audio data corresponding to each second audio data in the second audio dataset; Based on the annotation results, the scene configuration information corresponding to each second audio data in the second audio dataset, and the second audio dataset, a simulation test set is generated.

2. The method according to claim 1, characterized in that, The step of playing a preset first audio dataset in a vehicle within a preset vehicle scenario includes: Based on the scene configuration information in the first data table, the preset vehicle scene is constructed in the vehicle, and the first audio data in the first audio dataset is played in the vehicle.

3. The method according to claim 2, characterized in that, The step of storing the scene configuration information and the first audio data into a preset first data table includes: According to the preset format requirements, the first audio data and the scene configuration information are stored in the preset first data table.

4. The method according to claim 2, characterized in that, The step of constructing the preset vehicle scene in the vehicle based on the scene configuration information in the first data table, and playing the first audio data in the first audio dataset in the vehicle, includes: The vehicle is controlled according to the vehicle control parameters to construct the preset vehicle scenario within the vehicle. In the vehicle, the first audio data in the first audio dataset is played according to the playback control parameters.

5. The method according to claim 1, characterized in that, The step of labeling the second audio dataset based on the target recognition result, and generating a simulation test set based on the labeling result and the second audio dataset, includes: Based on the target recognition results, the time points of the target words and phrases in each second audio data in the second audio dataset are determined, and the time points of the target words and phrases in each second audio data are used as the annotation results; Based on the annotation results and the second audio dataset, a simulation test set is generated.

6. The method according to claim 1, characterized in that, The method further includes: Clean up the target storage space; the target storage space is used to store the log information and / or the second audio dataset.

7. The method according to claim 1, characterized in that, The method further includes: The simulation test was performed using the second audio dataset in the simulation test set, and the first simulation test result was obtained. The simulation test is evaluated based on the results of the first simulation test and the results of the identifiers in the simulation test set.

8. A device for generating a simulation test set, characterized in that, The device includes: The playback module is used to play a preset first audio dataset in a vehicle in a preset vehicle scenario by an automated voice test set conversion device. The collection module is used to obtain a second audio dataset collected from the vehicle in response to the first audio dataset, as well as log information generated by the vehicle during the playback of the first audio dataset. The extraction module is used to obtain the target recognition result of the vehicle for the second audio dataset from the log information; The generation module is used to annotate the second audio dataset according to the target recognition result, and generate a simulation test set according to the annotation result and the second audio dataset; The device further includes: The scene information acquisition module is used to acquire preset first audio data and scene configuration information set for each first audio data; store the scene configuration information and the first audio data into a preset first data table to obtain a first audio dataset, wherein the scene configuration information includes vehicle control parameters and playback control parameters; The generation module is specifically used to determine the scene configuration information corresponding to the first audio data corresponding to each second audio data in the second audio dataset from the first data table; and to generate a simulation test set based on the annotation results, the scene configuration information corresponding to each second audio data in the second audio dataset, and the second audio dataset.

9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the method for generating a simulation test set as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the method for generating a simulation test set as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech recognition effect test method and system

    CN110415681A