Method and apparatus for splicing audio data, electronic device, and storage medium

By performing scale transformation and splicing on the sub-audio in the training data of the speech recognition model, the problem of insufficient model training data in the multi-person speaking scenario is solved, and the recognition accuracy and robustness of the model are improved.

CN115440188BActive Publication Date: 2025-06-20JINGDONG TECH HLDG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110615043.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-02
Publication Date
2025-06-20
Estimated Expiration
2041-06-02

AI Technical Summary

Technical Problem

In the multi-person speaking scenario, the number of training data of the speech recognition model is insufficient, resulting in poor robustness of the model in the multi-person speaking scenario.

Method used

By obtaining the annotation sample set of voice information of multiple target users, the sub-audio is scale-transformed using the target scheme to generate new audio data, and splice it with the original text annotation to form rich training data.

Benefits of technology

Without adding new labeled sample data, the training data is enriched, the target model's recognition accuracy of the audio data of multiple people is improved, and the model's robustness is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115440188B_ABST
    Figure CN115440188B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for splicing audio data, an electronic device, and a storage medium. The method includes: obtaining an annotation sample set of voice information of multiple target users, where the annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio; performing transformation processing on a first sub-audio using a first target scheme, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range, and the first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is an audio generated after the first sub-audio is processed by the first target scheme; splicing the second sub-audio and the first sub-text annotation to obtain spliced audio group data. The present application solves the problems of insufficient training data quantity of the model and poor robustness of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech data processing, and in particular, to a method and device for splicing audio data, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of deep learning and artificial intelligence technologies, automatic speech recognition based on end-to-end deep neural networks has gradually become the mainstream technology in the current speech recognition field.

[0003] Generally, when performing automatic speech recognition, the actual performance of a speech recognition model often depends on a large amount of manually labeled speech training data. However, the collection and manual annotation of large-scale training data are very difficult and costly. On the one hand, compared with the acquisition of data such as images and texts, speech data is usually difficult to obtain a large amount of original speech data easily because it may involve information such as privacy and copyright. On the other hand, since the annotation of audio data requires manual listening at least once, the annotation cost is also often very high.

[0004] To solve the problems of collection and annotation of large-scale training data, data augmentation algorithms are one of the most commonly used technologies. However, in many application scenarios of automatic speech recognition, such as a scenario where there are multiple oral statements in an audio data (such as a meeting scenario, a conversation in a public place, etc.), the related technologies do not record how to use data augmentation algorithms to increase the audio training data corresponding to multiple oral statements, resulting in the insufficient quantity of training data for the speech recognition model and poor robustness of the speech recognition model in a scenario where multiple people are speaking. Summary of the Invention

[0005] The present application provides a method and device for splicing audio data, a storage medium, and an electronic device to at least solve the problem that the quantity of training data for the speech recognition model in the related technologies is still insufficient and the robustness of the speech recognition model is poor in a scenario where multiple people are speaking.

[0006] According to one aspect of the embodiments of the present application, a method for splicing audio data is provided. The method includes: obtaining an annotation sample set of voice information of multiple target users, where the annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio; performing transformation processing on a first sub-audio using a first target scheme to obtain first transformation data, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range, and the first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is the audio generated after the first sub-audio is processed by the first target scheme; splicing the second sub-audio and the first sub-text annotation to obtain spliced audio group data.

[0007] According to another aspect of the embodiments of the present application, an apparatus for splicing audio data is further provided. The apparatus includes: a first acquisition unit, configured to obtain an annotation sample set of voice information of multiple target users, where the annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio; a transformation unit, configured to perform transformation processing on a first sub-audio using a first target scheme to obtain first transformation data, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range, and the first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is the audio generated after the first sub-audio is processed by the first target scheme; a splicing unit, configured to splice the second sub-audio and the first sub-text annotation to obtain spliced audio group data.

[0008] Optionally, the transformation unit includes: a determination module configured to determine the number of samples to be spliced according to the audio duration, where the audio duration is the audio duration corresponding to the first audio in the labeled sample set; a selection module configured to select a third audio that meets the number of samples to be spliced from all the first audios, where the third audio is a subset of the first audios; a first transformation module configured to perform transformation processing on the third sub-audio by using a first target scheme to obtain second transformation data, where the first target scheme is used to indicate selecting a third sub-audio different from the previously selected sub-audio from the third audio and performing scale transformation processing on the third sub-audio within a preset transformation range, and the second transformation data includes a fourth sub-audio and a second sub-text annotation corresponding to the fourth sub-audio, where the number of the fourth sub-audios, the number of the second sub-text annotations, and the number of the third audios are the same.

[0009] Optionally, the determination module includes: an acquisition subunit configured to acquire the maximum value and the minimum value of the audio duration; a determination subunit configured to determine the number of samples to be spliced by using the maximum value, the minimum value, and a preset threshold, where the preset threshold is a range critical value for selecting the number of samples to be spliced.

[0010] Optionally, the splicing unit includes: a generation module configured to generate a plurality of splicing interval durations by using a second target scheme, where the second target scheme is used to randomly generate a plurality of the splicing interval durations according to the number of the third audios, and the number of the splicing interval durations is the same as the number of the third audios; a splicing module configured to splice the second sub-audio, the first sub-text annotation, and the splicing interval durations to obtain the spliced audio group data.

[0011] Optionally, the first audio includes a plurality of sub-audios arranged in a preset order. The transformation unit includes: a second transformation module configured to, when performing transformation processing for the first time, acquire the sub-audio located first in the preset order from the first audio for transformation processing to obtain the first transformation data;

[0012] a third transformation module configured to, when not performing transformation processing for the first time, acquire the sub-audio adjacent to and after the sub-audio processed in the previous time from the first audio for transformation processing to obtain the first transformation data.

[0013] Optionally, the device further includes: a second obtaining unit, configured to obtain the audio duration corresponding to the first sub-audio located at the first position from the first audio according to the preset order after performing transformation processing on the first sub-audio by using the first target scheme to obtain first transformed data; a determining unit, configured to determine the total audio duration according to the audio duration corresponding to the first sub-audio, the audio duration corresponding to the corresponding sub-audio during this transformation processing, and the splicing interval duration; an ending unit, configured to end the transformation processing of the first audio when there is no unprocessed sub-audio in the first audio, or the total audio duration is greater than the maximum value of the audio duration, or the total number of the spliced audio group data is greater than the total number of audio of the voice information.

[0014] Optionally, the device further includes: a third obtaining unit, configured to obtain a plurality of the spliced audio group data to generate a data augmentation sample set; a mixing unit, configured to perform mixing processing on the labeled sample set and the data augmentation sample set to obtain a target training data set; a training unit, configured to train an initial model by using the target training data set to obtain a target model, where the target model is used to identify the voice information of a plurality of the target users.

[0015] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus; where the memory is configured to store a computer program; the processor is configured to execute the method steps of splicing audio data in any of the above embodiments by running the computer program stored on the memory.

[0016] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided, where a computer program is stored in the storage medium, and the computer program is configured to execute the method steps of splicing audio data in any of the above embodiments when running.

[0017] In an embodiment of the present application, by obtaining an annotation sample set of voice information of multiple target users, where the annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio; using a first target scheme to perform transformation processing on a first sub-audio to obtain first transformation data, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range, and the first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is an audio generated after the first sub-audio is processed by the first target scheme; by splicing the second sub-audio and the first sub-text annotation to obtain the spliced audio group data, since in the embodiment of the present application, each sub-audio in the first audio in the current existing annotation sample set is sequentially subjected to scale transformation processing to obtain transformed data, such as the first transformation data, and the obtained first transformation data is spliced in the time domain to obtain the spliced audio group data, in this way, without adding new annotation sample data, the training data is enriched, achieving the technical effect of improving the recognition accuracy of the target model for multi-person speech audio data, and further solving the problems that the number of training data of the speech recognition model in the related art is still insufficient and the robustness of the speech recognition model is poor in the multi-person speech scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 is a schematic diagram of the hardware environment of an optional method for splicing audio data according to an embodiment of the present invention;

[0021] Figure 2 is a schematic flow chart of an optional method for splicing audio data according to an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of an optional audio data and its corresponding text annotation according to an embodiment of the present application;

[0023] Figure 4 is a schematic overall flow chart of an optional method for splicing audio data according to an embodiment of the present application;

[0024] Figure 5 It is a structural block diagram of an optional splicing device for audio data according to an embodiment of the present application;

[0025] Figure 6 It is a structural block diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0026] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] In recent years, with the rapid development of deep learning and artificial intelligence technologies, automatic speech recognition based on end-to-end deep neural networks (such as ASR (Automatic Speech Recognition)) has gradually become the mainstream technology in the current speech recognition field. Taking the ASR automatic speech recognition model as an example, the end-to-end ASR model can directly convert a segment of audio data at the input end into a text recognition result at the output end, which is easy to train, infer, and optimize the model. Currently, it has been applied in many actual scenarios such as e-commerce, logistics, and finance.

[0029] Under normal circumstances, due to the large number of parameters in the end-to-end ASR model, the actual performance of the model often depends on a large amount of manually annotated speech training data. However, the collection and manual annotation of large-scale ASR data are very difficult and costly. On the one hand, compared with the acquisition of data such as images and texts, it is usually difficult to easily obtain a large amount of original speech data because speech data may involve information such as privacy and copyright. On the other hand, since the annotation of audio data requires manual listening at least once, the annotation cost is also often very high.

[0030] To solve the problems of collection and annotation of large-scale ASR training data, data augmentation algorithms are one of the most commonly used techniques. A speech data augmentation algorithm refers to making appropriate transformations to audio data based on the original annotated data. In this way, one manually annotated data sample will generate multiple samples after being transformed by data augmentation, thus enriching the features and quantity of training data. Existing ASR data augmentation algorithms can be mainly divided into time-domain-based data augmentation and frequency-domain-based data augmentation according to the different data domains they act on, as follows:

[0031] (1) Time-domain-based data augmentation: This type of method expands the training data by making transformations such as speed change, noise addition, and reverberation addition to the audio time-domain data sequence, modifying the speech rate, background noise, and simulating far-field sound effects of the audio data respectively.

[0032] (2) Frequency-domain-based data augmentation: This type of method enriches the training data by making transformations such as noise addition, distortion, and masking to the spectral features of the audio data, modifying the background noise, perturbation characteristics, and masking local spectral features of the audio data respectively.

[0033] However, in many speech automatic recognition application scenarios, such as scenarios where there are multiple oral statements in a piece of audio data (such as meeting scenarios, public conversations, etc.), the features at different times of each piece of audio data often exhibit non-stationary characteristics. For example, in a meeting scenario, there may be multiple people speaking. For example, in a meeting scenario audio, in the first half, Zhang San who is closer to the microphone is speaking, and in the second half, Li Si who is farther from the microphone is speaking. Due to the large differences in the pronunciation characteristics and sound source distances of different people, the accuracy of the ASR model trained with existing data based on the assumption of stationary distribution is poor. Although there have been studies attempting to use speaker segmentation technology to avoid the problem of multiple speakers in a piece of audio, however, the effect of this type of technology in actual application scenarios (especially complex speech scenarios) is still difficult to guarantee, and at the same time, the problems brought by the robustness of ASR have not been well solved. Therefore, the related technology does not record how to use data augmentation algorithms to increase the audio training data corresponding to multiple oral statements.

[0034] To solve the above problems, according to one aspect of the embodiments of the present application, a method for splicing audio data is provided. Optionally, in this embodiment, the above method for splicing audio data can be applied to, for example, Figure 1 the hardware environment shown. As Figure 1 shown, the terminal 102 may include a memory 104, a processor 106, and a display 108 (optional component). The terminal 102 can communicate with the server 112 through the network 110. The server 112 can be used to provide services for the terminal or the client installed on the terminal (such as game services, application services, etc.). A database 114 can be set on the server 112 or independently of the server 112 for providing data storage services for the server 112. In addition, a processing engine 116 can run in the server 112, and the processing engine 116 can be used to execute the steps performed by the server 112.

[0035] Optionally, the terminal 102 can be, but is not limited to, a terminal that can calculate data, such as a mobile terminal (such as a mobile phone, a tablet computer), a laptop computer, a PC (Personal Computer) machine, etc. The above network can include, but is not limited to, a wireless network or a wired network. Among them, the wireless network includes: Bluetooth, WIFI (Wireless Fidelity), and other networks that implement wireless communication. The above wired network can include, but is not limited to: a wide area network, a metropolitan area network, and a local area network. The above server 112 can include, but is not limited to, any hardware device that can perform calculations.

[0036] In addition, in this embodiment, the above method for splicing audio data can also be, but is not limited to, applied to a powerful independent processing device without data interaction. For example, the processing device can be, but is not limited to, a powerful terminal device, that is, each operation in the above method for splicing audio data can be integrated in an independent processing device. The above is only an example, and this embodiment does not make any limitation in this regard.

[0037] Optionally, in this embodiment, the above method for splicing audio data can be executed by the server 112, or can be executed by the terminal 102, or can also be jointly executed by the server 112 and the terminal 102. Among them, the terminal 102 executing the... method of the embodiments of the present application can also be executed by the client installed on it.

[0038] Taking running on the server as an example, Figure 2 is a schematic flowchart of an optional method for splicing audio data according to the embodiments of the present application. As Figure 2 shown, the process of this method can include the following steps:

[0039] Step S201: Obtain an annotation sample set of voice information of multiple target users. The annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio.

[0040] Optionally, first, the embodiments of the present application need to use a server to obtain an annotation sample set of voice information of multiple target users. Here, the target users are the users who make sounds in the current voice acquisition scenario. In the embodiments of the present application, the current voice acquisition scenario is usually a certain meeting, a certain public place, etc. Therefore, it is necessary to obtain multiple users who make sounds at this time.

[0041] The annotation sample set here includes the first audio of the voice information and the first text annotation corresponding to the first audio. It should be noted that the annotation sample set in the embodiments of the present application can be manually annotated or pre-annotated by a machine, as long as the data in the sample set has been annotated. In addition, since the number of target users obtained in the embodiments of the present application is at least one, the number of first audios corresponding to the voice information is also at least one.

[0042] Optionally, the annotation sample set can be denoted as: Φ o ={x i , y i |i = 1, 2,..., N}, where N is the number of all samples in the annotation sample set, x i is audio data with a duration of t i , and y i is the text annotation corresponding to the audio data. As Figure 3 shown, in Figure 3 , a certain audio data is x i , and its corresponding text annotation is y i : "What's the weather like today?".

[0043] Step S202: Use a first target scheme to perform transformation processing on a first sub-audio to obtain first transformation data. The first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audios from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range. The first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio. The second sub-audio is the audio generated after the first sub-audio is processed by the first target scheme.

[0044] Optionally, the embodiments of the present application perform transformation processing on the first sub-audio selected from the first audio using the first target scheme. More specifically, the embodiments of the present application randomly assign corresponding scale transformation factors and sequence numbers for processing each sub-audio in the first audio, and then perform transformation processing on the sub-audio according to the sequence numbers according to the scale transformation factors.

[0045] Among them, the above-mentioned first sub-audio is a sub-audio selected from the first audio by the first target scheme, which is different from the previously selected sub-audio. At the same time, the first sub-audio also needs to be ranked high when the sequence number is selected, that is, the sub-audio ranked first among the unselected sub-audios.

[0046] In addition, the scale transformation factor s in the embodiment of the present application is k ∈[-S,S] constitutes a preset transformation range. In the embodiment of the present application, S=10, so the preset transformation range is: [-10dB, 10dB]. Therefore, the embodiment of the present application transforms the first sub-audio within the preset transformation range. At this time, the first transformation data will be obtained, and the first transformation data includes the second sub-audio after the transformation of the first sub-audio and the first sub-text annotation corresponding to the second sub-audio.

[0047] Step S203: splice the second sub-audio and the first sub-text annotation to obtain spliced ​​audio group data.

[0048] Optionally, the server of the embodiment of the present application splices the acquired second sub-audio and its corresponding first sub-text annotation, and then generates a spliced ​​audio group data.

[0049] It can be understood that the above embodiment is that after the first target scheme transforms the first sub-audio, the second sub-audio obtained after the processing and its corresponding first sub-text annotation are spliced ​​to generate a spliced ​​audio group data. The embodiment of the present application also needs to use the first target scheme to transform the other sub-audios of the first audio, and then splice the obtained transformed data, so that a plurality of spliced ​​audio group data can be obtained, so that a plurality of transformed training sample data can be obtained, and a data enhanced sample set can be obtained.

[0050] In an embodiment of the present application, by obtaining an annotation sample set of voice information of multiple target users, where the annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio; using a first target scheme to perform transformation processing on the first sub-audio to obtain first transformation data, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range, the first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is an audio generated after the first sub-audio is processed by the first target scheme; splicing the second sub-audio and the first sub-text annotation to obtain spliced audio group data. Since in the embodiment of the present application, scale transformation processing is sequentially performed on each sub-audio in the first audio in the current existing annotation sample set to obtain transformed data, such as the first transformation data, and the obtained first transformation data is spliced in the time domain to obtain spliced audio group data, in this way, without adding new annotation sample data, the training data is enriched, achieving the technical effect of improving the recognition accuracy of the target model for multi-person speech audio data, and further solving the problems that the number of training data of the speech recognition model in the related art is still insufficient and the robustness of the speech recognition model is poor in the scenario of multi-person speech.

[0051] As an alternative embodiment, using the first target scheme to perform transformation processing on the first sub-audio to obtain the first transformation further includes:

[0052] Determine the number of samples to be spliced according to the audio duration, where the audio duration is the audio duration corresponding to the first audio in the annotation sample set;

[0053] Select a third audio that meets the number of samples to be spliced from all the first audios, and the third audio is a subset of the first audio;

[0054] Use the first target scheme to perform transformation processing on the third sub-audio to obtain second transformation data, where the first target scheme is used to indicate selecting a third sub-audio different from the previously selected sub-audio from the third audio and performing scale transformation processing on the third sub-audio within a preset transformation range, the second transformation data includes a fourth sub-audio and a second sub-text annotation corresponding to the fourth sub-audio, and the fourth sub-audio is an audio generated after the third sub-audio is processed by the first target scheme, and the number of the fourth sub-audios, the number of the second sub-text annotations, and the number of the third audios are the same.

[0055] Optionally, embodiments of the present application may first obtain a part of the audio from the first audio as a splicing sample. More specifically, embodiments of the present application may determine the number of samples to be spliced (such as K) according to the audio duration corresponding to each audio sample in the first audio, then select a third audio that meets the number of samples to be spliced from all the first audio, and then perform transformation processing on the third sub-audio in the third audio using the first target scheme.

[0056] It can be understood that since the third audio is selected from the first audio, the third audio is a subset of the first audio, and the number of sub-audios included in the third audio should be the same as the number of samples to be spliced, both being K; the way the first target scheme performs transformation processing on the third sub-audio in the third audio is the same as the way in the foregoing embodiments, that is, performing transformation processing on the third sub-audio within a preset transformation range to obtain second transformation data. However, when embodiments of the present application perform transformation on the third sub-audio, K scale transformation factors need to be randomly generated by simulating a uniform distribution. Therefore, the number of fourth sub-audios included in the generated second transformation data and the number of second sub-text annotations corresponding to the fourth sub-audios are also both K.

[0057] In embodiments of the present application, the process of random sampling can be simulated, the number of samples to be spliced can be determined using the audio duration, and then scale transformation processing is performed on each sub-audio in the third audio that meets the number of samples to be spliced. In this way, without increasing the training data in the original annotation sample set, the training data is made more abundant.

[0058] As an optional embodiment, determining the number of samples to be spliced according to the audio duration includes:

[0059] Obtain the maximum value and minimum value of the audio duration;

[0060] Use the maximum value, minimum value, and a preset threshold to determine the number of samples to be spliced, where the preset threshold is the range critical value for selecting the number of samples to be spliced.

[0061] Optionally, when embodiments of the present application determine the number of samples to be spliced, they can use the audio duration of each audio in the annotation sample set. For example, obtain the maximum value T max , minimum value T min , and then according to the preset threshold set by the server in advance, such as the value 2, where 2 is the range critical value for selecting the number of samples to be spliced.

[0062] Use T max , T min and 2 to generate the number of samples to be spliced. For example, the number of samples to be spliced Among them, is the floor function. For example, T max takes the value of 15, Tmin If the value is 2, then T max / T min = 15 / 2 = 7.5, then Therefore, K ∈ [2, 7], and K can take any value among 2, 3, 4, 5, 6, and 7.

[0063] In the embodiment of the present application, the maximum value, minimum value, and preset threshold of the audio duration are used to determine the number of samples to be spliced. In this way, the embodiment of the present application first performs transformation processing on audio samples within a certain range to increase the quantity of audio data.

[0064] As an alternative embodiment, the second sub-audio and the first sub-text annotation are spliced to obtain the spliced audio group data, including:

[0065] Generate multiple splicing interval durations using the second target scheme, where the second target scheme is used to randomly generate multiple splicing interval durations according to the number of the third audio, and the number of splicing interval durations is the same as the number of the third audio;

[0066] Splice the second sub-audio, the first sub-text annotation, and the splicing interval duration to obtain the spliced audio group data.

[0067] Optionally, the embodiment of the present application uses the second target scheme, such as the uniform distribution method, to generate multiple splicing interval durations. Among them, the number of splicing interval durations is the same as the number of audio samples in the third audio, both are K, and the splicing interval duration d k ∈ [0, D], where D is the maximum splicing time interval. In the embodiment of the present application, D = 0.5s.

[0068] Then, splice the second sub-audio and the first sub-text annotation obtained in the above embodiment with the splicing interval duration obtained in the embodiment of the present application to obtain the spliced audio group data. Optionally, between any two second sub-audios, add a splicing interval duration corresponding to the audio interval length. In this way, the audio group data spliced by the second sub-audio, the first sub-text annotation, and the splicing interval duration is used as the training data in the data enhancement sample set.

[0069] As an alternative embodiment, the first audio includes multiple sub-audios arranged in a preset order. Among them, the first sub-audio is transformed using the first target scheme to obtain the first transformed data, including:

[0070] When performing the transformation processing for the first time, obtain the sub-audio located at the first position in the preset order from the first audio for transformation processing to obtain the first transformed data;

[0071] When it is not the first time to perform the transformation process, obtain the sub-audio adjacent to and after the sub-audio processed in the previous time from the first audio for transformation processing to obtain the first transformation data.

[0072] Optionally, when performing the first transformation process on the sub-audio in the first audio, at this time, it is necessary to obtain the sub-audio arranged in the first position in the preset order from the first audio, and then perform transformation processing on the sub-audio to obtain the first transformation data. At this time, set the loop parameter k = 1 (1 ≤ k ≤ K), the obtained sub-audio is the sampling sample x1 in the labeled sample set, and the corresponding text label is y1. Then set the loop parameter k = k + 1 and perform the second round of transformation processing. At this time, k = 2;

[0073] When performing the second round (k = 2), the third round (k = 3)... the Kth round (k = K) of transformation processing, only need to obtain the sub-audio adjacent to and after the sub-audio processed in the previous time from the first audio for transformation processing to obtain the first transformation data.

[0074] In the embodiment of the present application, perform cyclic transformation processing on the sub-audio in the first audio to ensure that each sub-audio has undergone transformation processing and ensure the richness of the training data in the enhanced sample set.

[0075] As an alternative embodiment, after using the first target scheme to perform transformation processing on the first sub-audio to obtain the first transformation data, the method further includes:

[0076] Obtain the audio duration corresponding to the sub-audio located in the first position from the first audio according to the preset order;

[0077] Determine the total audio duration according to the audio duration corresponding to the first sub-audio, the audio duration corresponding to the corresponding sub-audio during this transformation processing, and the splicing interval duration;

[0078] In the case that there is no unprocessed sub-audio in the first audio, or the total audio duration is greater than the maximum value of the audio duration, or the total number of the spliced audio group data is greater than the total number of audio of the voice information, end the transformation processing of the first audio.

[0079] Optionally, in the embodiment of the present application, it is necessary to obtain the audio duration t1 corresponding to the sub-audio sampling sample x1 when k = 1, and then obtain the audio duration (such as t 100 ) corresponding to the corresponding sub-audio (such as x 100 ) during this transformation processing. Since there is a corresponding splicing interval duration d k added between any two second sub-audios, so there are K splicing interval durations d 100 between x1 and x k, at this time, add t1 + t 100 + K × d k = the total audio duration.

[0080] To balance the scales of the labeled sample set and the data augmentation sample set, set the number of sample training data in these two data sets to be the same, both being M.

[0081] Set condition 1: There is no unprocessed sub-audio in the first audio; condition 2: The total audio duration is greater than the maximum value T of the audio duration max ; condition 3: The total number of the spliced audio group data (such as M) is greater than the total number of audio of the voice information (such as M) as the condition to end the transformation process of the first audio. Meeting any of the above conditions can end the transformation process of the first audio. Otherwise, splice the transformed data obtained from this transformation process with the splicing interval duration to obtain new audio group data.

[0082] As an optional embodiment, after splicing the second sub-audio and the first sub-text annotation to obtain the spliced audio group data, the method further includes:

[0083] Obtain multiple spliced audio group data and generate a data augmentation sample set;

[0084] Mix the labeled sample set and the data augmentation sample set to obtain a target training data set;

[0085] Use the target training data set to train the initial model to obtain a target model, where the target model is used to identify the voice information of multiple target users.

[0086] Optionally, from multiple spliced audio group data, generate an audio group set L. At this time, for each audio group data in the audio group set L, perform audio time domain splicing to obtain a new data augmentation sample set Φ a . Specifically, it includes:

[0087] (1) The duration of each data augmentation sample is: the total duration of all sub-audios in the audio group data plus the total sum of all splicing interval durations.

[0088] (2) The audio data sequence of the data augmentation sample is: the splicing of all audios in the audio group data, and between any two sub-audios, add the splicing interval duration corresponding to the sub-audio interval length.

[0089] (3) The audio data annotation of the data augmentation sample is: the splicing of the sub-text annotations corresponding to all sub-audios in the audio group data.

[0090] Through the above method, M data augmentation samples can be obtained to obtain the data augmentation data set Φ a .

[0091] Then, the data augmentation dataset Φ a is mixed with the labeled sample set Φ o to obtain the target training dataset Φ = Φ a ∪Φ o . By using the dataset Φ to train the initial model, a target model is obtained. Among them, the target model is used to recognize the speech information of multiple target users, and the initial model is trained by the training data in the labeled sample set Φ o . The initial model can be an ASR model.

[0092] In the embodiments of the present application, without adding new labeled data, the training data is enriched, and the recognition accuracy of the target model for multi-person speech audio data is improved.

[0093] As an alternative embodiment, as Figure 4 , Figure 4 is a schematic diagram of the overall process of an alternative method for splicing audio data according to the embodiments of the present application. The specific parts include:

[0094] (1) Audio group generation and processing part: From the existing ASR training data, according to the duration characteristics of each data sample, multiple audios are randomly selected to form an audio group; all the audios in the audio group are randomly assigned corresponding scale transformation factors and sequence numbers, and transformation processing is carried out in sequence according to the numbers.

[0095] (2) Audio splicing part: The processed data of the audios in the audio group are spliced into a whole audio in the time domain to achieve data augmentation.

[0096] (3) Model training part: The data-augmented audio and the original training audio are mixed to train the ASR model (i.e., the initial model) to obtain the final target model.

[0097] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present application.

[0099] According to another aspect of the embodiments of the present application, there is also provided an audio data splicing device for implementing the above audio data splicing method. Figure 5 FIG. is a structural block diagram of an optional audio data splicing device according to an embodiment of the present application, as Figure 5 shown. The device may include:

[0100] A first acquisition unit 501, configured to acquire an annotation sample set of voice information of multiple target users. The annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio;

[0101] A transformation unit 502, connected to the first acquisition unit 501, configured to perform transformation processing on a first sub-audio by using a first target scheme to obtain first transformation data. The first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range. The first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is an audio generated after the first sub-audio is processed by the first target scheme;

[0102] A splicing unit 503, connected to the transformation unit 502, configured to splice the second sub-audio and the first sub-text annotation to obtain spliced audio group data.

[0103] It should be noted that the first acquisition unit 501 in this embodiment can be used to execute the above step S201, the transformation unit 502 in this embodiment can be used to execute the above step S202, and the splicing unit 503 in this embodiment can be used to execute the above step S203.

[0104] Through the above module, by obtaining an annotation sample set of voice information of multiple target users, where the annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio; using a first target scheme to perform transformation processing on a first sub-audio to obtain first transformation data, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range, and the first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is an audio generated after the first sub-audio is processed by the first target scheme; by splicing the second sub-audio and the first sub-text annotation to obtain the spliced audio group data, since in the embodiment of the present application, each sub-audio in the first audio in the currently existing annotation sample set is sequentially subjected to scale transformation processing to obtain transformed data, such as the first transformation data, and the obtained first transformation data is spliced in the time domain to obtain the spliced audio group data, in this way, without adding new annotation sample data, the training data is enriched, achieving the technical effect of improving the recognition accuracy of the target model for multi-person speech audio data, and further solving the problems that the number of training data of the speech recognition model in the related technology is still insufficient and the robustness of the speech recognition model is poor in the scenario of multi-person speech.

[0105] As an alternative embodiment, the transformation unit 502 includes: a determination module, configured to determine the number of samples to be spliced according to the audio duration, where the audio duration is the audio duration corresponding to the first audio in the annotation sample set; a selection module, configured to select a third audio that meets the number of samples to be spliced from all the first audios, and the third audio is a subset of the first audio; a first transformation module, configured to use the first target scheme to perform transformation processing on the third sub-audio to obtain second transformation data, where the first target scheme is used to indicate selecting a third sub-audio different from the previously selected sub-audio from the third audio and performing scale transformation processing on the third sub-audio within a preset transformation range, and the second transformation data includes a fourth sub-audio and a second sub-text annotation corresponding to the fourth sub-audio, and the fourth sub-audio is an audio generated after the third sub-audio is processed by the first target scheme, and the number of the fourth sub-audios, the number of the second sub-text annotations, and the number of the third audios are the same.

[0106] As an alternative embodiment, the determination module includes: an acquisition subunit, configured to acquire the maximum value and the minimum value of the audio duration; a determination subunit, configured to use the maximum value, the minimum value, and a preset threshold to determine the number of samples to be spliced, where the preset threshold is the range critical value for selecting the number of samples to be spliced.

[0107] As an alternative embodiment, the splicing unit 503 includes: a generation module configured to generate a plurality of splicing interval durations by using a second target scheme, where the second target scheme is configured to randomly generate a plurality of splicing interval durations according to the number of third audio signals, and the number of splicing interval durations is the same as the number of third audio signals; and a splicing module configured to splice the second sub-audio, the first sub-text annotation, and the splicing interval durations to obtain spliced audio group data.

[0108] As an alternative embodiment, the first audio includes a plurality of sub-audio signals arranged in a preset order. The transformation unit includes: a second transformation module configured to, when performing the transformation process for the first time, obtain the sub-audio signal that is the first in the preset order from the first audio and perform the transformation process to obtain first transformation data;

[0109] a third transformation module configured to, when not performing the transformation process for the first time, obtain the sub-audio signal that is adjacent to and after the sub-audio signal processed in the previous time from the first audio and perform the transformation process to obtain first transformation data.

[0110] As an alternative embodiment, the apparatus further includes: a second obtaining unit configured to, after obtaining first transformation data by performing the transformation process on the first sub-audio by using a first target scheme, obtain the audio duration corresponding to the sub-audio signal that is the first in the preset order from the first audio according to the preset order; a determination unit configured to determine the total audio duration according to the audio duration corresponding to the first sub-audio signal, the audio duration corresponding to the corresponding sub-audio signal when performing the transformation process this time, and the splicing interval durations; and an end unit configured to end the transformation process of the first audio when there is no unprocessed sub-audio signal in the first audio, or the total audio duration is greater than the maximum value of the audio duration, or the total number of the spliced audio group data is greater than the total number of audio signals of the voice information.

[0111] As an alternative embodiment, the apparatus further includes: a third obtaining unit configured to obtain a plurality of spliced audio group data and generate a data augmentation sample set; a mixing unit configured to mix the annotation sample set and the data augmentation sample set to obtain a target training data set; and a training unit configured to train an initial model by using the target training data set to obtain a target model, where the target model is configured to recognize the voice information of a plurality of target users.

[0112] It should be noted here that the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as a part of the apparatus, can run in a hardware environment as shown in Figure 1 and can be implemented by software or by hardware, where the hardware environment includes a network environment.

[0113] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above audio data splicing method, and the electronic device may be a server, a terminal, or a combination thereof.

[0114] Figure 6 is a structural block diagram of an optional electronic device according to the embodiments of the present application. As Figure 6 shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604. Among them,

[0115] The memory 603 is used to store computer programs;

[0116] The processor 601, when executing the computer program stored on the memory 603, implements the following steps:

[0117] S1, obtaining an annotation sample set of voice information of multiple target users, where the annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio;

[0118] S2, performing transformation processing on the first sub-audio using a first target scheme to obtain first transformation data, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range. The first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is the audio generated after the first sub-audio is processed by the first target scheme;

[0119] S3, splicing the second sub-audio and the first sub-text annotation to obtain spliced audio group data.

[0120] Optionally, in this embodiment, the above communication bus may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0121] The communication interface is used for communication between the above electronic device and other devices.

[0122] The memory may include RAM, or may also include non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0123] As an example, as Figure 6 shown, the above-mentioned memory 603 may but is not limited to include the first acquisition unit 501, transformation unit 502, and splicing unit 503 in the splicing device of the above audio data. In addition, it may also include but is not limited to other module units in the splicing device of the above audio data, which will not be elaborated in this example.

[0124] The above-mentioned processor may be a general-purpose processor, which may include but is not limited to: CPU (Central Processing Unit, central processor), NP (Network Processor, network processor), etc.; it may also be a DSP (Digital Signal Processing, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field-Programmable Gate Array, field programmable gate array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0125] In addition, the above-mentioned electronic device further includes: a display for displaying the splicing result of audio data.

[0126] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiment, and will not be elaborated here.

[0127] Those of ordinary skill in the art can understand that Figure 6 the structure shown is only schematic. The device for implementing the above audio data splicing method may be a terminal device, and the terminal device may be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD and other terminal devices. Figure 6 It does not limit the structure of the above-mentioned electronic device. For example, the terminal device may further include more or fewer components (such as a network interface, a display device, etc.) than Figure 6 shown, or have a different configuration from Figure 6 shown.

[0128] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the relevant hardware of the terminal device, and the program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, ROM, RAM, a magnetic disk, or an optical disc, etc.

[0129] According to another aspect of the embodiments of the present application, a storage medium is also provided. Optionally, in this embodiment, the above storage medium can be used to execute the program code of the audio data splicing method.

[0130] Optionally, in this embodiment, the above storage medium can be located on at least one of the multiple network devices in the network shown in the above embodiments.

[0131] Optionally, in this embodiment, the storage medium is set to store the program code for executing the following steps:

[0132] S1, obtain an annotation sample set of the voice information of multiple target users, where the annotation sample set includes the first audio of the voice information and the first text annotation corresponding to the first audio;

[0133] S2, perform transformation processing on the first sub-audio using the first target scheme to obtain first transformation data, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range. The first transformation data includes a second sub-audio and the first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is the audio generated after the first sub-audio is processed by the first target scheme;

[0134] S3, splice the second sub-audio and the first sub-text annotation to obtain the spliced audio group data.

[0135] Optionally, the specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be elaborated herein.

[0136] Optionally, in this embodiment, the above storage medium can include but is not limited to: various media such as a USB flash drive, ROM, RAM, a mobile hard disk, a magnetic disk, or an optical disc that can store program code.

[0137] According to another aspect of the embodiments of the present application, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; the processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the audio data splicing method in any of the above embodiments.

[0138] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0139] If the integrated units in the above embodiments are implemented in the form of software function units and sold or used as independent products, they can be stored in the above computer-readable storage media. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the audio data splicing method in various embodiments of the present application.

[0140] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0141] In the several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0142] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution provided in this embodiment.

[0143] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software function units.

[0144] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for splicing audio data, characterized in that, The method includes: Obtaining an annotation sample set of voice information of multiple target users, where the annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio; Performing transformation processing on a first sub-audio using a first target scheme to obtain first transformation data, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing scale transformation processing on the first sub-audio within a preset transformation range, and the first transformation data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is an audio generated after the first sub-audio is processed by the first target scheme; Splicing the second sub-audio and the first sub-text annotation to obtain spliced audio group data.

2. The method according to claim 1, characterized in that, The performing transformation processing on a first sub-audio using a first target scheme to obtain first transformation further includes: Determining the number of samples to be spliced according to the audio duration, where the audio duration is the audio duration corresponding to the first audio in the annotation sample set; Selecting third audios that meet the number of samples to be spliced from all the first audios, and the third audios are a subset of the first audios; Performing transformation processing on a third sub-audio using a first target scheme to obtain second transformation data, where the first target scheme is used to indicate selecting a third sub-audio different from the previously selected sub-audio from the third audio and performing scale transformation processing on the third sub-audio within a preset transformation range, and the second transformation data includes a fourth sub-audio and a second sub-text annotation corresponding to the fourth sub-audio, and the fourth sub-audio is an audio generated after the third sub-audio is processed by the first target scheme, and the number of the fourth sub-audios, the number of the second sub-text annotations, and the number of the third audios are the same.

3. The method according to claim 2, characterized in that, The determining the number of samples to be spliced according to the audio duration includes: Obtaining the maximum value and the minimum value of the audio duration; Determining the number of samples to be spliced using the maximum value, the minimum value, and a preset threshold, where the preset threshold is a range critical value for selecting the number of samples to be spliced.

4. The method according to claim 2, characterized in that, The splicing the second sub-audio and the first sub-text annotation to obtain spliced audio group data includes: Generating multiple splicing interval durations using a second target scheme, where the second target scheme is used to randomly generate multiple splicing interval durations according to the number of the third audios, and the number of the splicing interval durations is the same as the number of the third audios; Splicing the second sub-audio, the first sub-text annotation, and the splicing interval durations to obtain the spliced audio group data.

5. The method according to claim 4, characterized in that, The first audio includes multiple sub-audios arranged in a preset order in sequence, where the performing transformation processing on a first sub-audio using a first target scheme to obtain first transformation data includes: When performing transformation processing for the first time, obtaining the sub-audio located at the first position in the preset order from the first audio for transformation processing to obtain the first transformation data; When it is not the first time to perform the transformation process, obtain the sub-audio adjacent to and after the sub-audio processed in the previous time from the first audio for transformation processing to obtain the first transformed data.

6. The method according to claim 5, characterized in that, After performing the transformation process on the first sub-audio using the first target scheme to obtain the first transformed data, the method further includes: Obtain the audio duration corresponding to the sub-audio located first from the first audio according to the preset order; Determine the total audio duration according to the audio duration corresponding to the first sub-audio, the audio duration corresponding to the corresponding sub-audio during this transformation process, and the splicing interval duration; End the transformation process of the first audio when there is no unprocessed sub-audio in the first audio, or the total audio duration is greater than the maximum value of the audio duration, or the total number of the spliced audio group data is greater than the total number of audio of the voice information.

7. The method according to any one of claims 1 to 6, characterized in that, After splicing the second sub-audio and the first sub-text annotation to obtain the spliced audio group data, the method further includes: Obtain a plurality of the spliced audio group data to generate a data augmentation sample set; Perform a mixing process on the annotation sample set and the data augmentation sample set to obtain a target training data set; Use the target training data set to train an initial model to obtain a target model, where the target model is used to recognize the voice information of a plurality of the target users.

8. An audio data splicing device, characterized in that, The device includes: A first acquisition unit, configured to acquire an annotation sample set of voice information of a plurality of target users, where the annotation sample set includes a first audio of the voice information and a first text annotation corresponding to the first audio; A transformation unit, configured to perform transformation processing on a first sub-audio using a first target scheme to obtain first transformed data, where the first target scheme is used to indicate selecting a first sub-audio different from the previously selected sub-audio from the first audio and performing a scale transformation process on the first sub-audio within a preset transformation range, and the first transformed data includes a second sub-audio and a first sub-text annotation corresponding to the second sub-audio, and the second sub-audio is an audio generated after the first sub-audio is processed by the first target scheme; A splicing unit, configured to splice the second sub-audio and the first sub-text annotation to obtain the spliced audio group data.

9. An electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein, The processor, the communication interface, and the memory complete communication with each other through the communication bus, and it is characterized in that The memory is used to store a computer program; The processor is configured to execute the steps of the audio data splicing method according to any one of claims 1 to 7 by running the computer program stored on the memory.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, where the computer program is set to execute the steps of the audio data splicing method according to any one of claims 1 to 7 when running.

Citation Information

Patent Citations

  • Voice synthesis model training method and system for automatically adding corpora

    CN110390928A

  • Voice recognition model applied to power industry

    CN110930995A