Audio data archiving method and related apparatus

CN121506093BActive Publication Date: 2026-10-09HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411053425.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-10-09
Estimated Expiration
2044-08-01

AI Technical Summary

Benefits of technology

[0007]By matching the audio features of an audio file with those in a feature comparison library, and obtaining the feature number of the audio file, audio files with the same feature number are archived together based on the feature number. This ensures that audio files from the same speaker are grouped together in the final archive set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506093B_ABST
    Figure CN121506093B_ABST
Patent Text Reader

Abstract

The application discloses an audio data archiving method and related equipment. The method comprises the following steps: dividing a first audio to obtain at least one second audio, the second audio being an audio segment containing human voice in the first audio; matching an audio feature of the second audio with an audio feature in a feature comparison library to obtain a feature number corresponding to the second audio, the feature comparison library comprising a mapping relationship between the audio feature and the feature number; and archiving the second audio based on the feature number of the second audio, the second audio being used for training a speech model, and the speech model being used for speech synthesis. The method can accurately archive the audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and in particular to audio data archiving methods and related equipment. Background Technology

[0002] Speech synthesis, also known as text-to-speech (TTS), is an important research direction in the field of speech processing, aiming to enable machines to generate natural and pleasant human speech. Speech synthesis has evolved through general emotional speech synthesis and emotionalized speech synthesis. With the rapid development of digital virtual humans and short videos, personalization is receiving increasing attention, and users have a strong demand for personalized visual avatars, personalized content generation, and personalized speech synthesis.

[0003] Currently, personalized speech can be synthesized using personalized voice assistant models. This technology relies on training the personalized voice assistant model with a large amount of audio to learn speaker features. This allows the personalized voice assistant model to synthesize the speaker's audio based on the user's input audio during application.

[0004] Before training a personalized voice assistant model, massive amounts of audio data need to be archived by different speakers. This archived audio data is then used to train the personalized voice assistant model. For example, if audio data from a blogger's past videos (containing different speakers) is used to train a personalized voice assistant model, the blogger's audio data needs to be archived before training. Archiving involves categorizing this audio data according to different speakers; specifically, speaker A could be audio 1, audio 2, and audio 3; speaker B could be audio 4. How to accurately archive this data is a problem that needs to be solved. Summary of the Invention

[0005] This application provides an audio data archiving method and related equipment, which can accurately archive audio data.

[0006] Firstly, some embodiments of this application provide an audio data archiving method. This audio data archiving method may include: segmenting a first audio file to obtain at least one second audio file, wherein the second audio file is an audio segment of the first audio file containing human voice; matching the audio features of the second audio file with audio features in a feature comparison library to obtain a feature number corresponding to the second audio file, wherein the feature comparison library includes a mapping relationship between audio features and feature numbers; archiving the second audio file based on its feature number, wherein the second audio file is used to train a speech model, and the speech model is used for speech synthesis.

[0007] By matching the audio features of an audio file with those in a feature comparison library, and obtaining the feature number of the audio file, audio files with the same feature number are archived together based on the feature number. This ensures that audio files from the same speaker are grouped together in the final archive set.

[0008] In one possible implementation, the audio features of the second audio are matched with audio features in a feature comparison library to obtain the feature number corresponding to the second audio. This includes: if the similarity between the first audio feature in the feature comparison library and the audio features of the second audio is higher than a first threshold, then the feature number corresponding to the first audio feature is used as the feature number corresponding to the second audio; if no audio feature in the feature comparison library has a similarity higher than the first threshold with the audio features of the second audio, then a new number is added as the feature number corresponding to the second audio.

[0009] In this way, audio features are associated with the speaker. By matching audio features in the feature comparison library with those of the second audio, the feature numbers of similar audio features are used as the feature numbers of the second audio, thus making the feature numbers associated with the speaker of the second audio. For second audio without similar audio features, new feature numbers are added to identify the second audio as audio associated with a different speaker.

[0010] In one possible implementation, the newly added number is used as the feature number corresponding to the second audio, including: adding the newly added number as the feature number corresponding to the second audio, and adding the audio features of the second audio and the feature number corresponding to the second audio to the feature comparison library.

[0011] By using the above method, after adding a new number to the second audio, the new number is added to the feature comparison library and the feature comparison library is dynamically updated so that subsequent audio with the same speaker as the second audio can obtain the same feature number as the second audio, thus archiving the second audio under the same speaker.

[0012] In one possible implementation, the number of successful matches for each audio feature in the feature comparison library is recorded. The number of successful matches is the number of times that the audio features in the feature comparison library are matched with a similarity higher than a first threshold. The audio features of the second audio are matched with the audio features in the feature comparison library in descending order of the number of successful matches for each audio feature in the feature comparison library.

[0013] By using the above method, audio features that have been matched successfully the most times in the feature comparison library are prioritized to improve matching efficiency and thus determine the feature number more quickly.

[0014] In one possible implementation, the audio features of the second audio are matched with the audio features in the feature comparison library in descending order of the number of successful matches corresponding to each audio feature. This includes: matching the audio features of the second audio with the audio features in the first N audio features included in the feature comparison library in descending order of the number of successful matches corresponding to each audio feature in the feature comparison library, where N is an integer greater than 1.

[0015] By using the above method, only the first N audio features with a high number of successful matches are matched, thus further improving the matching efficiency.

[0016] In one possible implementation, when the number of feature IDs in the feature comparison library is higher than a second threshold, audio features and feature IDs with fewer than a third threshold of successful matches are removed from the feature comparison library.

[0017] By using the above method, audio features in the feature comparison library are dynamically removed, avoiding an excessive number of audio features in the feature comparison library that would lead to low matching efficiency.

[0018] In one possible implementation, before segmenting the first audio to obtain at least one second audio, the method further includes: acquiring multiple third audio segments with a time length less than a fourth threshold; splicing the multiple third audio segments with a time length less than the fourth threshold to obtain the first audio, wherein there are silent segments among the various third audio segments in the first audio.

[0019] By using the above method, short videos are pre-stitched and then processed, making it possible to accurately extract audio features from the audio.

[0020] Secondly, this application provides an audio data archiving device, which can be an electronic device, a device within an electronic device, or a device compatible with an electronic device. The audio data archiving device can also be a chip system, capable of executing the methods performed by the electronic device in the first aspect. The functions of the audio data archiving device can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more units corresponding to the aforementioned functions. These units can be software and / or hardware. The operations performed by the audio data archiving device and its beneficial effects are described in the first aspect above, and will not be repeated here.

[0021] Thirdly, this application provides an electronic device including one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, including computer instructions, which, when executed by the one or more processors, cause the electronic device to perform the audio data archiving method of any possible implementation of the first aspect described above.

[0022] Fourthly, this application provides a chip system including a processor and an interface, the processor and the interface being coupled; the interface is used to receive or output signals, and the processor is used to execute code instructions to perform the audio data archiving method in any possible implementation of the first aspect above.

[0023] Fifthly, this application provides a computer-readable storage medium storing a computer program / instructions that, when the computer program product is run on a computer, cause the computer to perform the audio data archiving method in any possible implementation of the first aspect described above.

[0024] Sixthly, this application provides a computer program product that, when run on a computer, causes the computer to execute the audio data archiving method in any possible implementation of the first aspect above. Attached Figure Description

[0025] Figure 1 A flowchart illustrating an audio data archiving method provided in an embodiment of this application;

[0026] Figure 2A A schematic diagram of a first audio signal provided in an embodiment of this application;

[0027] Figure 2B A schematic diagram of a feature comparison library provided in an embodiment of this application;

[0028] Figure 2C A schematic diagram of archived audio data provided in an embodiment of this application;

[0029] Figure 3 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0030] Figure 4 A schematic diagram of the software structure of an electronic device provided in an embodiment of this application;

[0031] Figure 5 A schematic diagram of an audio data archiving device provided in an embodiment of this application;

[0032] Figure 6This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0034] It should be understood that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0035] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0036] To facilitate understanding of the solutions provided in the embodiments of this application, the relevant concepts involved in the embodiments of this application are introduced below:

[0037] I. Speech Synthesis

[0038] Speech synthesis, also known as Text-to-Speech (TTS), is an important research direction in the field of speech processing. It aims to enable machines to generate natural and pleasant human speech, converting text information into sound. A typical TTS system consists of a text processing module, a speech synthesis module, and a prosodic module. The text processing module analyzes linguistic features, such as grammar and semantics, from the input text; the speech synthesis module generates the speech based on these features; and the prosodic module adjusts the speech rate, pitch, and pauses to make the synthesized speech more natural and fluent. Modern TTS systems widely employ deep learning techniques, such as Tacotron and WaveNet. TTS applications can include reading audiobooks, news broadcasting, and smart hardware.

[0039] II. Personalized Speech Synthesis

[0040] Personalized speech synthesis refers to synthesizing the voice of the speaker input by the user. For example, in a news reading scenario, the user inputs their own voice, and the terminal performs text-to-speech conversion based on that user's own voice, so that the final product is the user's own voice reading the news. Personalized speech synthesis has the potential for widespread application in a wide range of intelligent electronic terminals, including computers, mobile phones, e-readers, MP3 players, car navigation systems, car phones, smart homes, intelligent transportation, virtual robots, vehicle-to-everything (V2X) systems, and the Internet of Things (IoT), with a vast array of application scenarios.

[0041] The following is combined with Figure 1 This application provides a detailed description of an audio data archiving method based on embodiments. For example... Figure 1 As shown, the audio data archiving method is as follows: steps 101-103. Figure 1 The method shown can be executed by an electronic device. Or Figure 1 The subject of the method shown can be a chip or chip system in an electronic device, but this application does not limit it. Figure 5 The method will be explained using an electronic device as the executing entity. Specifically:

[0042] 101. An electronic device segments a first audio file to obtain at least one second audio file, wherein the second audio file is an audio segment of the first audio file that contains human voice.

[0043] Optionally, the first audio may include one or more second audios. Different second audios can be audios corresponding to different speakers, or they can be second audios corresponding to the same speaker. For example, the first audio may include three second audios: second audio A, second audio B, and second audio C. Second audio A and second audio B are audios corresponding to speaker A, and second audio C is audio corresponding to speaker C.

[0044] Optionally, there may be silent segments between the various second audio segments in the first audio. For example, such as... Figure 2A As shown, Figure 2A 201 in the first audio is a second audio digit. Figure 2A 202 in the text is a silent segment.

[0045] Optionally, the electronic device performs Voice Activity Detection (VAD) on the first audio and obtains the VAD result; the electronic device segments the first audio based on the VAD result to obtain at least one second audio.

[0046] This VAD is primarily used to determine the presence of speech activity in an audio signal, with its core purpose being to distinguish between speech and non-speech signals. In this application, the VAD is mainly used to identify valid human voice segments. For example... Figure 2A As shown, silent segments correspond to a low level in the VAD results, while active vocal segments correspond to a high level in the VAD results.

[0047] In one possible embodiment, before the electronic device segments the first audio to obtain at least one second audio, the method further includes: the electronic device acquiring multiple third audio segments with a time length less than a fourth threshold; splicing the multiple third audio segments with a time length less than the fourth threshold to obtain the first audio, wherein there are silent segments among the various third audio segments in the first audio.

[0048] The fourth threshold is a pre-set threshold, which can also be determined based on the audio duration requirements of the audio recognition structure. For example, if the audio recognition structure requires the audio to be at least 30 seconds long, then the fourth threshold is 30 seconds.

[0049] Optionally, the duration of the first audio obtained after concatenation is greater than or equal to the fourth threshold. For example, if there are three third audios: third audio A (5 seconds), third audio B (15 seconds), and third audio C (14 seconds), concatenating these three third audios will yield the first audio (at least 34 seconds).

[0050] Optionally, the multiple third audio files to be spliced ​​can be audio files from the same speaker.

[0051] 102. The electronic device matches the audio features of the second audio with the audio features in the feature comparison library to obtain the feature number corresponding to the second audio. The feature comparison library includes the mapping relationship between audio features and feature numbers.

[0052] Optionally, the feature numbers in this feature comparison library are related to the speaker of the audio, with different feature numbers corresponding to different speakers. This feature comparison library includes audio features, feature numbers, and the mapping relationship between audio features and feature numbers. There is a one-to-one correspondence between audio features and feature numbers.

[0053] Optionally, before matching the audio features of the second audio with the audio features in the feature comparison library, the electronic device performs feature extraction on the second audio to obtain the audio features of the second audio.

[0054] Optionally, the feature extraction can be performed by sliding window detection on the second audio to extract its audio features. Alternatively, feature extraction can be performed after frame segmentation and windowing.

[0055] Optionally, features of the second audio can be extracted using methods such as short-time Fourier transform, deep neural networks, and convolutional neural networks. This application does not restrict the method used for feature extraction.

[0056] Optionally, the audio features of the second audio and the audio features in the feature comparison library are vectors of the same dimension. For example, the audio features of the second audio can be a 192-dimensional vector, and the audio features in the feature comparison library are also all 192-dimensional vectors.

[0057] Optionally, the audio features of the second audio files corresponding to the same speaker can be different. However, the similarity of the audio features of the second audio files corresponding to the same speaker is higher than that of the audio features of the second audio files corresponding to different speakers. For example, if second audio files A and B are from the same speaker, and second audio files A and C are from different speakers, the similarity between second audio files A and B is higher than the similarity between second audio files A and C.

[0058] Optionally, the feature comparison library includes pre-set audio features and feature numbers, as well as historically matched audio features and feature numbers.

[0059] In one possible embodiment, the electronic device matches the audio features of the second audio with audio features in a feature comparison library to obtain a feature number corresponding to the second audio. This includes: if the similarity between the first audio feature in the feature comparison library and the audio feature of the second audio is higher than a first threshold, then the electronic device uses the feature number corresponding to the first audio feature as the feature number corresponding to the second audio; if no audio feature in the feature comparison library has a similarity higher than the first threshold with the audio feature of the second audio, then the electronic device adds a new number as the feature number corresponding to the second audio.

[0060] The first threshold can be preset; for example, if the first threshold is 0.7, then the similarity between the audio features of the first audio feature and the audio features of the second audio feature is higher than 0.7. This newly added number is different from the feature numbers in the current feature comparison library.

[0061] Optionally, the new number can be added incrementally. For example, if the current feature numbers are role1, role2, and role3, then the new number would be role4.

[0062] For example, the audio feature comparison library includes: combine_wav1 (audio feature)-role1 (feature number), combine_wav2 (audio feature)-role2 (feature number), and combine_wav3 (audio feature)-role3 (feature number). If the similarity between the second audio and combine_wav2 is higher than 0.7, then the feature number corresponding to the second audio is role2. If the similarity between the second audio and combine_wav1, combine_wav2, and combine_wav3 is all less than 0.7, that is, if there is no audio feature in the feature comparison library with an audio feature similarity higher than 0.7 to the second audio, then a new number role4 will be added as the feature number corresponding to the second audio.

[0063] Optionally, the newly added number of the electronic device is used as the feature number corresponding to the second audio, including: the newly added number of the electronic device is used as the feature number corresponding to the second audio, and the audio features of the second audio and the feature number corresponding to the second audio are added to the feature comparison library.

[0064] In other words, the feature comparison library is dynamic. As the matching progresses, new audio features and their corresponding feature numbers will be added to the feature comparison library.

[0065] For example, the audio feature comparison library includes: combine_wav1 (audio feature)-role1 (feature number), combine_wav2 (audio feature)-role2 (feature number), and combine_wav3 (audio feature)-role3 (feature number). If there is no audio feature in the feature comparison library with an audio feature similarity higher than 0.7 to the second audio, a new number role4 will be added as the feature number corresponding to the second audio, and combine_wav4 (audio feature of the second audio)-role4 will be added to the audio feature comparison library.

[0066] Optionally, after feature extraction of the second audio, a temporary number is assigned to the second audio. After matching the temporary number with the feature comparison library, the temporary number is updated to the feature number.

[0067] For example, the following is combined with Figure 2B The feature comparison library will be further introduced.

[0068] First, the feature comparison library is initialized, which can be based on the first processed audio data, specifically the first audio file: wav1. wav1 contains two second audio files: wav1_role1 and wav1_role2. Since the feature comparison library is empty at this point (i.e., it lacks audio features and feature IDs), the mapping relationship between the audio features combine_wav1 and role1 (feature ID) of wav1_role1 is stored in the feature comparison library. Similarly, the mapping relationship between the audio features combine_wav2 and role2 (feature ID) of wav1_role2 is stored in the feature comparison library.

[0069] Then, when the first audio of the second phase arrives, that is, when processing wav2 (which includes two second audio files: wav2_role1 and wav2_role2), based on the previously initialized feature comparison library, the feature numbers corresponding to wav2_role1 and wav2_role2 are determined. Specifically, the audio features of wav2_role1 are matched with the two audio features (combine_wav1 and combine_wav2) in the feature comparison library. If the audio features of wav2_role1 match combine_wav1 successfully, then the feature number of combine_wav1 is used as the feature number of wav2_role1. If the audio features of wav2_role2 do not match either combine_wav1 or combine_wav2, then the newly added number role3 is used as the feature number of wav2_role2, that is, wav2_role2 (role2 in wav2_role2 is a temporary number) is updated to wav2_role3. Furthermore, the feature comparison library is updated, and the mapping relationship between the audio features of wav2_role2, combine_wav3 and role3, is stored in the feature comparison library.

[0070] Finally, when the first audio of the third phase arrives, that is, when processing wav3 (which includes three second audios: wav3_role1, wav3_role2, and wav3_role3), the feature numbers corresponding to wav3_role1, wav3_role2, and wav3_role3 are determined based on the updated feature comparison library. Specifically, the audio features of wav3_role1 are matched with the three audio features (combine_wav1, combine_wav2, and combine_wav3) in the feature comparison library. If the audio features of wav3_role1 match combine_wav2 successfully, then the feature number of combine_wav2 is used as the feature number of wav3_role1, that is, wav3_role1 (role1 in wav3_role1 is a temporary number) is updated to wav3_role2.

[0071] If the audio features of wav3_role2 do not match those of combine_wav1, combine_wav2, and combine_wav3, then the newly added number role4 will be used as the feature number for wav3_role2. In other words, wav3_role2 (where role2 is a temporary number) will be updated to wav3_role4. Furthermore, the feature comparison database will be updated, storing the mapping relationship between the audio features combine_wav4 and role4 of wav3_role2.

[0072] If the audio features of wav3_role3 do not match those of combine_wav1, combine_wav2, combine_wav3, and combine_wav4, then the newly added number role5 will be used as the feature number for wav3_role3. In other words, wav3_role3 (where role3 is a temporary number) will be updated to wav3_role5. Furthermore, the feature comparison database will be updated, storing the mapping relationship between the audio features combine_wav5 and role5 of wav3_role3.

[0073] In summary, Figure 2BInitialization in the first phase refers to adding the audio features from the first audio to the feature comparison database; updating the mapping table refers to updating the temporary number corresponding to the second audio to the feature number. The mapping table will only be updated if there is an audio feature that can be successfully matched in the feature comparison database; updating the feature comparison database refers to adding a new feature number as the feature number of the second audio. The feature comparison database will only be updated if there is no audio feature that can be successfully matched in the feature comparison database.

[0074] In one possible embodiment, the electronic device records the number of successful matches corresponding to each audio feature in the feature comparison library. The number of successful matches is the number of times that the audio features in the feature comparison library are matched with a similarity higher than a first threshold. The audio features of the second audio are matched with the audio features in the feature comparison library in descending order of the number of successful matches corresponding to each audio feature in the feature comparison library.

[0075] Optionally, the number of successful matches refers to the number of times an audio feature with a similarity higher than a first threshold is matched. For example, the feature comparison library includes the following two audio features: audio feature A and audio feature B. First, the audio feature of the second audio A matches audio feature A successfully (the similarity between the audio feature of the second audio A and audio feature A is higher than the first threshold); second, the audio feature of the second audio B matches audio feature A successfully; third, the audio feature of the second audio C matches audio feature A successfully; fourth, the audio feature of the second audio D matches audio feature B successfully. Then, the number of successful matches for audio feature A is 4, and the number of successful matches for audio feature B is 1.

[0076] Optionally, the second audio feature is prioritized for matching with audio features that have a high number of successful matches in the feature comparison library.

[0077] For example, the feature comparison library includes the following audio features: audio feature A (4 successful matches), audio feature B (20 successful matches), and audio feature C (16 successful matches). The feature comparison library, sorted by the number of successful matches, is: audio feature B, audio feature C, and audio feature A. The second audio file is first matched against audio feature B. If no match is found, the second audio file is then matched against audio feature C. This process continues.

[0078] In one possible embodiment, the electronic device matches the audio features of the second audio with the audio features in the feature comparison library in descending order of the number of successful matches corresponding to each audio feature. This includes: the electronic device matches the audio features of the second audio with the audio features in the first N audio features included in the feature comparison library in descending order of the number of successful matches corresponding to each audio feature in the feature comparison library, where N is an integer greater than 1.

[0079] N can be preset.

[0080] For example, if the feature comparison library includes 30 audio features, then the audio features are matched only with the first 10 audio features in the feature comparison library, in descending order of the number of successful matches for each audio feature.

[0081] In one possible embodiment, when the number of feature IDs in the feature comparison library is higher than a second threshold, audio features and feature IDs with fewer than a third threshold number of successful matches are removed from the feature comparison library.

[0082] Optionally, both the second and third thresholds are preset.

[0083] For example, the second threshold is 50 and the third threshold is 1. When the number of feature IDs in the feature comparison library is higher than 50, audio features and feature IDs that have never been successfully matched are removed from the feature comparison library.

[0084] Optionally, if the number of feature IDs in the removed feature comparison library is still higher than the second threshold, then audio features and feature IDs with fewer than five successful matches are removed from the feature comparison library. The fifth threshold is the third threshold + 1.

[0085] In other words, if the number of feature IDs in the feature comparison library is still higher than the second threshold after the first removal process, the third threshold is automatically incremented by 1 to obtain the fifth threshold, and the feature comparison library is removed again based on this fifth threshold. Similarly, if the number of feature IDs in the feature comparison library is still higher than the second threshold after the second removal process, the third threshold is increased, and so on.

[0086] Optionally, the feature comparison database can be removed based on the first cycle. That is, the electronic device can be triggered to remove features from the feature comparison database if the number of feature IDs in the database exceeds a second threshold, or it can be done automatically on a regular basis. This application does not impose any restrictions on this.

[0087] By using the above method, audio features in the feature comparison library are dynamically removed, avoiding an excessive number of audio features in the feature comparison library that would lead to low matching efficiency.

[0088] 103. The electronic device archives the second audio based on the feature number of the second audio. The second audio is used to train the speech model, and the speech model is used for speech synthesis.

[0089] Optionally, the electronic device may archive audio files with the same identifier in the same directory. For example, such as... Figure 2C As shown, all audio files under Spk_0 (wav1.wav, wav4.wav, wav10.wav, wav11.wav, etc.) have the same feature number.

[0090] Optionally, the speech model can be a personalized voice assistant model, which is used for personalized speech synthesis. This personalized speech synthesis can be found in the description above.

[0091] To better understand the audio data archiving method provided in this application, the application scenarios of personalized voice assistant models will be further introduced below.

[0092] The training objective of this personalized voice assistant model is to enable it to better extract audio features, that is, to better learn the features of different speakers, so that the extracted audio features can represent different speakers.

[0093] In application, this personalized voice assistant model can extract features from user-input audio. Then, based on these audio features, it combines them with text to synthesize corresponding audio, thereby achieving personalized speech synthesis.

[0094] Therefore, the focus of this personalized voice assistant model lies in the extraction of audio features. The more representative the extracted audio is to the speaker, the closer the personalized voice synthesized by the model will be to the speaker in the user-input audio. The audio data archiving method described in this application can accurately archive audio data. This allows the personalized voice assistant model to be better trained based on the archived audio, resulting in more accurate and representative audio features.

[0095] The hardware structure of electronic devices is described below:

[0096] Please see Figure 3 , Figure 3 This is a schematic diagram of the hardware structure of the electronic device 100 provided in the embodiments of this application.

[0097] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0098] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0099] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0100] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0101] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. The processor 110 retrieves the instructions or data stored in the memory, causing the electronic device 100 to execute the imaging method performed by the electronic device in the following method embodiments.

[0102] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0103] The charging management module 140 is used to receive charging input from the charger. The charger can be a wireless charger or a wired charger.

[0104] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc. In some other embodiments, the power management module 141 may also be located in the processor 110.

[0105] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0106] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0107] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0108] A modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor.

[0109] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), BLE broadcasting, global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0110] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with networks and other devices through wireless communication technology.

[0111] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0112] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. In some embodiments, electronic device 100 may include one or N displays screens 194, where N is a positive integer greater than 1. Display screen 194 may include OLED screens.

[0113] Optionally, the display 194 may further include: an OLED glass layer, an OLED light-emitting unit, a fingerprint recognition sensor, a microlens array, etc. The display 194 supports optical in-display fingerprint recognition.

[0114] Electronic device 100 can perform shooting functions through an ISP, camera 193, video codec, GPU, display screen 194, and application processor. The ISP processes data fed back by the camera 193. The camera 193 captures still images or videos. The camera 193 may include a front-facing camera and a rear-facing camera; the front-facing camera is located on the display area of ​​the screen, and the rear-facing camera is located on the back area of ​​the screen. The digital signal processor processes digital signals, including digital image signals and other digital signals. The video codec is used to compress or decompress digital video. Electronic device 100 may support one or more video codecs.

[0115] NPU stands for Neural-Network (NN) Computing Processor. By drawing inspiration from the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can quickly process input information and continuously learn on its own.

[0116] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage functions.

[0117] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as a sound playback function), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data), etc. Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as flash memory devices.

[0118] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0119] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0120] A speaker 170A, also called a "loudspeaker," is used to convert audio electrical signals into sound signals. A receiver 170B, also called a "handpiece," is used to convert audio electrical signals into sound signals. A microphone 170C, also called a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. A headphone jack 170D is used to connect wired headphones. A pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, the pressure sensor 180A may be located on the display screen 194. A gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. A barometric pressure sensor 180C is used to measure barometric pressure. A magnetic sensor 180D includes a Hall effect sensor. An accelerometer 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). A distance sensor 180F is used to measure distance. A proximity sensor 180G may include, for example, a light-emitting diode (LED) and a photosensor. An ambient light sensor 180L is used to sense ambient light intensity. A fingerprint sensor 180H is used to collect fingerprints. Temperature sensor 180J is used to detect temperature. Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. Bone conduction sensor 180M can acquire vibration signals. Buttons 190 include power button, volume buttons, etc. Motor 191 can generate vibration prompts. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card.

[0121] Furthermore, an operating system runs on top of the aforementioned components. Examples include iOS and Android. The operating system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100. It should be noted that although this application embodiment uses the Android system as an example for illustration, its basic principles are equally applicable to electronic devices with other operating systems.

[0122] Figure 4This is a schematic diagram of the software structure of an electronic device 100 provided in an embodiment of this application. The software structure adopts a layered architecture, which divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In this embodiment, the operating system (taking Android as an example, with Android running on an AP) can be divided into three layers, from top to bottom: the application layer (APP), the application framework layer (framework, FWK), and the kernel layer.

[0123] The application layer can include a series of application packages. For example, such as... Figure 4 As shown, the application package may include personalized voice assistants, etc. The electronic device may also include more applications, such as news applications, map applications, etc. This embodiment uses a personalized voice assistant as an example.

[0124] The application framework layer provides application developers with an application programming interface (API) framework and various services and management tools to access core functionalities, including interface management, data access, application-layer messaging, application package management, telephony management, and location management. The application framework layer includes some predefined functions. For example, ... Figure 4 As shown, the application framework layer may include, but is not limited to, a content playback module (TextToSpeech, TTS) and a speech synthesis management module (UtteranceProgressListener). This TTS may include a personalized speech synthesis model. Wherein:

[0125] TextToSpeech is an important technology for realizing text-to-speech functionality. It enables devices to express text information in the form of speech. TextToSpeech provides functions such as adjusting the tone of pronunciation and presets.

[0126] UtteranceProgressListener is used to listen for and process events in the TextToSpeech engine to improve speech synthesis efficiency. Through timely feedback, UtteranceProgressListener allows developers to take appropriate actions based on different stages of speech synthesis (onStart, onDone, errors). Developers can accurately grasp the speech synthesis status of each text segment based on the information obtained from callback methods, thereby gaining better control over the application's behavior.

[0127] The kernel layer is the layer between hardware and software. At a minimum, it includes speaker drivers, Bluetooth drivers, etc. Synthesized speech can be played directly from the electronic device's speaker, or it can be played from other devices via Bluetooth, such as Bluetooth headsets or Bluetooth speakers.

[0128] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an audio data archiving device 500 provided in an embodiment of this application. Figure 5 The audio data archiving device shown can be an electronic device, a device within an electronic device, or a device that can be used in conjunction with an electronic device. Figure 5 The audio data archiving device shown may include a processing unit 501 and a matching unit 502.

[0129] in:

[0130] The processing unit 501 is used to segment the first audio to obtain at least one second audio, wherein the second audio is an audio segment of the first audio that contains human voice;

[0131] The matching unit 502 is used to match the audio features of the second audio with the audio features in the feature comparison library to obtain the feature number corresponding to the second audio. The feature comparison library includes the mapping relationship between audio features and feature numbers.

[0132] The processing unit 501 is also used to archive the second audio based on the feature number of the second audio, the second audio is used to train the speech model, and the speech model is used for speech synthesis.

[0133] In one possible implementation, the processing unit 501 is further configured to, if the similarity between the first audio feature in the feature comparison library and the audio feature of the second audio is higher than a first threshold, use the feature number corresponding to the first audio feature as the feature number corresponding to the second audio.

[0134] The processing unit 501 is further configured to, if there is no audio feature in the feature comparison library whose similarity to the audio feature of the second audio is higher than the first threshold, add a new number as the feature number corresponding to the second audio.

[0135] In one possible implementation, the processing unit 501 is further configured to add a new number as the feature number corresponding to the second audio, and add the audio features of the second audio and the feature number corresponding to the second audio to the feature comparison library.

[0136] In one possible implementation, the processing unit 501 is also used to record the number of successful matches corresponding to each audio feature in the feature comparison library. The number of successful matches is the number of times that the audio features in the feature comparison library are matched with a similarity higher than a first threshold.

[0137] The matching unit 502 is also used to match the audio features of the second audio with the audio features in the feature comparison library in descending order of the number of successful matches corresponding to each audio feature in the feature comparison library.

[0138] In one possible implementation, the matching unit 502 is further configured to match the audio features of the second audio with the audio features in the first N audio features included in the feature comparison library, in descending order of the number of successful matches corresponding to each audio feature in the feature comparison library, where N is an integer greater than 1.

[0139] In one possible implementation, the processing unit 501 is further configured to remove audio features and feature numbers with fewer than a third threshold from the feature comparison library when the number of feature numbers in the feature comparison library is higher than a second threshold.

[0140] In one possible implementation, the processing unit 501 is further configured to acquire multiple third audio segments with a duration less than a fourth threshold; and to splice the multiple third audio segments with a duration less than the fourth threshold to obtain a first audio segment, wherein there are silent segments among the third audio segments in the first audio segment.

[0141] For cases where the audio data archiving device can be a chip or a chip system, see [link to relevant documentation]. Figure 6 The diagram shows the structure of the chip. Figure 6 The chip 600 shown includes a processor 601 and an interface 602. Optionally, it may also include a memory 603. The number of processors 601 can be one or more, and the number of interfaces 602 can be multiple.

[0142] For cases where the chip is used to implement the electronic device in the embodiments of this application:

[0143] The interface 602 is used to receive or output signals;

[0144] The processor 601 is used to perform data processing operations of the electronic device.

[0145] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0146] It is understood that some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current solution on which they are based, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in certain scenarios. Accordingly, the audio data archiving device given in the embodiments of this application can also implement these features or functions, which will not be elaborated here.

[0147] It should be understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0148] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0149] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed on an electronic device, implement the functions of any of the above method embodiments.

[0150] This application also provides a computer program product that, when run on a computer, enables the computer to perform the functions of any of the above method embodiments.

[0151] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0152] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An audio data archiving method, characterized in that, The method includes: The first audio is segmented to obtain at least one second audio, wherein the second audio is an audio segment of the first audio that contains human voice; The audio features of the second audio are matched with the audio features in the feature comparison library to obtain the feature number corresponding to the second audio. The feature comparison library includes the mapping relationship between audio features and feature numbers. The feature numbers in the feature comparison library are related to the speaker of the audio, and different feature numbers correspond to different speakers. Based on the feature number of the second audio, the second audio is archived, and the second audio is used to train a speech model, which is used for speech synthesis. Record the number of successful matches for each audio feature in the feature comparison library. The number of successful matches is the number of times that the audio feature in the feature comparison library is matched with a similarity higher than a first threshold. The audio features of the second audio are matched with the audio features in the feature comparison library in descending order of the number of successful matches for each audio feature in the feature comparison library. When the number of feature IDs in the feature comparison library is higher than the second threshold, audio features and feature IDs with fewer than the third threshold of successful matches are removed from the feature comparison library. If the number of feature IDs in the removed feature comparison library is still higher than the second threshold, then audio features and feature IDs with fewer than five successful matches will be removed from the feature comparison library. The fifth threshold is the third threshold plus 1.

2. The method according to claim 1, characterized in that, The step of matching the audio features of the second audio with the audio features in the feature comparison library to obtain the feature number corresponding to the second audio includes: If the similarity between the first audio feature in the feature comparison library and the audio feature of the second audio is higher than the first threshold, then the feature number corresponding to the first audio feature is used as the feature number corresponding to the second audio. If no audio feature in the feature comparison library has a similarity to the audio feature of the second audio that is higher than the first threshold, then a new number is added as the feature number corresponding to the second audio.

3. The method according to claim 2, characterized in that, The newly added number serves as the feature number corresponding to the second audio, including: The newly added number is used as the feature number corresponding to the second audio, and the audio features of the second audio and the feature number corresponding to the second audio are added to the feature comparison library.

4. The method according to claim 3, characterized in that, The step of matching the audio features of the second audio with the audio features in the feature comparison library according to the order of the number of successful matches for each audio feature in the feature comparison library from largest to smallest includes: According to the order of the number of successful matches corresponding to each audio feature in the feature comparison library from largest to smallest, the audio features of the second audio are matched with the audio features in the first N audio features included in the feature comparison library, where N is an integer greater than 1.

5. The method according to any one of claims 1-4, characterized in that, Before segmenting the first audio to obtain at least one second audio, the method further includes: Retrieve multiple third audio clips with a duration shorter than the fourth threshold; The first audio is obtained by splicing together the multiple third audio segments with a time length less than the fourth threshold, and there are silent segments among the various third audio segments in the first audio.

6. An electronic device comprising one or more memories and one or more processors, characterized in that, The memory is used to store a computer program; the processor is used to invoke the computer program to cause the electronic device to perform the method of any one of claims 1-5.

7. A chip system for use in electronic devices, characterized in that, The chip system includes at least one processor and an interface for receiving instructions and transmitting them to the at least one processor; the at least one processor executes the instructions to cause the electronic device to perform the method as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1-5.

9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Voice data processing method and device

    CN110931013A

  • Audio training data processing method and device, equipment and storage medium

    CN112614478A

  • Speech recognition method and device

    CN116153291A