Audio generation methods, devices, electronic devices, and storage media
By generating a mixture of virtual and real audio in the metaverse platform and using registration random code encoding, the problem of limited user voice settings is solved, the voiceprint recognition rate and the accuracy of identity verification are improved, and user privacy is protected.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VOICEAI TECH CO LTD
- Filing Date
- 2022-12-21
- Publication Date
- 2026-06-30
AI Technical Summary
In the Metaverse platform, user-defined voice settings are limited, which may result in different users having the same voice, affecting the accuracy of voiceprint verification and the convenience of interaction, and also posing security risks.
By obtaining virtual voice parameters and real audio from registered identity information, a mixed audio of virtual and real audio is generated, which is then encoded using a registration random code to increase voice diversity, and the recognition rate is improved through voiceprint verification.
It improves the diversity of audio in the metaverse and the voiceprint recognition rate, protects users' physiological privacy, and enhances the accuracy and security of identity verification.
Smart Images

Figure CN116030818B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and more specifically, to an audio generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, speech processing technology is applied in various fields. For example, it can be applied to the metaverse. The metaverse is a virtual world constructed by humans using digital technology, which maps to or transcends the real world, allows interaction with it, and provides a digital living space with a new social system. With the development of metaverse platforms, related technologies have increasingly diverse requirements for audio within the metaverse. Summary of the Invention
[0003] In view of the above problems, this application proposes an audio generation method, apparatus, electronic device and storage medium, which can generate mixed audio through dual audio, increasing the diversity of sound and improving the recognition rate of mixed voiceprints.
[0004] In a first aspect, embodiments of this application provide an audio generation method, the method comprising: obtaining registration identity information and obtaining virtual sound parameters corresponding to the registration identity information; obtaining virtual audio corresponding to the registration identity information based on the virtual sound parameters; obtaining real audio corresponding to the registration identity information; obtaining a registration random code corresponding to the registration identity information based on the virtual audio and the real audio; obtaining a target registration audio based on the registration random code, the virtual audio, and the real audio; and storing the target registration audio in association with the registration identity information.
[0005] Secondly, embodiments of this application provide an audio generation apparatus, comprising: a registration identity information acquisition module, a virtual audio acquisition module, a real audio acquisition module, a registration random code acquisition module, a target registration audio acquisition module, and a storage module. The registration identity information acquisition module is used to acquire registration identity information and virtual sound parameters corresponding to the registration identity information; the virtual audio acquisition module is used to acquire virtual audio corresponding to the registration identity information based on the virtual sound parameters; the real audio acquisition module is used to acquire real audio corresponding to the registration identity information; the registration random code acquisition module is used to acquire a registration random code corresponding to the registration identity information based on the virtual audio and the real audio; the target registration audio acquisition module is used to acquire target registration audio based on the registration random code, the virtual audio, and the real audio; and the storage module is used to associate and store the target registration audio with the registration identity information.
[0006] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory is coupled to the processor, the memory stores instructions, and when the instructions are executed by the processor, the processor performs the above-described method.
[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing program code, which can be invoked by a processor to execute the above-described method.
[0008] The audio generation method, apparatus, electronic device, and storage medium provided in this application obtain registration identity information and virtual sound parameters corresponding to the registration identity information. Based on the virtual sound parameters, they obtain virtual audio corresponding to the registration identity information and real audio corresponding to the registration identity information. Based on the virtual audio and real audio, they obtain a registration random code corresponding to the registration identity information. Based on the registration random code, virtual audio, and real audio, they obtain a target registration audio. They then associate and store the target registration audio with the registration identity information. This generates mixed audio through dual audio, increasing the diversity of sound and improving the recognition rate of mixed voiceprints. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A schematic flowchart of an audio generation method provided in an embodiment of this application is shown;
[0011] Figure 2 A schematic flowchart of an audio generation method provided in an embodiment of this application is shown;
[0012] Figure 3 A schematic flowchart of an audio generation method provided in an embodiment of this application is shown;
[0013] Figure 4 A schematic flowchart of an audio generation method provided in an embodiment of this application is shown;
[0014] Figure 5 A schematic flowchart of an audio generation method provided in an embodiment of this application is shown;
[0015] Figure 6 A schematic flowchart of an audio generation method provided in an embodiment of this application is shown;
[0016] Figure 7 A schematic flowchart of an audio generation method provided in an embodiment of this application is shown;
[0017] Figure 8 A block diagram of an audio generation apparatus according to an embodiment of this application is shown;
[0018] Figure 9 A block diagram of an electronic device for performing an audio generation method according to an embodiment of this application is shown;
[0019] Figure 10 An embodiment of this application shows a storage unit for storing or carrying program code that implements the audio generation method according to an embodiment of this application. Detailed Implementation
[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0021] Among related technologies, Metaverse is a new type of internet application and social form that integrates multiple new technologies, blending the virtual and real worlds. It provides immersive experiences based on extended reality technology and generates a mirror image of the real world using digital twin technology. It builds an economic system through blockchain technology, closely integrating the virtual and real worlds in economic, social, and identity systems, and allowing each user to produce and edit content. Currently, Metaverse is demonstrating strong development potential, and the Metaverse platform is nurturing numerous cutting-edge technologies. Among these, the realistic experience pursued by Metaverse inevitably includes voice communication and interaction. In Metaverse, users can customize their transformed voices, freely setting gender, accent, style, etc.; however, the available transformation settings are limited, making it possible for different users to have the same voice. This situation is detrimental to the promotion of voiceprint verification, the convenience of Metaverse interaction, and also provides opportunities for criminals.
[0022] Therefore, with the development of metaverse platforms, related technologies require a diversity of audio content within the metaverse.
[0023] To address the aforementioned problems, the inventors, through long-term research, discovered and proposed the audio generation method, apparatus, electronic device, and storage medium provided in the embodiments of this application. By generating mixed audio from dual audio sources, the diversity of sounds is increased, and the recognition rate of mixed voiceprints is improved. The specific audio generation method will be described in detail in subsequent embodiments.
[0024] Please see Figure 1 , Figure 1A schematic flowchart of an audio generation method according to an embodiment of this application is shown. This audio generation method generates mixed audio through dual audio, increasing the diversity of sounds and improving the recognition rate of mixed voiceprints. In a specific embodiment, this audio generation method can be applied to, for example... Figure 8 The audio generating device 200 and the electronic device 100 configured with the audio generating device 200 are shown. Figure 9 The following will use an electronic device as an example to illustrate the specific process of this embodiment. Of course, it is understood that the electronic device used in this embodiment may include smartphones, tablets, wearable electronic devices, etc., and is not limited thereto. The following will focus on... Figure 1 The process shown will be explained in detail. The audio generation method may specifically include the following steps:
[0025] Step S110: Obtain registration identity information and obtain the virtual voice parameters corresponding to the registration identity information.
[0026] In some implementations, the electronic device may include a camera, which can be used to capture a facial image of a user operating the electronic device. The electronic device can then obtain the user's identity information based on the facial image and use the identity information as registration identity information.
[0027] In some implementations, the electronic device may include a screen, on which an authentication interface can be generated and displayed. Optionally, a processor in the electronic device may detect presses on the screen to obtain registration identity information entered by the user based on the authentication interface; the processor in the electronic device may also detect button presses on the electronic device to obtain registration identity information entered by the user based on the authentication interface.
[0028] In some implementations, the operating system of the electronic device may include a metaverse system. The camera in the electronic device can capture a facial image of the user operating the device, and further, the electronic device can obtain the user's identity information based on the facial image. The electronic device can also obtain the registration identity information entered by the user through an authentication interface generated by the electronic device. Furthermore, the electronic device can match the user's identity information obtained from the facial image with the registration identity information entered by the user through the authentication interface. If the match is successful, meaning the user's identity information obtained from the facial image is the same as the registration identity information entered by the user through the authentication interface, then the user is determined to have passed authentication, and the metaverse system in the electronic device can obtain the registration identity information.
[0029] Furthermore, after obtaining the registered identity information, the electronic device can generate a virtual voice parameter setting interface and display it on the screen. The electronic device can detect screen presses to obtain the virtual voice parameters input by the user based on the virtual voice parameter setting interface; the processor in the electronic device can also detect button presses to obtain the virtual voice parameters input by the user based on the virtual voice parameter setting interface. In some embodiments, after obtaining the registered identity information, the electronic device can generate an interface for setting the virtual voice parameters of the virtual character corresponding to the registered identity information in the metaverse system. These virtual voice parameters may include the virtual character's gender, timbre, volume, etc.
[0030] Step S120: Obtain the virtual audio corresponding to the registered identity information based on the virtual sound parameters.
[0031] In some implementations, after acquiring virtual sound parameters, the electronic device can obtain virtual audio corresponding to the registered identity information based on the virtual sound parameters. Specifically, the electronic device's metaverse system can generate virtual audio based on these virtual parameters and a first preset content.
[0032] Step S130: Obtain the real audio corresponding to the registered identity information.
[0033] In some implementations, the electronic device may include a microphone, and the processor within the electronic device can acquire the audio captured by the microphone. Specifically, after acquiring registration identity information, the electronic device can use the human voice captured by the microphone as the actual audio corresponding to that registration identity information.
[0034] Optionally, after obtaining the registered identity information, the electronic device can generate and display a voice input interface; wherein the voice input interface may include a second preset content. Further, the electronic device can acquire the voice input by the user based on the voice input interface, including the second preset content, as the actual audio corresponding to the registered identity information. That is, in some embodiments, the electronic device can acquire the voice input by the user based on the second target content corresponding to the registered identity information, and determine the actual audio corresponding to the registered identity information based on the voice. Wherein, after acquiring the voice, the electronic device can perform noise reduction processing on the voice, or perform human voice extraction processing on the voice, thereby obtaining the actual audio corresponding to the registered identity information.
[0035] In some implementations, after obtaining the registered identity information, the electronic device can obtain the real audio corresponding to the registered identity information from the associated cloud or electronic device through wireless communication technology (such as Bluetooth, WiFi, Zigbee, etc.); or it can obtain the real audio corresponding to the registered identity information from the associated electronic device through a serial communication interface (such as a serial peripheral interface, etc.).
[0036] Virtual audio and real audio can be acquired simultaneously or separately. The text content in the virtual audio can be the same as or different from the text content in the real audio. That is, the first preset content can be the same as or different from the second preset content. Simultaneously acquired virtual audio and real audio can be defined as audio with the same text content obtained through simultaneous acquisition. Separately acquired virtual audio and real audio can be defined as audio obtained through point-to-point acquisition, where the text content of the virtual audio and real audio obtained through point-to-point acquisition can be the same as or different.
[0037] Step S140: Obtain a registration random code corresponding to the registration identity information based on the virtual audio and the real audio.
[0038] In some implementations, after obtaining virtual audio, the electronic device can perform word segmentation on the virtual audio to obtain a first word segment; after obtaining real audio, the electronic device can perform word segmentation on the real audio to obtain a second word segment. Furthermore, the electronic device can obtain a registration random code corresponding to the registered identity information based on the number of the first word segment and the number of the second word segment.
[0039] In some implementations, the electronic device may have a pre-set word segmentation model that can perform word segmentation processing on audio. This word segmentation model may be a Hidden Markov Model (HMM), a Recurrent Neural Network (RNN) model, a Long Short-Term Memory (LSTM) model, or the like.
[0040] Optionally, the electronic device may generate a random code with the same number of bits as the number of the first word segment as the registration random code corresponding to the registration identity information; the electronic device may also generate a random code with the same number of bits as the number of the second word segment as the registration random code corresponding to the registration identity information; the electronic device may also generate a random code with the same number of bits as the sum of the number of the first word segment and the number of the second word segment as the registration random code corresponding to the registration identity information.
[0041] Step S150: Obtain the target registration audio based on the registration random code, the virtual audio, and the real audio.
[0042] In some implementations, after obtaining a registration random code, the electronic device can obtain a target registration audio based on the registration random code, virtual audio, and real audio. Specifically, the electronic device can obtain the target registration audio by filling the virtual audio encoding bits of the registration random code with virtual audio and filling the real audio encoding bits of the registration random code with real audio, wherein the sum of the number of virtual audio encoding bits and the number of real audio encoding bits of the registration random code is the same as the number of bits in the registration random code.
[0043] For example, the virtual audio obtained by the electronic device is A1, and the real human audio is A2. Furthermore, the electronic device can perform word segmentation processing on the virtual audio A1 and the real human audio A2 based on a word segmentation model to obtain a set of speech segments corresponding to the virtual audio, that is, the first segmented word segment. Obtain the set of speech segments corresponding to the actual audio, that is, the second word segment. The virtual audio encoding bits of the registration random code can be an odd number of bits, while the real audio encoding bits can be an even number of bits. The number of bits in the registration random code is the same as the number of the first and second word segments. The electronic device can then use the first word segment... The virtual audio encoding bits that fill the registration random code, the second word segment The real audio encoding bits of the registration random code are filled to obtain the target registration audio A. Then, by interleaving and splicing the two audio segments, a mixed audio is generated, which increases the diversity of the sound and improves the recognition rate of the mixed voiceprint.
[0044] Step S160: Associate and store the target registered audio with the registered identity information.
[0045] In some implementations, after obtaining the target registration audio, the electronic device can associate and store the target registration audio with the registration identity information. Optionally, after obtaining the target registration audio, the electronic device can register a voiceprint corresponding to the registration identity information based on the target registration audio, and associate and store the target registration audio with the registration identity information, registration random code, real audio, and virtual audio, so that the electronic device can perform voiceprint verification based on the target registration audio.
[0046] For example, please refer to Figure 2This document illustrates a flowchart of an audio generation method provided in an embodiment of this application. The electronic device may include a metaverse system. The electronic device can acquire registration identity information input by the user for authentication while operating the metaverse system, and acquire virtual voice parameters of the virtual character corresponding to the registration identity information in the metaverse system, set by the user. Further, the electronic device acquires the registration identity information and the virtual voice parameters corresponding to the registration identity information, and can register the identity of the virtual character using audio. The metaverse system in the electronic device can obtain virtual audio A1 corresponding to the registration identity information based on a first preset content and the virtual voice parameters; it can also acquire the voice input by the user based on a second preset content, and determine the real audio A2 corresponding to the registration identity information based on the voice; wherein the first preset content and the second preset content are the same. Further, the electronic device can perform word segmentation processing on the virtual audio A1 and the real audio A2 based on a word segmentation model to obtain a set of voice segments corresponding to the two audio segments, which are respectively the first word segmentation segments. Second participle fragment Furthermore, the metaverse system of electronic devices can be based on the first word segment. The quantity, and the second participle segment The number of bits generated and the first segmented word fragment The same number of random codes are used as registration random codes. This can be understood as the first preset content being the same as the second preset content, i.e., the first word segment. Quantity and second participle The number is the same. Furthermore, the electronic device can sequentially fill the odd-numbered bits with the corresponding first word segment based on the parity of the registered random code. Fill the even-numbered positions with the corresponding second word segment. This allows for the interleaving and splicing of two audio segments to generate the target registered audio, also known as mixed audio A.
[0047] Furthermore, electronic devices can register the voiceprint of the target registered audio with the voiceprint of the user corresponding to the registered identity information, and associate and store the target registered audio with the registered identity information, the registration random code, the virtual audio and the real audio, so as to perform voiceprint verification based on the target registered audio.
[0048] Understandably, mixing real and virtual audio to obtain a mixed voiceprint effectively protects the physiological privacy of users' real voices, avoids the leakage of personal information, and adds real voice components to the target registration audio, increasing the diversity of sounds, improving the accuracy of audio-based identity verification in the metaverse, and enhancing the user experience.
[0049] An embodiment of this application provides an audio generation method that obtains registration identity information and virtual sound parameters corresponding to the registration identity information. Based on the virtual sound parameters, it obtains virtual audio corresponding to the registration identity information and real audio corresponding to the registration identity information. Based on the virtual audio and real audio, it obtains a registration random code corresponding to the registration identity information. Based on the registration random code, virtual audio, and real audio, it obtains a target registration audio. It then associates and stores the target registration audio with the registration identity information. In this way, it generates mixed audio through dual audio, which increases the diversity of sound and improves the recognition rate of mixed voiceprints.
[0050] Please see Figure 3 , Figure 3 A schematic flowchart of an audio generation method according to an embodiment of this application is shown. This method is applied to the aforementioned electronic device, and will be discussed below. Figure 3 The process shown will be explained in detail. The audio generation method may specifically include the following steps:
[0051] Step S210: Obtain the registration identity information and the virtual voice parameters corresponding to the registration identity information.
[0052] Step S220: Obtain the virtual audio corresponding to the registered identity information based on the virtual sound parameters.
[0053] Step S230: Obtain the real audio corresponding to the registered identity information.
[0054] For a detailed description of steps S210-S230, please refer to the previous description of steps S110-S130, which will not be repeated here.
[0055] Step S240: Obtain a registration random code corresponding to the registration identity information based on the virtual audio and the real audio.
[0056] In some implementations, after acquiring virtual and real audio, the electronic device can obtain a registration random code corresponding to the registered identity information based on the virtual and real audio. Specifically, the metaverse system within the electronic device can generate a random code corresponding to the registered identity information as the registration random code based on the virtual and real audio.
[0057] For example, if virtual audio and real audio are acquired simultaneously, meaning the text content of the virtual audio is the same as the text content of the real audio, the metaverse system of the electronic device can further generate a random code with a length (in bits) equal to the number of segmented segments obtained after segmenting the virtual audio, as a registration random code. It is understood that the text content of the virtual audio is the same as the text content of the real audio, and the number of segmented segments obtained after segmenting the virtual audio is the same as the number of segmented segments obtained after segmenting the real audio.
[0058] For example, if virtual audio and real audio are obtained through point-based acquisition, and the text content of the virtual audio differs from that of the real audio, the electronic device's metaverse system can further generate a random code with a length (in bits) equal to or equal to the number of segmented segments obtained from processing the virtual audio, as a registration random code. It is understood that the text content of the virtual audio differs from that of the real audio, and the number of segmented segments obtained from processing the virtual audio differs from the number of segmented segments obtained from processing the real audio.
[0059] Specifically, if the text content of the virtual audio is longer than the text content of the real audio, the electronic device's metaverse system can generate a random code with a length (in bits) equal to the number of segmented segments obtained after segmenting the virtual audio, as the registration random code; or if the text content of the virtual audio is shorter than the text content of the real audio, the electronic device's metaverse system can generate a random code with a length (in bits) equal to the number of segmented segments obtained after segmenting the real audio, as the registration random code. In other words, the electronic device can set the length (in bits) of the registration random code to be the same as the number of segmented segments of the longest text.
[0060] In some implementations, step S240 may include steps S241-S243, or steps S241, S242 and S244.
[0061] Step S241: Perform word segmentation on the virtual audio based on a preset word segmentation model to obtain the first word segment.
[0062] In some implementations, the electronic device may have a pre-set word segmentation model, which may be an HMM model, an RNN model, an LSTM model, etc. Furthermore, after obtaining virtual audio, the electronic device can perform word segmentation processing on the virtual audio based on the pre-set word segmentation model to obtain the corresponding word segment, i.e., the first word segment; wherein the first word segment may include one or more word segments, that is, the number of first word segments can be one or more.
[0063] Step S242: Perform word segmentation on the real audio based on the preset word segmentation model to obtain a second word segment.
[0064] In some implementations, after obtaining real audio, the electronic device can perform word segmentation on the real audio based on a preset word segmentation model to obtain the word segmentation fragment corresponding to the real audio, namely the second word segmentation fragment; wherein, the second word segmentation fragment may include one or more word segmentation fragments, that is, the number of second word segmentation fragments may be one or more.
[0065] Step S243: If the number of the first word segment is greater than or equal to the number of the second word segment, then generate a random code with the same number of bits as the number of the first word segment as the registration random code corresponding to the registration identity information.
[0066] In some implementations, after the electronic device obtains the first word segment and the second word segment, it can compare the number of the first word segment with the number of the second word segment. If the number of the first word segment is greater than or equal to the number of the second word segment, a random code with the same number of bits as the number of the first word segment is generated as the registration random code corresponding to the registration identity information.
[0067] Step S244: If the number of the second word segment is greater than or equal to the number of the first word segment, then generate a random code with the same number of bits as the number of the second word segment as the registration random code corresponding to the registration identity information.
[0068] In some implementations, after the electronic device obtains the first word segment and the second word segment, it can compare the number of the first word segment with the number of the second word segment. If the number of the second word segment is greater than or equal to the number of the first word segment, a random code with the same number of bits as the number of the second word segment is generated as the registration random code corresponding to the registration identity information.
[0069] Step S250: Obtain the target registration audio based on the registration random code, the virtual audio, and the real audio.
[0070] In some implementations, after obtaining a registration random code, virtual audio, and real audio, the electronic device can obtain the target registration audio based on the registration random code, virtual audio, and real audio. Specifically, the electronic device can fill the virtual encoding bits of the registration random code with the first segmented word fragment corresponding to the virtual audio, and fill the real encoding bits of the registration random code with the second segmented word fragment corresponding to the real audio, thus obtaining the target registration audio.
[0071] In some implementations, the text content of the virtual audio is the same as the text content of the real audio, and the number of the first segmented segments is the same as the number of the second segmented segments. Furthermore, the electronic device obtains a registration random code with the same number of bits as either the number of the first or the number of the second segmented segments, based on the number of the first and second segmented segments.
[0072] Furthermore, the electronic device can sequentially use the odd-numbered bits of the registered random code as virtual encoding bits and the even-numbered bits as real encoding bits, based on the parity of the registered random code. Further, it can use the odd-numbered bits of the first word segment in the virtual audio as the audio to fill the virtual encoding bits of the registered random code, and the even-numbered bits of the second word segment in the real audio as the audio to fill the real encoding bits of the registered random code. This allows the two segments to be interleaved and spliced together to generate a mixed audio, i.e., the target registered audio.
[0073] In some implementations, virtual audio and real audio are acquired simultaneously. The electronic device can directly replace the virtual encoding bit of the corresponding registration random code of the real audio with the virtual audio to obtain the target registration audio; the electronic device can also directly replace the real encoding bit of the corresponding registration random code of the virtual audio with the real audio to obtain the target registration audio.
[0074] In some implementations, please refer to Figure 4 Step S250 may include steps S2510 and S2530.
[0075] Step S2510: If the number of bits in the registered random code is greater than the number of the first word segment, then supplement the number of the first word segment to be the same as the number of bits in the registered random code, and obtain a new first word segment.
[0076] In some implementations, the number of bits in the registered random code is the same as the number of word segments in the longest text segment in both the virtual and real audio. If the number of bits in the registered random code is greater than the number of the first word segments, the electronic device can supplement the number of the first word segments to match the number of bits in the registered random code, thus obtaining a new first word segment.
[0077] Specifically, the electronic device supplements the number of first word segments to match the number of bits in the registered random code. Obtaining a new first word segment can be achieved by calculating the difference between the number of bits in the registered random code and the number of first word segments. Further, the electronic device can randomly select a target word segment from the first word segments, matching the difference in number, and concatenate this target word segment with the first word segments to obtain a new first word segment. The target word segment can be concatenated at the beginning, end, or middle of the virtual audio text.
[0078] In some implementations, please refer to Figure 5 Step S2510 may include steps S2511-S2513.
[0079] Step S2511: If the number of bits in the registration random code is greater than the number of the first word segment, then obtain the difference between the number of bits in the registration random code and the number of the first word segment.
[0080] In some implementations, after the electronic device obtains the registration random code and the first word segment, it can compare the number of bits in the registration random code with the number of the first word segment. If the number of bits in the registration random code is greater than the number of the first word segment, the difference between the number of bits in the registration random code and the number of the first word segment is obtained.
[0081] Step S2512: According to the text of the first segmented segment from front to back, obtain the target segmented segment with the same number of differences from the first segmented segment.
[0082] Furthermore, after the electronic device obtains the difference between the number of digits of the registration random code and the number of the first segmented fragments, it can extract the target segmented fragments from the first segmented fragments in the order of the text of the first segmented fragment from beginning to end, with the number of segments equal to the difference.
[0083] Step S2513: Add the target word segment to the end of the first word segment to obtain the new first word segment.
[0084] Furthermore, after obtaining the target word segment, the electronic device can add the target word segment to the end of the first word segment to obtain a new first word segment.
[0085] Understandably, if virtual audio and real audio are obtained through point-to-point acquisition, and the text of the virtual audio is different from that of the real audio, the electronic device can copy and concatenate the word segments corresponding to the shortest audio segments in the virtual and real audio, so that the number of word segments corresponding to the shortest audio segment is the same as the number of word segments corresponding to the longest audio segment. In this way, the corresponding audio segments of the two audio segments can be replaced to obtain mixed audio.
[0086] Step S2520: Fill the first encoding bit of the registration random code based on the new first word segment, and fill the second encoding bit of the registration random code based on the second word segment to obtain the target registered audio, wherein the sum of the number of the first encoding bit and the number of the second encoding bit is equal to the number of bits of the registration random code.
[0087] In some implementations, after obtaining a new first word segment, the electronic device can fill in the first encoding bits of the registration random code based on the new first word segment, and fill in the second encoding bits of the registration random code based on the second word segment, to obtain the target registered audio. The number of new first word segments is the same as the number of bits in the registration random code, and also the same as the number of second word segments. The first encoding bit can be understood as the encoding bit of the corresponding virtual audio of the registration random code, i.e., a virtual encoding bit; the second encoding bit can be understood as the encoding bit of the corresponding real audio of the registration random code, i.e., a real encoding bit; the sum of the number of first encoding bits and the number of second encoding bits equals the number of bits in the registration random code.
[0088] In some implementations, the electronic device can sequentially use the odd and even bits of the registered random code, with the odd bits as the first encoding bits corresponding to the new first word segment and the even bits as the second encoding bits corresponding to the second word segment, to achieve the interleaving and splicing of the two segments to generate mixed audio, that is, the target registered audio.
[0089] In some implementations, see 6, step S250 may include steps S2530 and S2540.
[0090] Step S2530: If the number of bits in the registered random code is greater than the number of the second word segment, then the number of the second word segment is increased to be the same as the number of bits in the registered random code, and a new second word segment is obtained.
[0091] After obtaining the registration random code and the second word segment, the electronic device can compare the number of digits in the registration random code with the number of second word segments. If the number of digits in the registration random code is greater than the number of second word segments, the number of second word segments can be increased to match the number of digits in the registration random code, thus obtaining a new second word segment. Specifically, increasing the number of second word segments to match the number of digits in the registration random code to obtain a new second word segment can be achieved by: obtaining the difference between the number of digits in the registration random code and the number of second word segments; further, after obtaining the difference, the electronic device can extract target word segments from the second word segment in the order of the text from beginning to end, with the number of segments matching the difference; and finally, after obtaining the target word segments, the electronic device can add the target word segments to the end of the second word segment to obtain a new second word segment.
[0092] Step S2540: Fill the third encoding bit of the registration random code based on the first word segment and fill the fourth encoding bit of the registration random code based on the new second word segment to obtain the target registered audio, wherein the sum of the number of the third encoding bit and the number of the fourth encoding bit is equal to the number of bits of the registration random code.
[0093] In some implementations, after obtaining a new second word segment, the electronic device can fill in the third encoding bit of the registration random code based on the first word segment, and fill in the fourth encoding bit of the registration random code based on the new second word segment to obtain the target registered audio. The number of new second word segments is the same as the number of bits in the registration random code and also the same as the number of first word segments. The third encoding bit can be understood as the encoding bit of the corresponding virtual audio of the registration random code, i.e., the virtual encoding bit; the fourth encoding bit can be understood as the encoding bit of the corresponding real audio of the registration random code, i.e., the real encoding bit; the sum of the number of the third and fourth encoding bits is equal to the number of bits in the registration random code. The first encoding bit can be the same as the third encoding bit or the fourth encoding bit; the second encoding bit can be the same as the fourth encoding bit or the third encoding bit.
[0094] In some implementations, the electronic device can sequentially use the odd and even bits of the registered random code, with odd bits as the third encoding bit corresponding to the first word segment and even bits as the fourth encoding bit corresponding to the new second word segment, to achieve interleaving and splicing of the two segments to generate mixed audio, i.e., the target registered audio. The addition of real audio components to the target registered audio increases the diversity of sound.
[0095] Step S260: Associate and store the target registered audio with the registered identity information.
[0096] Furthermore, after obtaining the target registration audio, the electronic device can register a voiceprint corresponding to the registration identity information based on the target registration audio, and associate and store the target registration audio with the registration identity information, registration random code, real human audio, and virtual audio.
[0097] Understandably, registering a hybrid voiceprint using dual audio synthesis can effectively protect the user's physiological voice privacy and prevent the leakage of personal information. Furthermore, by segmenting both real and virtual voice data and replacing or superimposing the virtual audio with a location-coded random code corresponding to the registered identity information, a completely new registration and verification voiceprint is formed. This random code is then used to interweave real and virtual audio, constructing a hybrid audio for voiceprint registration and verification. This increases audio diversity, enhances the security of voiceprint verification, and reduces the false recognition rate of voiceprint recognition based on the target registration audio.
[0098] Step S270: Obtain the target verification audio.
[0099] In some implementations, after the electronic device associates and stores the target registration audio with the registration identity information, it can obtain the target verification audio. Optionally, the electronic device can obtain the verification identity information, as well as the verification real audio and verification virtual audio corresponding to the verification identity information. For example, the metaverse system in the electronic device can obtain the target account entered by the user based on the metaverse system login interface, i.e., the verification identity information. Further, the electronic device can obtain the verification random code corresponding to the target account, as well as the verification real audio and verification virtual audio entered by the user based on the verification identity information.
[0100] Furthermore, the electronic device can generate target verification audio based on real verification audio, virtual verification audio, and a verification random code. Specifically, the electronic device can segment the virtual verification audio using a preset word segmentation model to obtain a first verification word segment, and segment the real verification audio using the same preset word segmentation model to obtain a second verification word segment. Further, if the number of digits in the verification random code is greater than the number of digits in the first verification word segment, the electronic device can supplement the number of first verification word segments to match the number of digits in the verification random code, obtaining a new first verification word segment. If the number of digits in the verification random code is greater than the number of digits in the first verification word segment, the electronic device can obtain the difference between the number of digits in the verification random code and the number of digits in the first verification word segment, and, following the text order of the first verification word segment, obtain a target verification word segment equal to the difference in number from the first verification word segment, adding this target verification word segment to the end of the first verification word segment to obtain a new first verification word segment. After obtaining a new first verification segment, the electronic device can fill the first verification code bit of the verification random code based on the new first verification segment and fill the second verification code bit of the verification random code based on the second verification segment to obtain the target verification audio. The sum of the number of the first verification code bit and the number of the second verification code bit is equal to the number of bits of the verification random code.
[0101] Step S280: Based on a preset similarity algorithm, obtain the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio.
[0102] In some implementations, after obtaining the target verification audio, the electronic device can obtain the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio based on a preset similarity algorithm. This preset similarity algorithm can be pre-set in the electronic device, obtained from a related cloud or electronic device via wireless communication technology, or obtained from a related electronic device via a serial communication interface.
[0103] In some implementations, please refer to Figure 7 Step S280 may include steps S281-S285.
[0104] Step S281: Obtain the product of the voiceprint features of the target registration audio and the voiceprint features of the target verification audio as a first product, and obtain the product of the norm of the voiceprint features of the target registration audio and the norm of the voiceprint features of the target verification audio as a second product.
[0105] In some implementations, after obtaining the target verification audio, the electronic device may obtain the product of the voiceprint features of the target registration audio and the voiceprint features of the target verification audio as a first product, and obtain the product of the norm of the voiceprint features of the target registration audio and the norm of the voiceprint features of the target verification audio as a second product.
[0106] Step S282: Obtain the first parameter by dividing the first product by the second product.
[0107] Furthermore, after obtaining the first product and the second product, the electronic device can obtain the first parameter by dividing the first product by the second product.
[0108] Step S283: Obtain the product of the voiceprint features of the real audio corresponding to the target registration audio and the voiceprint features of the target verification audio as the third product, and obtain the product of the norm of the voiceprint features of the real audio corresponding to the target registration audio and the norm of the voiceprint features of the target verification audio as the fourth product.
[0109] In some implementations, after obtaining the target verification audio, the electronic device may obtain the product of the voiceprint features of the real audio corresponding to the target registration audio and the voiceprint features of the target verification audio as a third product, and obtain the product of the norm of the voiceprint features of the real audio corresponding to the target registration audio and the norm of the voiceprint features of the target verification audio as a fourth product.
[0110] Step S284: Obtain the second parameter by dividing the third product by the fourth product.
[0111] Furthermore, after obtaining the third and fourth products, the electronic device can obtain the second parameter by dividing the third product by the fourth product.
[0112] Step S285: Calculate the product of the first parameter and the common logarithm with the second parameter as the argument, and use it as the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio.
[0113] Furthermore, after obtaining the first parameter and the second parameter, the electronic device can calculate the product of the first parameter and the common logarithm with the second parameter as the argument, and use this product as the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio.
[0114] For example, let A be the voiceprint feature of the target registration audio, B be the voiceprint feature of the target verification audio, ||A|| be the norm of the voiceprint feature of the target registration audio, ||B|| be the norm of the voiceprint feature of the target verification audio, A1 be the voiceprint feature of the real audio corresponding to the target registration audio, and ||A1|| be the norm of the voiceprint feature of the real audio corresponding to the target registration audio. The similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio is...
[0115] Step S290: If the similarity is greater than or equal to the similarity threshold, then the target verification audio is determined to have passed verification, and the target verification audio is stored as a new target registration audio associated with the registration identity information.
[0116] In some implementations, a similarity threshold may be preset in the electronic device. After the electronic device obtains the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio, it can compare the similarity with the similarity threshold. If the similarity is greater than or equal to the similarity threshold, it is determined that the target verification audio has passed the verification, and the target verification audio is stored as a new target registration audio associated with the registration identity information.
[0117] Once the electronic device determines that the target verification audio verification is successful, it can replace the registration random code with the verification random code as the new registration random code, or it can replace the target registration audio with the target verification audio as the new target registration audio. The new target registration audio is then associated and stored with the registration identity information, the new registration random code, and the real verification audio corresponding to the new target registration audio, so that the electronic device can perform voiceprint verification based on the new target registration audio.
[0118] It is understandable that artificially generated voiceprints possess both real and virtual voiceprint information, making the mixed audio similar to the two independent voiceprints, which can easily lead to recognition confusion. To circumvent this situation and ensure the uniqueness of the mixed voiceprint, this application uses a low-confusion scoring formula, i.e., a preset similarity algorithm, to obtain the similarity between the target verification voiceprint and the target registration voiceprint. This reduces the probability of electronic devices incorrectly accepting real and virtual voiceprints, effectively preventing malicious actors from bypassing voiceprint verification through brute-force enumeration. Furthermore, since the mixed audio target verification audio contains segments of real and virtual verification audio, directly calculating the similarity between the real and virtual verification audio and the mixed audio target registration audio will yield a high similarity value. Therefore, based on the similarity threshold, it is impossible to effectively verify user information and resist malicious attacks. This application's embodiment uses dual audio for voiceprint registration and verification, effectively avoiding the problem of malicious attacks bypassing voiceprint verification. The low-obfuscation scoring formula uses real audio as a comparison object for similarity calculation, making the comparison between virtual audio and real human audio significantly different. Furthermore, it performs logarithmic processing on the voiceprint to reduce the overall similarity value between the target registration audio and the target verification audio. This effectively separates the mixed audio used to maliciously attack voiceprint verification from the target verification audio, greatly avoiding malicious attacks by criminals and improving the security of voiceprint verification.
[0119] An embodiment of this application provides an audio generation method that, compared to... Figure 1The audio generation method shown in this embodiment can further obtain target verification audio after associating and storing the target registration audio with the registration identity information. Based on a preset similarity algorithm, the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio is obtained. If the similarity is greater than or equal to the similarity threshold, the target verification audio is determined to have passed verification. The target verification audio is then used as a new target registration audio and associated with the registration identity information for storage. This generates mixed audio through dual audio, increasing the diversity of sounds and improving the recognition rate of mixed voiceprints. In addition, voiceprint verification is performed on the obtained verification audio based on the target registration audio, improving the security of voiceprint verification. At the same time, the verified target verification audio is used as a new target registration audio and associated with the registration identity information for storage. The stored registration audio is updated in real time, improving the accuracy of voiceprint verification and enhancing the user experience.
[0120] Please see Figure 8 , Figure 8 A block diagram of an audio generation apparatus according to an embodiment of this application is shown. This audio generation apparatus 200 is applied to the aforementioned electronic device. The following will focus on... Figure 8 The process is described in detail below. The audio generation device 200 includes: a registration identity information acquisition module 210, a virtual audio acquisition module 220, a real audio acquisition module 230, a registration random code acquisition module 240, a target registration audio acquisition module 250, and a storage module 260, wherein:
[0121] The registration identity information acquisition module 210 is used to acquire registration identity information and the virtual voice parameters corresponding to the registration identity information.
[0122] The virtual audio acquisition module 220 is used to obtain the virtual audio corresponding to the registered identity information based on the virtual sound parameters.
[0123] The real audio acquisition module 230 is used to acquire the real audio corresponding to the registered identity information.
[0124] The registration random code acquisition module 240 is used to obtain a registration random code corresponding to the registration identity information based on the virtual audio and the real audio.
[0125] The target registration audio acquisition module 250 is used to obtain the target registration audio based on the registration random code, the virtual audio, and the real audio.
[0126] Storage module 260 is used to associate and store the target registered audio with the registered identity information.
[0127] Furthermore, after associating and storing the target registration audio with the registration identity information, the audio generation device 200 may further include: a target verification audio acquisition module, a similarity acquisition module, and a storage update module, wherein:
[0128] The target verification audio acquisition module is used to acquire the target verification audio.
[0129] The similarity acquisition module is used to obtain the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio based on a preset similarity algorithm.
[0130] The storage update module is used to determine that the target verification audio has passed verification if the similarity is greater than or equal to the similarity threshold, and to store the target verification audio as a new target registration audio associated with the registration identity information.
[0131] Further, the similarity acquisition module may include: a first product and a second product acquisition unit, a first parameter acquisition unit, a third product and a fourth product acquisition unit, a second parameter acquisition unit, and a similarity acquisition unit, wherein:
[0132] The first product and second product acquisition unit are used to acquire the product of the voiceprint features of the target registration audio and the voiceprint features of the target verification audio as the first product, and to acquire the product of the norm of the voiceprint features of the target registration audio and the norm of the voiceprint features of the target verification audio as the second product.
[0133] The first parameter obtaining unit is used to obtain the first parameter by dividing the first product by the second product.
[0134] The third and fourth product acquisition units are used to obtain the product of the voiceprint features of the real audio corresponding to the target registration audio and the voiceprint features of the target verification audio as the third product, and to obtain the product of the norm of the voiceprint features of the real audio corresponding to the target registration audio and the norm of the voiceprint features of the target verification audio as the fourth product.
[0135] The second parameter obtaining unit is used to obtain the second parameter by dividing the third product by the fourth product.
[0136] The similarity acquisition unit is used to calculate the product of the first parameter and the common logarithm with the second parameter as the argument, as the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio.
[0137] Further, the registration random code acquisition module 240 may include: a first word segment acquisition unit, a second word segment acquisition unit, and a registration random code acquisition unit, wherein:
[0138] The first segmentation fragment acquisition unit is used to perform segmentation processing on the virtual audio based on a preset segmentation model to obtain the first segmentation fragment.
[0139] The second segmentation unit is used to perform segmentation processing on the real audio based on the preset segmentation model to obtain the second segmentation.
[0140] The registration random code acquisition unit is configured to generate a random code with the same number of bits as the number of the first word segments as the registration random code corresponding to the registration identity information if the number of the first word segments is greater than or equal to the number of the second word segments; or
[0141] If the number of the second word segment is greater than or equal to the number of the first word segment, a random code with the same number of bits as the number of the second word segment is generated as a registration random code corresponding to the registration identity information.
[0142] Furthermore, the target registration audio acquisition module 250 may include: a new first segmentation fragment acquisition unit and a target registration audio acquisition first unit, wherein:
[0143] The new first word segment obtaining unit is used to supplement the number of the first word segments to the same number of bits as the registered random code if the number of bits of the registered random code is greater than the number of the first word segments, thereby obtaining a new first word segment.
[0144] The target registered audio obtains a first unit, which is used to fill the first encoding bit of the registration random code based on the new first word segment and to fill the second encoding bit of the registration random code based on the second word segment to obtain the target registered audio, wherein the sum of the number of the first encoding bit and the number of the second encoding bit is equal to the number of bits of the registration random code.
[0145] Furthermore, the new first segmentation fragment acquisition unit may include: a difference acquisition unit, a target segmentation fragment acquisition unit, and a new first segmentation fragment acquisition subunit, wherein:
[0146] The difference acquisition unit is used to acquire the difference between the number of bits of the registration random code and the number of the first word segment if the number of bits of the registration random code is greater than the number of the first word segment.
[0147] The target segmentation fragment acquisition unit is used to acquire target segmentation fragments from the first segmentation fragment in the order of the text of the first segmentation fragment from beginning to end, with the number of such fragments being the same as the difference value.
[0148] A new first segmentation fragment acquisition subunit is used to add the target segmentation fragment to the end of the first segmentation fragment to obtain the new first segmentation fragment.
[0149] Furthermore, the target registration audio acquisition module 250 may include: a new second word segment acquisition unit and a target registration audio acquisition second unit, wherein:
[0150] The new second word segment obtaining unit is used to supplement the number of second word segments to the same number of bits as the registered random code if the number of bits of the registered random code is greater than the number of second word segments, thereby obtaining a new second word segment.
[0151] The target registered audio obtains a second unit, which is used to fill the third encoding bit of the registration random code based on the first word segment and to fill the fourth encoding bit of the registration random code based on the new second word segment to obtain the target registered audio, wherein the sum of the number of the third encoding bit and the number of the fourth encoding bit is equal to the number of bits of the registration random code.
[0152] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0153] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0154] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0155] Please see Figure 9 This document illustrates a structural block diagram of an electronic device according to an embodiment of this application. The electronic device 100 can be a smartphone, tablet computer, e-reader, or other electronic device capable of running applications. The electronic device 100 in this application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications can be stored in the memory 120 and configured to be executed by one or more processors 110, and the one or more applications are configured to perform the methods described in the foregoing method embodiments.
[0156] The processor 110 may include one or more processing cores. The processor 110 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120. Optionally, the processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 110 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 110 and may be implemented separately using a communication chip.
[0157] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the electronic device 100 during use (such as phonebook data, audio and video data, chat log data, etc.).
[0158] Please see Figure 10 This illustration shows a structural block diagram of a computer-readable storage medium according to an embodiment of this application. The computer-readable medium 300 stores program code that can be invoked by a processor to execute the methods described in the above method embodiments.
[0159] The computer-readable storage medium 300 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 300 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 300 has storage space for program code 310 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 310 may be compressed, for example, in a suitable form.
[0160] In summary, the audio generation method, apparatus, electronic device, and storage medium provided in this application obtain registration identity information and virtual sound parameters corresponding to the registration identity information. Based on the virtual sound parameters, they obtain virtual audio corresponding to the registration identity information and real audio corresponding to the registration identity information. Based on the virtual audio and real audio, they obtain a registration random code corresponding to the registration identity information. Based on the registration random code, virtual audio, and real audio, they obtain a target registration audio. The target registration audio is then associated with and stored with the registration identity information. This process generates mixed audio through dual audio, increasing the diversity of sounds and improving the recognition rate of mixed voiceprints.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An audio generation method, characterized by, The method includes: Obtain the registered identity information and the virtual voice parameters corresponding to the registered identity information; Based on the virtual sound parameters, obtain the virtual audio corresponding to the registered identity information; Obtain the actual audio corresponding to the registered identity information; Based on the virtual audio and the real audio, a registration random code corresponding to the registered identity information is obtained; Based on the registration random code, the virtual audio, and the real audio, the target registration audio is obtained; The target registered audio is associated with and stored in relation to the registered identity information; The step of obtaining a registration random code corresponding to the registration identity information based on the virtual audio and the real audio includes: The virtual audio is segmented to obtain the first segmented word fragment; The real audio is segmented into words to obtain a second segmented word fragment; Based on the number of the first word segment and the number of the second word segment, a registration random code corresponding to the registration identity information is obtained; The step of obtaining the target registration audio based on the registration random code, the virtual audio, and the real audio includes: The virtual audio encoding bits of the registration random code are filled with virtual audio, and the real audio encoding bits of the registration random code are filled with real audio to obtain the target registration audio. The sum of the number of virtual audio encoding bits and the number of real audio encoding bits is the same as the number of bits of the registration random code.
2. The method of claim 1, wherein, After associating and storing the target registered audio with the registered identity information, the method further includes: Obtain the target verification audio; Based on a preset similarity algorithm, the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio is obtained; If the similarity is greater than or equal to the similarity threshold, the target verification audio is determined to have passed verification, and the target verification audio is stored as a new target registration audio associated with the registration identity information.
3. The method of claim 2, wherein, The step of obtaining the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio based on a preset similarity algorithm includes: The product of the voiceprint features of the target registration audio and the voiceprint features of the target verification audio is obtained as a first product, and the product of the norm of the voiceprint features of the target registration audio and the norm of the voiceprint features of the target verification audio is obtained as a second product. The first parameter is obtained by dividing the first product by the second product; The product of the voiceprint features of the real audio corresponding to the target registration audio and the voiceprint features of the target verification audio is obtained as the third product, and the product of the norm of the voiceprint features of the real audio corresponding to the target registration audio and the norm of the voiceprint features of the target verification audio is obtained as the fourth product. The second parameter is obtained by dividing the third product by the fourth product; The product of the first parameter and the common logarithm with the second parameter as the argument is calculated as the similarity between the voiceprint of the target verification audio and the voiceprint of the target registration audio.
4. The method of claim 1, wherein, The step of obtaining a registration random code corresponding to the registration identity information based on the virtual audio and the real audio includes: The virtual audio is segmented based on a preset word segmentation model to obtain a first segmented word fragment. The real audio is segmented based on the preset word segmentation model to obtain a second segmented word fragment. If the number of the first word segment is greater than or equal to the number of the second word segment, then a random code with the same number of bits as the number of the first word segment is generated as the registration random code corresponding to the registration identity information; or If the number of the second word segment is greater than or equal to the number of the first word segment, then a random code with the same number of bits as the number of the second word segment is generated as the registration random code corresponding to the registration identity information.
5. The method of claim 4, wherein, The process of obtaining the target registration audio based on the registration random code, the virtual audio, and the real audio includes: If the number of bits in the registration random code is greater than the number of bits in the first word segment, then the number of the first word segment is increased to be the same as the number of bits in the registration random code to obtain a new first word segment. The target registered audio is obtained by filling the first encoding bit of the registration random code with the new first word segment and filling the second encoding bit of the registration random code with the second word segment, wherein the sum of the number of the first encoding bit and the number of the second encoding bit is equal to the number of bits of the registration random code.
6. The method of claim 5, wherein, If the number of bits in the registered random code is greater than the number of the first word segment, then the number of the first word segment is increased to be the same as the number of bits in the registered random code to obtain a new first word segment, including: If the number of bits in the registration random code is greater than the number of the first word segment, then the difference between the number of bits in the registration random code and the number of the first word segment is obtained; According to the text order of the first segmented segment from beginning to end, obtain the target segmented segment with the same number of differences from the first segmented segment; The target word segment is added to the end of the first word segment to obtain the new first word segment.
7. The method according to claim 4, characterized in that, The process of obtaining the target registration audio based on the registration random code, the virtual audio, and the real audio includes: If the number of bits in the registration random code is greater than the number of the second word segment, then the number of the second word segment is increased to be the same as the number of bits in the registration random code to obtain a new second word segment. The target registered audio is obtained by filling the third encoding bit of the registration random code based on the first word segment and filling the fourth encoding bit of the registration random code based on the new second word segment, wherein the sum of the number of the third encoding bit and the number of the fourth encoding bit is equal to the number of bits of the registration random code.
8. An audio generation apparatus, characterized in that, The device includes: The registration identity information acquisition module is used to acquire registration identity information and the virtual voice parameters corresponding to the registration identity information; A virtual audio acquisition module is used to obtain the virtual audio corresponding to the registered identity information based on the virtual sound parameters; The real audio acquisition module is used to acquire the real audio corresponding to the registered identity information; The registration random code acquisition module is used to obtain a registration random code corresponding to the registration identity information based on the virtual audio and the real audio. The target registration audio acquisition module is used to obtain the target registration audio based on the registration random code, the virtual audio, and the real audio. A storage module is used to associate and store the target registration audio with the registration identity information; Specifically, the registration random code acquisition module is used to perform word segmentation on the virtual audio to obtain a first word segment; perform word segmentation on the real audio to obtain a second word segment; and obtain a registration random code corresponding to the registration identity information based on the number of the first word segment and the number of the second word segment. The target registration audio acquisition module is specifically used to fill the virtual audio encoding bits of the registration random code based on virtual audio, and fill the real audio encoding bits of the registration random code based on real audio to obtain the target registration audio. The sum of the number of virtual audio encoding bits and the number of real audio encoding bits is the same as the number of bits of the registration random code.
9. An electronic device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Speech synthesis system for online game and implementation method thereof
CN102479506A
Arrangements for Using Voice Biometrics in Internet Based Activities
US20090187405A1