A method, apparatus, and electronic device for determining the accuracy of voiceprint features.

By judging the accuracy of short speech voiceprint feature groups by two-dimensional voiceprint distribution density, the problem of inaccurate voiceprint features caused by non-stationary noise in the existing technology is solved, and the accuracy of voiceprint recognition is improved.

CN120766683BActive Publication Date: 2026-06-30HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-06-18
Publication Date
2026-06-30

Smart Images

  • Figure CN120766683B_ABST
    Figure CN120766683B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, and electronic device for judging the accuracy of voiceprint features, relating to the field of voiceprint recognition technology. The method, applied to an electronic device, includes: acquiring a first long speech segment corresponding to a target object; segmenting the first long speech segment into multiple short speech segments; extracting voiceprint features from each short speech segment to obtain a first voiceprint feature group; and performing an accuracy judgment operation based on the first voiceprint feature group. This accuracy judgment operation includes: calculating a first distribution density of the first voiceprint feature group in a two-dimensional circle corresponding to the first voiceprint feature group; and judging whether the first voiceprint feature group is accurate based on the first distribution density. This application judges the accuracy of short speech voiceprint feature groups by using the distribution density of the short speech voiceprint feature groups in their corresponding two-dimensional circles, which can achieve the accuracy judgment of long speech voiceprint features, thereby ensuring the accuracy of voiceprint recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voiceprint recognition technology, and in particular to a method, apparatus and electronic device for judging the accuracy of voiceprint features. Background Technology

[0002] Voiceprint recognition technology, as a biometric technology, has been widely used in various fields such as identity authentication, voice assistants, and security monitoring. Voiceprint recognition primarily identifies a target by analyzing the voiceprint characteristics of long-form speech produced by that target.

[0003] Current techniques typically obtain the voiceprint features of long speech samples from a target object by using a sliding window to segment the long speech into several short speech samples, extracting the voiceprint features of each short speech sample separately, and finally summing and averaging the results to obtain the voiceprint features of the long speech. However, because long speech samples usually contain non-stationary noise, the extracted voiceprint features of the short speech samples may be inaccurate, which in turn leads to inaccurate voiceprint features of the obtained long speech (i.e., the voiceprint features obtained through this conventional method may deviate from the true voiceprint features), affecting the accuracy of voiceprint recognition. Summary of the Invention

[0004] This application provides a method, apparatus, and electronic device for judging the accuracy of voiceprint features. It mainly judges whether the short speech voiceprint feature group is accurate by the two-dimensional voiceprint distribution density of the short speech voiceprint feature group, and then judges whether the voiceprint features of the long speech are accurate, so as to ensure the accuracy of voiceprint recognition.

[0005] In a first aspect, embodiments of this application provide a method for accurately determining the accuracy of voiceprint features, which is applied to an electronic device. This method can be executed by the electronic device, or by a chip, module, or unit configured within the electronic device.

[0006] Specifically, the method includes: acquiring a first long speech corresponding to the target object; segmenting the first long speech into multiple short speech segments; extracting the voiceprint features of each short speech segment to obtain a first voiceprint feature group; performing an accuracy judgment operation based on the first voiceprint feature group, the accuracy judgment operation including: calculating a first distribution density of the first voiceprint feature group in the two-dimensional circle corresponding to the first voiceprint feature group; and judging whether the first voiceprint feature group is accurate based on the first distribution density.

[0007] Understandably, the denser the distribution of voiceprint features in a short speech voiceprint feature group (e.g., the first voiceprint feature group mentioned above), the more accurate the voiceprint features of long speech obtained based on that voiceprint feature group will be. However, the voiceprint features in a short speech voiceprint feature group are usually distributed in a high-dimensional space, which is not conducive to calculating the distribution density. Therefore, it is necessary to use a dimensionality reduction method to transform the high-dimensional space into a two-dimensional plane, and then calculate the distribution density based on the two-dimensional plane.

[0008] This application determines the accuracy of short speech voiceprint feature groups by measuring their distribution density within a corresponding two-dimensional circle. This allows for the accurate assessment of long speech voiceprint features, thereby ensuring high accuracy in voiceprint recognition. For example, if a short speech voiceprint feature group is deemed accurate, the resulting long speech voiceprint features can be considered accurate as well, and voiceprint recognition can then be performed based on these long speech features, thus guaranteeing high accuracy. Alternatively, if a short speech voiceprint feature group is deemed inaccurate, the resulting long speech voiceprint features can be considered inaccurate as well, and voiceprint recognition can no longer be performed based on these long speech features, thus avoiding impacting the accuracy of voiceprint recognition.

[0009] In conjunction with the first aspect, in some implementations, calculating the first distribution density of the first voiceprint feature group in the two-dimensional circle corresponding to the first voiceprint feature group includes: calculating the centroid of the first voiceprint feature group; calculating the cosine distance of each voiceprint feature in the first voiceprint feature group from the centroid to obtain a first cosine distance group; determining the two-dimensional circle based on the voiceprint feature corresponding to the largest cosine distance in the first cosine distance group and the centroid; mapping each voiceprint feature in the first voiceprint feature group to the two-dimensional circle; calculating the area of ​​the two-dimensional circle; and determining the first distribution density based on the ratio of the area of ​​the two-dimensional circle to the number of voiceprint features contained in the first voiceprint feature group.

[0010] This application utilizes the angle between high-dimensional vectors (i.e., cosine distance) to map the high-dimensional voiceprint space containing short speech voiceprint feature groups into a three-dimensional voiceprint space, and then further maps the three-dimensional voiceprint space into a two-dimensional circle, so that the area that cannot be calculated in the high-dimensional space is mapped into the area of ​​a circle that can be calculated in the two-dimensional space.

[0011] Furthermore, using the area of ​​a two-dimensional circle divided by the number of voiceprint features it contains as the two-dimensional voiceprint distribution density allows for greater focus on removing voiceprint features at the edge of the two-dimensional circle during subsequent feature correction.

[0012] In conjunction with the first aspect, in some implementations, determining the two-dimensional circle based on the voiceprint feature corresponding to the largest cosine distance in the first cosine distance group and the centroid includes: mapping the voiceprint feature corresponding to the largest cosine distance in the first cosine distance group to a radius, with the centroid as the center; and obtaining the two-dimensional circle based on the center and the radius. Based on this, a high-dimensional voiceprint space can be mapped onto a two-dimensional circle.

[0013] In conjunction with the first aspect, in some implementations, the area of ​​the two-dimensional circle satisfies the following relationship:

[0014] S=2πd max (2-d max )

[0015] Where S is the area of ​​the two-dimensional circle, d max It is the largest cosine distance in the first cosine distance group.

[0016] In conjunction with the first aspect, in some implementations, determining whether the first voiceprint feature group is accurate based on the first distribution density includes: considering the first voiceprint feature group to be accurate when the first distribution density is greater than a first threshold; or, performing a first operation when the first distribution density is less than or equal to the first threshold, the first operation including: removing p% of the voiceprint features from the first voiceprint feature group to obtain a second voiceprint feature group, wherein the initial value of p is ρ, and the cosine distance of each voiceprint feature removed from the centroid is greater than the cosine distance of any voiceprint feature in the second voiceprint feature group from the centroid; calculating the second distribution density of the second voiceprint feature group in the two-dimensional circle corresponding to the second voiceprint feature group when the number of voiceprint features in the second voiceprint feature group is greater than or equal to a second threshold; updating the first voiceprint feature group based on the second voiceprint feature group when the second distribution density is greater than or equal to the first distribution density, updating the current p value to the initial value, and then re-performing the accuracy determination operation based on the updated first voiceprint feature group.

[0017] It should be understood that when the first distribution density is greater than the first threshold, the voiceprint features in the first voiceprint feature group are considered to be relatively densely distributed, and there is no need to further modify the first voiceprint feature group. When the first distribution density is less than or equal to the first threshold, the voiceprint features at the edge of the two-dimensional circle can be removed to obtain the second voiceprint feature group. When the second voiceprint feature group meets certain conditions (i.e., when the number of voiceprint features in the second voiceprint feature group is greater than or equal to the second threshold, and the second distribution density is greater than or equal to the first distribution density), the first voiceprint feature group is updated based on the second voiceprint feature group, thereby improving the voiceprint distribution density and making the voiceprint features of the final long speech closer to the real voiceprint features, thus improving the accuracy of voiceprint recognition.

[0018] It should be understood that this application does not limit the threshold (e.g., a first threshold, a second threshold, or a third threshold) or the initial value of p. For example, the threshold and the initial value of p can be determined by the researchers or users based on experience.

[0019] In conjunction with the first aspect, in some implementations, the first operation further includes: when the number of voiceprint features in the second voiceprint feature group is less than the second threshold, determining whether the current p value is equal to the initial value; when the current p value is equal to the initial value, decreasing the p value, and re-executing the first operation based on the decreased p value; when the current p value is not equal to the initial value, performing a second operation, the second operation including: determining whether the first distribution density is less than a third threshold, the third threshold being less than or equal to the first threshold; when the first distribution density is less than the third threshold, considering the first voiceprint feature group to be inaccurate; when the first distribution density is greater than or equal to the third threshold, considering the first voiceprint feature group to be accurate.

[0020] It should be understood that, for cases where the number of voiceprint features in the second voiceprint feature group is less than the second threshold, considering that this situation may be due to the removal of too many voiceprint features previously, it is possible to first determine whether the current p value is the initial value. If so, the p value can be reduced, and the first operation can be re-executed based on the reduced p value. If not, it means that the current p value has already been adjusted, and the final judgment process (i.e., the second operation) can be directly entered.

[0021] In conjunction with the first aspect, in some implementations, the first operation further includes: when the second distribution density is less than the first distribution density, determining whether the current p value is equal to the initial value; when the current p value is equal to the initial value, increasing the p value and re-executing the first operation based on the increased p value; and when the current p value is not equal to the initial value, performing the second operation.

[0022] It should be understood that, in the case where the second distribution density is less than the first distribution density, we can continue to determine whether the current p value is the initial value. If it is, considering that this situation may be due to the current p value being low, we can increase the p value and re-execute the first operation based on the increased p value. If not, it means that the current p value has already been adjusted and has not yielded a better result, so we can directly enter the final judgment process (i.e., execute the second operation).

[0023] In conjunction with the first aspect, in some implementations, the reduced p value is ρ / 2, and the increased p value is 3ρ / 2.

[0024] In conjunction with the first aspect, in some implementations, the method further includes: when the first voiceprint feature set is considered accurate, using the mean of the first voiceprint feature set as the voiceprint feature of the first long speech. Based on this, the accuracy of the voiceprint features of the first long speech can be guaranteed, thereby ensuring the accuracy of voiceprint recognition.

[0025] In conjunction with the first aspect, in some implementations, when the first voiceprint feature set is deemed inaccurate, a second long speech corresponding to the target object is obtained, and the accuracy of the short speech voiceprint feature set corresponding to the second long speech is determined. Based on this, the accuracy of the voiceprint features of the long speech corresponding to the target object can be guaranteed, thereby ensuring the accuracy of voiceprint recognition.

[0026] Secondly, embodiments of this application provide a device for determining the accuracy of voiceprint features. This device may be an electronic device, or it may be a chip, module, or unit configured in an electronic device. Specifically, the device includes at least one processor coupled to a memory, which can be used to execute instructions in the memory to implement the methods in the first aspect and any possible implementation thereof.

[0027] Optionally, the device further includes a memory. Optionally, the device further includes a communication interface, to which the processor is coupled.

[0028] Thirdly, embodiments of this application provide an electronic device including at least one processor, the at least one processor being configured to perform the methods described in the first aspect and any possible implementation thereof.

[0029] Fourthly, a computer program product is provided, the computer program product comprising: a computer program (also referred to as code or instructions), which, when the computer program is run, causes a computer to perform the methods described in the first aspect and any possible implementation thereof.

[0030] Fifthly, a computer-readable medium is provided that stores a computer program (also referred to as code or instructions) that, when run on a computer, causes the computer to perform the methods described in the first aspect and any possible implementation thereof. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of a voiceprint recognition system provided in an embodiment of this application;

[0032] Figure 2 This is a schematic diagram illustrating a method for extracting voiceprint features from long speech.

[0033] Figure 3This is a schematic diagram illustrating the distribution of short speech voiceprint features provided in an embodiment of this application;

[0034] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0035] Figure 5 This is a software structure block diagram of an electronic device provided in an embodiment of this application;

[0036] Figure 6 This is a schematic flowchart illustrating a method for determining the accuracy of voiceprint features provided in an embodiment of this application;

[0037] Figure 7 This is a schematic flowchart illustrating the calculation of the distribution density of short speech voiceprint features, provided in an embodiment of this application.

[0038] Figure 8 This is a schematic diagram illustrating a method for determining a two-dimensional circle according to an embodiment of this application;

[0039] Figure 9 This is a schematic flowchart illustrating how to determine whether the current first voiceprint feature group is accurate, provided in an embodiment of this application.

[0040] Figure 10 This is a schematic diagram illustrating the correction of the current first voiceprint feature group provided in an embodiment of this application;

[0041] Figure 11 This is a schematic structural diagram of a device for judging the accuracy of voiceprint features provided in an embodiment of this application;

[0042] Figure 12 This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation

[0043] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" do not necessarily imply that they are different.

[0044] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0045] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0046] The following is a detailed description of the proposed solution with reference to the accompanying drawings.

[0047] Figure 1 This is a schematic diagram of a voiceprint recognition system provided in an embodiment of this application. Figure 1 As shown, the voiceprint recognition system 10 mainly includes a registration system 11 and an authentication system 12.

[0048] The registration system 11 includes: an audio acquisition module, a preprocessing module, a registration card control module, a feature extraction module, and a voiceprint feature library.

[0049] The audio acquisition module is used to acquire the voice signal (i.e. long speech) emitted by the object to be registered.

[0050] The preprocessing module performs endpoint detection and noise cancellation on the acquired long audio streams. Endpoint detection involves analyzing the input audio stream and automatically removing invalid parts such as silence or non-human voices, retaining only valid speech. The noise cancellation stage filters out background noise to meet user needs in different environments.

[0051] The registration and control module is used to intercept and control voice messages that are still unqualified after preprocessing (e.g., excessive noise).

[0052] The feature extraction module is used to acquire uninterrupted speech signals from the registration and control module, and extract spectral feature parameters from the speech signals that can characterize the specific organ structure or behavioral habits of the speaker. It should be understood that these feature parameters are relatively stable for the same speaker, do not change with time or environment, are consistent across different utterances of the same speaker, and have strong noise resistance and are not easily imitated.

[0053] The voiceprint feature library is used to store the voiceprint features extracted by the feature extraction module.

[0054] The authentication system 12 includes: an audio acquisition module, a preprocessing module, a feature extraction module, a matching module, and an authentication result output module.

[0055] The audio acquisition module is used to acquire the voice signal (i.e., long voice) emitted by the object to be authenticated.

[0056] The preprocessing module performs endpoint detection and noise cancellation on the acquired long audio streams. Endpoint detection involves analyzing the input audio stream and automatically removing invalid parts such as silence or non-human voices, retaining only valid speech. The noise cancellation stage filters out background noise to meet user needs in different environments.

[0057] The feature extraction module is used to acquire the preprocessed speech signal from the preprocessing module and extract spectral feature parameters from the speech signal that can characterize the specific organ structure or behavioral habits of the speaker. It should be understood that these feature parameters are relatively stable for the same speaker, do not change with time or environment, are consistent across different utterances of the same speaker, and have strong noise resistance and are not easily imitated.

[0058] The matching module is used to match the voiceprint features extracted by the feature extraction module with the voiceprint features stored in the voiceprint feature library.

[0059] The authentication result output module is used to output the authentication result based on the matching result of the matching module, that is, to output whether the current object to be authenticated is an object registered by the system.

[0060] Current technologies generally extract speaker features from long speech segments as follows: A sliding window is used to segment the long speech segment into several shorter segments, and then the speaker features of each short segment are extracted to obtain a set of short speech speaker features. Finally, the speaker features of the long speech segment are obtained by summing and averaging these features. Figure 2 As shown. However, because long speech often contains non-stationary noise, the extracted speaker features from short speech may be inaccurate. For example, the speaker features of short speech may be scattered (e.g., Figure 3 (as shown in (a)) or causes abnormalities in the voiceprint features of certain short speech words (such as...) Figure 3As shown in (b) in the figure), this leads to inaccurate voiceprint features of the obtained long speech (i.e., the voiceprint features obtained through this conventional method may deviate from the real voiceprint features, such as...). Figure 3 (As shown), this affects the accuracy of subsequent voiceprint recognition.

[0061] Based on this, this application proposes a method for judging the accuracy of voiceprint features. It mainly judges whether the short speech voiceprint feature group is accurate by the two-dimensional voiceprint distribution density of the short speech voiceprint feature group, and then judges whether the voiceprint features of the long speech are accurate, so as to ensure the accuracy of voiceprint recognition.

[0062] The solution of this application can be applied to electronic devices, which can also be called user equipment (UE), terminal equipment, or user terminal, etc.; the electronic device can be, but is not limited to, mobile phones, tablets, desktops, laptops, handheld computers, notebook computers, in-vehicle devices, ultra-mobile personal computers (UMPC), netbooks, cellular phones, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, etc., and the embodiments of this application do not limit this.

[0063] For example, Figure 4 This is a schematic diagram of the hardware structure of an electronic device 100 provided in an embodiment of this application. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a USB interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0064] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0065] Processor 110 may include one or more processing units, such as application processors, satellite communication processors, modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0066] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0067] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0068] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0069] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0070] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0071] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals.

[0072] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the same device as at least some modules of the processor 110.

[0073] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0074] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with networks and other devices through wireless communication technology.

[0075] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0076] The display screen 194 is used to display images, display videos, and receive swipe operations, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0077] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0078] The ISP is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0079] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0080] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0081] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0082] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0083] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0084] Internal memory 121 can be used to store executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located within the processor.

[0085] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0086] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0087] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, touch operations applied to different applications (such as taking photos, playing audio, etc.) can correspond to different vibration feedback effects. Touch operations applied to different areas of the display screen 194 can also correspond to different vibration feedback effects from motor 191. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0088] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0089] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0090] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.

[0091] Figure 5 This is a software structure block diagram of an electronic device 100 provided in an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0092] The application layer can include a series of application packages.

[0093] like Figure 5 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0094] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0095] like Figure 5 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0096] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0097] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0098] A view system includes visual controls, such as controls for displaying characters and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text message notification icon could include a view for displaying characters and a view for displaying images.

[0099] A phone manager is used to provide communication functions for electronic devices. For example, it manages call status (including connection and disconnection).

[0100] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0101] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0102] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0103] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0104] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0105] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0106] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0107] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0108] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0109] A 2D graphics engine is a graphics engine for 2D drawing.

[0110] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0111] It should be noted that although the embodiments of this application are described using the Android system, the solutions of this application are also applicable to electronic devices with operating systems such as iOS or Windows.

[0112] Figure 6 This is a schematic flowchart illustrating a method for determining the accuracy of voiceprint features provided in an embodiment of this application. Figure 6 As shown, the method 600 may include steps S610 to S640, and each step of the method is described in detail below.

[0113] S610, obtain the first long speech corresponding to the target object.

[0114] Optionally, method 600 can be applied to the feature extraction module in the registration system 11, in which case the target object is the object to be registered; or it can be applied to the feature extraction module in the authentication system 12, in which case the target object is the object to be authenticated.

[0115] This application does not limit the form of acquiring the first long speech. For example, the first long speech may be the speech made by the target object based on system prompts, or it may be the sound made by the target object according to its own will.

[0116] S620 segments the first long speech into multiple short speech segments.

[0117] Specifically, the first long speech segment can be divided into multiple short speech segments using a sliding window, such as... Figure 2 As shown.

[0118] S630, extract the voiceprint features of each short speech segment to obtain the first voiceprint feature group.

[0119] The voiceprint features of each short speech segment can be extracted using a pre-trained voiceprint model to obtain the first voiceprint feature group.

[0120] S640, perform an accuracy judgment operation based on the first voiceprint feature group, wherein the accuracy judgment operation includes: step S1, calculating the first distribution density of the first voiceprint feature group in the two-dimensional circle corresponding to the first voiceprint feature group; step S2, judging whether the first voiceprint feature group is accurate based on the first distribution density.

[0121] Understandably, the denser the distribution of voiceprint features in a short speech voiceprint feature group (e.g., the first voiceprint feature group mentioned above), the more accurate the voiceprint features of long speech obtained based on that voiceprint feature group will be. However, the voiceprint features in a short speech voiceprint feature group are usually distributed in a high-dimensional space, which is not conducive to calculating the distribution density. Therefore, it is necessary to use a dimensionality reduction method to transform the high-dimensional space into a two-dimensional plane, and then calculate the distribution density based on the two-dimensional plane.

[0122] Specifically, step S1 above can be based on... Figure 7 The method shown transforms a high-dimensional space into a two-dimensional plane and calculates the first distribution density based on the two-dimensional plane, such as... Figure 7 As shown, method 700 includes steps S710 to S760.

[0123] S710, calculate the centroid of the first voiceprint feature group.

[0124] The first voiceprint feature group can be characterized as v = (v1, v2, ..., v N N is the number of voiceprint features included in the first voiceprint feature group.

[0125] The centroid can be obtained by summing the voiceprint features in the first voiceprint feature group and then taking the mean, as shown in the following formula:

[0126]

[0127] Where e is the centroid, v i It is the i-th voiceprint feature in the first voiceprint feature group.

[0128] S720, calculate the cosine distance from the centroid to each voiceprint feature in the first voiceprint feature group to obtain the first cosine distance group.

[0129] It's important to note that cosine distance is a vector-space-based metric that measures the similarity of two vectors by calculating the cosine of the angle between them. If two vectors are identical, the angle between them is 0 degrees, and the cosine value is 1, indicating they are completely identical. If two vectors are completely unrelated, the angle between them is 90 degrees, and the cosine value is 0, indicating they are completely different. The cosine distance between a voiceprint feature and its centroid is a commonly used similarity metric in voiceprint recognition. It is based on the concept of cosine similarity and is used to compare the degree of similarity between two voiceprint features.

[0130] In voiceprint recognition, the cosine distance from each voiceprint feature to the centroid is calculated. This distance reflects the distribution of voiceprint features and the similarity between each voiceprint feature and the whole.

[0131] The first cosine distance set can be represented as d = (d1, d2, ..., d...). N ).

[0132] S730, a two-dimensional circle is determined based on the voiceprint features and centroid corresponding to the largest cosine distance in the first cosine distance group.

[0133] The largest cosine distance can be characterized as d max =max(d1,d2,…,d N ).

[0134] Specifically, such as Figure 8 As shown, with the centroid e( Figure 8 Arrow 1) in the circle shown is the center of the circle. The largest cosine distance d in the first cosine distance group is... max The corresponding voiceprint feature is mapped to a radius r, and a two-dimensional circle is obtained based on the center and radius. Figure 8 (The shaded 2D circle shown).

[0135] This application utilizes the angle between high-dimensional vectors (i.e., cosine distance) to map the high-dimensional voiceprint space containing short speech voiceprint feature groups into a three-dimensional voiceprint space, and then further maps the three-dimensional voiceprint space into a two-dimensional circle, so that the area that cannot be calculated in the high-dimensional space is mapped into the area of ​​a circle that can be calculated in the two-dimensional space.

[0136] S740 maps each voiceprint feature in the first voiceprint feature group to a two-dimensional circle.

[0137] It is understandable that the radius of a two-dimensional circle is based on the maximum cosine distance d. max If the corresponding voiceprint features are determined, then it means that each voiceprint feature in the first voiceprint feature group can be mapped to the two-dimensional circle.

[0138] S750, calculate the area of ​​a two-dimensional circle.

[0139] First, based on the cosine distance formula d max =1-cosθ max Calculate cosθ max =1-d max .

[0140] Subsequently, the radius r of the two-dimensional planar circle is obtained as r = R·sinθ max =sinθ max .

[0141] Finally, the area of ​​the two-dimensional circle is calculated:

[0142] S=2πr 2 =2πsinθ max 2 =2π(1-cosθ) max 2 )=2π(1-(1-d max ) 2 )=2πd max (2-d max )

[0143] Based on this, the area of ​​a two-dimensional circle satisfies the following relationship:

[0144] S=2πd max (2-d max )

[0145] Where S is the area of ​​the two-dimensional circle, d max It is the largest cosine distance in the first cosine distance group.

[0146] S760, the first distribution density is determined based on the ratio of the area of ​​the two-dimensional circle to the number of voiceprint features contained in the first voiceprint feature group.

[0147] Right now, Where D is the first distribution density.

[0148] This application uses the area of ​​a two-dimensional circle divided by the number of voiceprint features it contains as the two-dimensional voiceprint distribution density, so that more attention can be paid to the removal of voiceprint features at the edge of the two-dimensional circle when performing feature correction in the future.

[0149] Regarding step S2 above, given the current first distribution density D, it can be specifically based on... Figure 9 The process shown determines whether the current first voiceprint feature group is accurate. For example... Figure 9 As shown, method 900 includes steps S901 to S913, which will be described in detail below.

[0150] S901, determine whether the current first distribution density D is greater than the first threshold D1.

[0151] When the current first distribution density D is greater than the first threshold D1, proceed directly to step S912, and consider the current first voiceprint feature group to be accurate. It should be understood that when the current first distribution density D is greater than the first threshold D1, it means that the voiceprint features in the current first voiceprint feature group are relatively densely distributed, so there is no need to further modify the first voiceprint feature group, and it can be directly considered that the first voiceprint feature group is accurate.

[0152] When the current first distribution density D is less than or equal to the first threshold D1, a first operation can be performed, which includes steps S902 to S910.

[0153] S902, remove p% of the voiceprint features in the current first voiceprint feature group to obtain the second voiceprint feature group, and then execute step S903.

[0154] The initial value of p is ρ, which can be, for example, a value of 10, 15 or 20, and this application does not limit it.

[0155] Among them, the cosine distance of each voiceprint feature removed from the centroid is greater than the cosine distance of any voiceprint feature in the second voiceprint feature group from the centroid. That is to say, the voiceprint features with larger corresponding cosine distances in the current first voiceprint feature group are removed.

[0156] As an example, we can first sort all voiceprint features in the current first voiceprint feature group by their cosine distance from the centroid; then, based on the sorting, remove the voiceprint features with a larger p% cosine distance in the current first voiceprint feature group.

[0157] S903, determine whether the number N1 of voiceprint features in the second voiceprint feature group is less than the second threshold N'.

[0158] When the number of voiceprint features N1 in the second voiceprint feature group is greater than or equal to the second threshold N', step S904 is executed.

[0159] When the number of voiceprint features N1 in the second voiceprint feature group is less than the second threshold N', step S907 is executed.

[0160] S904, calculate the second distribution density of the second voiceprint feature group in the two-dimensional circle corresponding to the second voiceprint feature group.

[0161] For details, please refer to the above method 700 to calculate the second distribution density D' of the second voiceprint feature group in the two-dimensional circle corresponding to the second voiceprint feature group, which will not be elaborated here.

[0162] S905, determine the relationship between the first distribution density D and the second distribution density D'.

[0163] When the second distribution density D' is greater than or equal to the first distribution density D, execute S906.

[0164] When the second distribution density D' is less than the first distribution density D, execute S909 to determine whether the current p value is equal to the initial value;

[0165] S906, update the current first voiceprint feature group based on the second voiceprint feature group, and update the current p value to the initial value.

[0166] That is to say, the second voiceprint feature group is used as the current first voiceprint feature group, and then the updated voiceprint feature group is used as the basis to re-enter step S901 to continue to determine whether the updated voiceprint feature group is accurate.

[0167] It is understandable that if the second distribution density D' is greater than or equal to the first distribution density D, it means that the voiceprint distribution density of the current first voiceprint feature group has been improved after removing the voiceprint features from the two-dimensional circular edge to obtain the second voiceprint feature group. Therefore, the current first voiceprint feature group can be updated based on the second voiceprint feature group. It should be understood that the update operation of the current first voiceprint feature group can also be understood as a correction operation of the current first voiceprint feature group.

[0168] It should be understood that, based on Figure 9 The process shown may yield an accurate judgment result without correction for the current first voiceprint feature group, or it may require one or more corrections to obtain an accurate judgment result. The specific situation needs to be determined based on the actual circumstances, and no limitation is made here.

[0169] For example, such as Figure 10 As shown, for the current first voiceprint feature group, voiceprint features v1 and v2 may be removed in the first correction; voiceprint features v3 and v4 may be removed in the second correction; and voiceprint features v5 and v6 may be removed in the third correction. By continuously removing voiceprint features from the edge of the two-dimensional circle, the voiceprint distribution density is improved, making the voiceprint features of long speech determined based on the remaining voiceprint features closer to the real voiceprint features, thus improving the accuracy of voiceprint matching.

[0170] S907, determine whether the current p value is equal to the initial value.

[0171] When the current p value is equal to the initial value, execute S908.

[0172] If the current value of p is not equal to the initial value, then a second operation is performed, which includes steps S911 to S913.

[0173] It should be understood that, for cases where the number of voiceprint features in the second voiceprint feature group is less than the second threshold, considering that this situation may be due to the removal of too many voiceprint features previously, it is possible to first determine whether the current p value is the initial value. If so, the p value can be reduced, and the first operation can be re-executed based on the reduced p value. If not, it means that the current p value has already been adjusted, and the final judgment process (i.e., the second operation) can be directly entered.

[0174] S908, reduce the p value, and return to step S902 based on the reduced p value.

[0175] Optionally, the reduced p value can be ρ / 2 or other values, without limitation.

[0176] S909, determine whether the current p value is equal to the initial value.

[0177] When the current p value is equal to the initial value, execute S910.

[0178] If the current value of p is not equal to the initial value, then a second operation is performed, which includes steps S911 to S913.

[0179] It should be understood that, in the case where the second distribution density is less than the first distribution density, we can continue to determine whether the current p value is the initial value. If it is, considering that this situation may be due to the current p value being low, we can increase the p value and re-execute the first operation based on the increased p value. If not, it means that the current p value has already been adjusted and has not yielded a better result, so we can directly enter the final judgment process (i.e., execute the second operation).

[0180] S910, increase the p value, and return to step S902 based on the increased p value.

[0181] Optionally, the increased p value can be 3ρ / 2 or other values, without limitation.

[0182] S911, determine whether the current first distribution density D is less than the third threshold D2, which is less than or equal to the first threshold D1.

[0183] When the current first distribution density D is greater than or equal to the third threshold D2, proceed to step S912.

[0184] When the current first distribution density D is less than the third threshold D2, proceed to step S913.

[0185] S912 considers the current first voiceprint feature group to be accurate.

[0186] If the current first voiceprint feature group is considered accurate, the mean of the current first voiceprint feature group can be used as the voiceprint feature of the first long speech.

[0187] S913 considers the current first voiceprint feature group to be inaccurate.

[0188] If the first voiceprint feature group is deemed inaccurate, the second long speech corresponding to the target object can be obtained, and the accuracy of the short speech voiceprint feature group corresponding to the second long speech can be determined based on the above scheme.

[0189] It should be understood that this application does not limit the threshold (e.g., a first threshold, a second threshold, or a third threshold) or the initial value of p. For example, the threshold and the initial value of p can be determined by the researchers or users based on experience.

[0190] In one possible implementation, the above accuracy judgment operation can be implemented based on an accuracy judgment module. The input data of this module is the first voiceprint feature group, and the output is the accuracy judgment result of the first voiceprint feature group. If the judgment is accurate, the voiceprint features of the corresponding long speech will also be output.

[0191] In summary, this application determines the accuracy of short speech voiceprint feature groups by measuring their distribution density in the corresponding two-dimensional circle, thereby enabling accurate judgment of long speech voiceprint features and ensuring the accuracy of voiceprint recognition.

[0192] The following, combined with Figure 11 This application describes the voiceprint feature accuracy determination device 1100 provided in the embodiments of this application. It should be understood that this device 1100 is applied to electronic devices. Figure 11 As shown, the device 1100 may include an acquisition module 1110 and a processing module 1120. The acquisition module 1110 and the processing module 1120 are used by the device 1100 to perform the corresponding processing steps in the above method embodiments. For example, the acquisition module 1110 is used to acquire a first long speech corresponding to the target object; the processing module 1120 is used to segment the first long speech into multiple short speech segments, then extract the voiceprint features of each short speech segment to obtain a first voiceprint feature group, and then perform an accuracy judgment operation based on the first voiceprint feature group. The accuracy judgment operation includes: calculating a first distribution density of the first voiceprint feature group in a two-dimensional circle corresponding to the first voiceprint feature group, and judging whether the first voiceprint feature group is accurate based on the first distribution density.

[0193] It should be understood that other related steps performed by the acquisition module 1110 and the processing module 1120 can be found in the description in the above method embodiments, and will not be repeated here.

[0194] In one possible implementation, the device 1100 may further include a storage module. This storage module is connected to the acquisition module 1110 and the processing module 1120 via a line. The storage module may include one or more memories, which can be devices in one or more devices or circuits used to store programs or data. The storage module can exist independently and be connected to the acquisition module 1110 and the processing module 1120 via a communication bus. Alternatively, the storage module may be integrated with the acquisition module 1110 and the processing module 1120.

[0195] The storage module may store computer-executable instructions for the methods in device 1100, causing device 1100 to execute the methods described in the above embodiments. The storage module may be a register, cache memory, or random access memory (RAM), etc. The storage module may also be a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions.

[0196] Figure 12 This is a schematic diagram of the structure of a chip 1200 provided in an embodiment of this application. For example... Figure 12 As shown, chip 1200 includes one or more (including two) processors 1201, communication lines 1202 and communication interfaces 1203. Optionally, chip 1200 also includes a memory 1204.

[0197] In some implementations, memory 1204 stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof.

[0198] The methods described in the embodiments of this application can be applied to processor 1201, or implemented by processor 1201. Processor 1201 may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above methods can be completed by integrated logic circuits in the hardware of processor 1201 or by instructions in software form. The processor 1201 may be a general-purpose processor (e.g., a microprocessor or conventional processor), a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components.

[0199] The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in mature storage media in the art, such as random access memory, read-only memory, programmable read-only memory, or electrically erasable programmable read-only memory (EEPROM). This storage medium is located in memory 1204, and processor 1201 reads information from memory 1204 and, in conjunction with its hardware, completes the steps of the above method.

[0200] The processor 1201, memory 1204 and communication interface 1203 can communicate with each other via communication line 1202.

[0201] In the above embodiments, the instructions stored in the memory for execution by the processor can be implemented in the form of a computer program product. This computer program product can be pre-written into the memory, or it can be downloaded and installed into the memory as software.

[0202] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. For example, available media may include magnetic media (e.g., floppy disk, hard disk, or magnetic tape), optical media (e.g., digital versatile disc (DVD)), or semiconductor media (e.g., solid-state disk (SSD)).

[0203] This application also provides a device for judging the accuracy of voiceprint features, including: one or more processors; one or more memories; the memories store one or more programs, and when one or more programs are executed by the processor, the device performs the technical solutions in the above embodiments.

[0204] This application provides a chip. The chip includes a processor, which is used to call a computer program in memory to execute the technical solutions in the above embodiments. Its implementation principle and technical effects are similar to those in the related embodiments described above, and will not be repeated here.

[0205] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program or instructions. When the computer program or instructions are executed by a processor, they implement the methods described above. The methods described in the above embodiments can be implemented wholly or partially by software, hardware, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted over the computer-readable medium. The computer-readable medium can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0206] As one possible design, computer-readable media may include compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; computer-readable media may include disk storage or other disk storage devices. Furthermore, any connecting cable may also be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. As used herein, disks and optical discs include optical discs (CD), laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0207] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0208] The above specific embodiments further illustrate the purpose, technical solution and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this application should be included within the scope of protection of this application.

Claims

1. A method for judging the accuracy of voiceprint features, characterized in that, The method is applied to an electronic device, and the method includes: Get the first long audio segment corresponding to the target object; The first long speech is divided into multiple short speech segments; Extract the voiceprint features of each short speech segment to obtain the first voiceprint feature group; An accuracy judgment operation is performed based on the first voiceprint feature group. The accuracy judgment operation includes: calculating the first distribution density of the first voiceprint feature group in the two-dimensional circle corresponding to the first voiceprint feature group. Based on the first distribution density, determine whether the first voiceprint feature group is accurate; The step of determining whether the first voiceprint feature group is accurate based on the first distribution density includes: When the first distribution density is greater than the first threshold, the first voiceprint feature group is considered accurate; When the first distribution density is less than or equal to the first threshold, a first operation is performed, the first operation including: Remove p% of the voiceprint features from the first voiceprint feature group to obtain the second voiceprint feature group. The initial value of p is ρ. The cosine distance of each voiceprint feature removed from the centroid of the first voiceprint feature group is greater than the cosine distance of any voiceprint feature in the second voiceprint feature group from the centroid. When the number of voiceprint features in the second voiceprint feature group is greater than or equal to the second threshold, calculate the second distribution density of the second voiceprint feature group in the two-dimensional circle corresponding to the second voiceprint feature group; When the second distribution density is greater than or equal to the first distribution density, the first voiceprint feature group is updated based on the second voiceprint feature group, and the current p value is updated to the initial value. Then, the accuracy judgment operation is re-executed based on the updated first voiceprint feature group.

2. The method according to claim 1, characterized in that, The calculation of the first distribution density of the first voiceprint feature group in the two-dimensional circle corresponding to the first voiceprint feature group includes: Calculate the centroid of the first voiceprint feature group; Calculate the cosine distance from each voiceprint feature in the first voiceprint feature group to the centroid to obtain the first cosine distance group; The two-dimensional circle is determined based on the voiceprint feature corresponding to the largest cosine distance in the first cosine distance group and the centroid. Each voiceprint feature in the first voiceprint feature group is mapped to the two-dimensional circle; Calculate the area of ​​the two-dimensional circle; The first distribution density is determined based on the ratio of the area of ​​the two-dimensional circle to the number of voiceprint features contained in the first voiceprint feature group.

3. The method according to claim 2, characterized in that, The step of determining the two-dimensional circle based on the voiceprint feature corresponding to the largest cosine distance in the first cosine distance group and the centroid includes: Using the centroid as the center, map the voiceprint feature corresponding to the largest cosine distance in the first cosine distance group as the radius; The two-dimensional circle is obtained based on the center and the radius.

4. The method according to claim 2, characterized in that, The area of ​​the two-dimensional circle satisfies the following relationship: in, Let the area of ​​the two-dimensional circle be . It is the largest cosine distance in the first cosine distance group.

5. The method according to claim 1, characterized in that, The first operation further includes: When the number of voiceprint features in the second voiceprint feature group is less than the second threshold, it is determined whether the current p value is equal to the initial value; When the current p value is equal to the initial value, the p value is decreased, and the first operation is re-executed based on the decreased p value; When the current value of p is not equal to the initial value, a second operation is performed, the second operation including: Determine whether the first distribution density is less than a third threshold, wherein the third threshold is less than or equal to the first threshold; If the first distribution density is less than the third threshold, then the first voiceprint feature group is considered inaccurate. When the first distribution density is greater than or equal to the third threshold, the first voiceprint feature group is considered accurate.

6. The method according to claim 5, characterized in that, The first operation further includes: When the second distribution density is less than the first distribution density, it is determined whether the current p value is equal to the initial value; When the current p value is equal to the initial value, increase the p value, and re-execute the first operation based on the increased p value; If the current value of p is not equal to the initial value, then the second operation is performed.

7. The method according to claim 6, characterized in that, The reduced p value is ρ / 2, and the increased p value is 3ρ / 2.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: When the first voiceprint feature set is deemed accurate, the mean of the first voiceprint feature set is used as the voiceprint feature of the first long speech; or... If the first voiceprint feature group is deemed inaccurate, the second long speech corresponding to the target object is obtained, and it is determined whether the short speech voiceprint feature group corresponding to the second long speech is accurate.

9. A device for judging the accuracy of voiceprint features, characterized in that, It includes at least one processor, said at least one processor being used to perform the method as described in any one of claims 1 to 8.

10. An electronic device, characterized in that, It includes at least one processor, said at least one processor being used to perform the method as described in any one of claims 1 to 8.

11. A computer-readable medium, characterized in that, Includes a computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voiceprint feature updating method and device, computer equipment and storage medium

    CN110660398A

  • Voiceprint feature validity detection method and device, and electronic equipment

    CN113571090A

  • Audio data display method and device

    CN116013326A