Voiceprint registration method, storage medium and electronic device

By clustering and confidence-based analysis of audio data from smart home devices, voiceprint registration is automatically completed, solving the problem of balancing efficiency and security in existing technologies and achieving efficient and secure voiceprint registration.

CN116612763BActive Publication Date: 2026-07-21HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
Filing Date
2022-12-29
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, voiceprint registration methods for smart home devices cannot balance efficiency and security. Unified registration faces cross-channel challenges, while individual registration is too time-consuming.

Method used

By acquiring the audio data of the first device, performing clustering processing, determining whether it meets the confidence conditions, and matching it with the audio data of registered devices, the voiceprint registration is automatically completed, ensuring the consistency of identity information.

Benefits of technology

This achieves improved voiceprint registration efficiency without the user's awareness, while ensuring the security and accuracy of the registration process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612763B_ABST
    Figure CN116612763B_ABST
Patent Text Reader

Abstract

The application discloses a voiceprint registration method, a storage medium and an electronic device, relates to the technical field of smart home / smart home, and the voiceprint registration method comprises the following steps: acquiring first audio data collected by a first device; in the case that the first audio data meets a first confidence condition corresponding to the first device, acquiring second audio data corresponding to at least one second device, wherein the second device is a device that has performed voiceprint registration through the second audio data; and in the case that the first audio data meets a second confidence condition corresponding to the second audio data, performing voiceprint registration on the first device based on the first audio data. The application solves the technical problem that the efficiency and safety of voiceprint registration cannot be considered in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home / intelligent home technology, and more specifically, to a voiceprint registration method, storage medium, and electronic device. Background Technology

[0002] In intelligent voice dialogue systems, one feasible method for different smart home devices is to register their respective voiceprint information. However, this method has limited effectiveness, is time-consuming, and inefficient. Another feasible method is to register the corresponding voiceprint information uniformly. However, on the one hand, different smart home devices cannot register voiceprint information uniformly when they use different voiceprint channels; on the other hand, even without considering cross-channel technical barriers, this method still involves mapping the same voice data to user identity information for different smart home devices, reducing user privacy and thus affecting the security of voiceprint registration. Therefore, related technologies suffer from the problem of not being able to balance the efficiency and security of voiceprint registration. Summary of the Invention

[0003] This application provides a voiceprint registration method, storage medium, and electronic device to at least solve the technical problem in the related art that cannot simultaneously achieve both voiceprint registration efficiency and security.

[0004] According to one aspect of the embodiments of this application, a voiceprint registration method is provided, comprising: acquiring first audio data collected by a first device; acquiring second audio data corresponding to at least one second device when the first audio data satisfies a first confidence condition corresponding to the first device, wherein the second device is a device that has already performed the voiceprint registration through the second audio data; and performing the voiceprint registration on the first device based on the first audio data when the first audio data satisfies a second confidence condition corresponding to the second audio data.

[0005] According to another aspect of the embodiments of this application, a voiceprint registration apparatus is also provided, comprising: a first acquisition unit, configured to acquire first audio data collected by a first device; a second acquisition unit, configured to acquire second audio data corresponding to at least one second device when the first audio data satisfies a first confidence condition corresponding to the first device, wherein the second device is a device that has already performed the voiceprint registration through the second audio data; and a registration unit, configured to perform the voiceprint registration on the first device based on the first audio data when the first audio data satisfies a second confidence condition corresponding to the second audio data.

[0006] As an optional solution, the above-mentioned device further includes: a first acquisition module, configured to acquire a first voiceprint object corresponding to the first audio data and at least one voiceprint object corresponding to the second audio data before performing the voiceprint registration on the first device based on the first audio data; and a determination module, configured to determine that the first audio data satisfies the second confidence condition if the first voiceprint object satisfies the confidence condition corresponding to the at least one voiceprint object before performing the voiceprint registration on the first device based on the first audio data.

[0007] As an optional solution, the determination module includes: a determination submodule, used to determine that the first audio data satisfies the second confidence condition when a second voiceprint object with a similarity greater than a second preset threshold is determined from the at least one voiceprint object.

[0008] As an optional solution, the first acquisition unit includes: a clustering module, used to perform clustering processing on the multiple audio data obtained to obtain at least one audio data set, wherein each data set in the at least one audio data set records audio data of the same type, and the first audio data is the same type of audio data obtained by the first device after performing the clustering processing on the multiple audio data collected; and a second acquisition module, used to determine the target audio data set from the at least one audio data set and acquire the audio data recorded in the target audio data set, wherein the audio data recorded in the target audio data set includes the first audio data.

[0009] As an optional solution, the above-mentioned device further includes: a third acquisition module, used to perform audio quality calculation on the new audio data when new audio data is recorded in the target audio data set, to obtain the quality confidence level corresponding to the new audio data; and a confidence module, used to determine that the first audio data satisfies the first confidence condition when the quality confidence level is greater than or equal to a preset quality threshold and the at least one audio data set contains audio data belonging to the same type of audio data as the new audio data.

[0010] As an optional solution, the above-mentioned device further includes: a determining module, configured to determine, before acquiring the second audio data corresponding to at least one second device, that the first audio data satisfies the first confidence condition if the amount of the first audio data is greater than or equal to a first preset threshold.

[0011] As an optional solution, the registration unit includes a registration module for registering the voiceprint of the first device using the first audio data and the second audio data.

[0012] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voiceprint registration method as described above.

[0013] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described voiceprint registration method through the computer program.

[0014] In this embodiment, first audio data collected by a first device is acquired; if the first audio data satisfies a first confidence condition corresponding to the first device, second audio data corresponding to at least one second device is acquired, wherein the second device is a device that has performed the voiceprint registration through the second audio data; if the first audio data satisfies a second confidence condition corresponding to the second audio data, the first device is registered for voiceprint based on the first audio data.

[0015] On the one hand, after the first device performs clustering processing, the first audio data of the same type is automatically judged to meet the first confidence condition and the second confidence condition in turn. The voiceprint registration of the first device is completed without the user's awareness, without the need for the user to perform additional voiceprint registration operations, thereby improving the voiceprint registration efficiency.

[0016] On the other hand, if the first audio data satisfies the first confidence condition, it is further determined whether the first audio data satisfies the second confidence condition of the second device that has been registered with voiceprints, so as to ensure that the identity information corresponding to the voiceprint information of the second device matches the first audio data. Based on the first audio, the first device is registered with voiceprints, thereby ensuring the accuracy of the identity information of the voiceprint registration of the first device and achieving the purpose of ensuring the security of voiceprint registration.

[0017] Thus, by using the above-mentioned voiceprint registration method, the goal of filtering high-confidence audio on different devices and registering already registered users in the background on different devices is achieved without the user's awareness. This achieves a technical effect that balances the efficiency and security of voiceprint registration, and solves the technical problem in related technologies that cannot balance the efficiency and security of voiceprint registration. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the hardware environment for a voiceprint registration method according to an embodiment of this application;

[0021] Figure 2 This is a schematic diagram of the flow of an optional voiceprint registration method according to an embodiment of this application;

[0022] Figure 3 This is a schematic diagram of an optional voiceprint registration method according to an embodiment of this application;

[0023] Figure 4 This is a schematic diagram of another optional voiceprint registration method according to an embodiment of this application;

[0024] Figure 5 This is a schematic diagram of another optional voiceprint registration method according to an embodiment of this application;

[0025] Figure 6 This is a schematic diagram of an optional information processing apparatus according to an embodiment of the present invention;

[0026] Figure 7 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] According to one aspect of the embodiments of this application, an interaction method for smart home devices is provided. This interaction method for smart home devices is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligencehouse ecosystems. Optionally, in this embodiment, the above-mentioned interaction method for smart home devices can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.

[0030] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.

[0031] Alternatively, as an alternative implementation method, such as Figure 2 As shown, the voiceprint registration method includes:

[0032] S202, acquire the first audio data collected by the first device;

[0033] S204, if the first audio data satisfies the first confidence condition corresponding to the first device, obtain the second audio data corresponding to at least one second device, wherein the second device is a device that has been registered with voiceprint through the second audio data;

[0034] S206, if the first audio data satisfies the second confidence condition corresponding to the second audio data, the first device is registered with a voiceprint based on the first audio data.

[0035] Optionally, in this embodiment, the execution entity of the above-mentioned voiceprint registration method may be, but is not limited to, a local client or a cloud server. Taking a cloud server as an example, the cloud server obtains the first audio data collected by the first device and determines whether the first audio data meets the first confidence condition corresponding to the first device. If the first audio data meets the first confidence condition corresponding to the first device, it obtains the second audio data corresponding to at least one second device. It further determines whether the first audio data meets the second confidence condition corresponding to the second audio data, and if the first audio data meets the second confidence condition corresponding to the second audio data, it performs voiceprint registration on the first device based on the first audio data.

[0036] In addition, the first device may also perform a judgment on whether the first audio data meets the first confidence condition corresponding to the first device, and then the cloud server obtains the registration request reported by the first device. The registration request is used to indicate that the first audio data collected by the first device meets the first confidence condition and to request voiceprint registration for the first device; further, the second audio data is obtained, and if the second audio data meets the second confidence condition, the voiceprint registration for the first device is performed.

[0037] Optionally, in this embodiment, the above-described voiceprint registration method can be applied, but is not limited to, to a smart dialogue system in a smart home environment. In a smart voice dialogue system, users can pre-register voiceprint information corresponding to controlling smart home devices to perform operations. When the same registered voiceprint information is emitted again, the smart home device automatically performs the corresponding operation, facilitating user operation and improving the user experience.

[0038] However, intelligent voice dialogue systems often involve multiple different smart home devices, such as smart water heaters, smart TVs, and smart refrigerators. Furthermore, due to the differences in front-end signal processing methods and the devices themselves, it is difficult to train a single voiceprint registration model to adapt to the audio characteristics of all device models, let alone to achieve unified voiceprint registration for multiple different smart home devices.

[0039] It should be noted that registering a unified voiceprint for multiple different smart home devices is insecure. Voiceprint information is the sound wave spectrum carrying speech information. Speakers' voices can vary depending on their environment, mindset, tone of voice, and health status. Therefore, if a user registers a unified voiceprint for multiple smart home devices and an error occurs in one voiceprint, it will lead to errors in the voiceprint information for all smart home devices.

[0040] However, if voiceprint registration is performed separately for each smart home device, a corresponding voiceprint registration model needs to be trained for each smart home device, which has limited effectiveness, takes too long, and is inefficient.

[0041] To address the insecurity of unified voiceprint registration and the inefficiency of individual voiceprint registration, the aforementioned voiceprint registration method uses the following approach: After clustering multiple audio data collected by a first device that has not yet registered its voiceprint, if the first audio data of the same type of audio data satisfies the first confidence condition corresponding to the first device, it automatically determines whether the first audio data and the second audio data corresponding to the second device that has already registered its voiceprint satisfy the second confidence condition, and then decides whether to register the voiceprint of the first device based on the first audio data.

[0042] On the one hand, after the first device performs clustering processing, the first audio data of the same type is automatically judged to meet the first confidence condition and the second confidence condition in turn. The voiceprint registration of the first device is completed without the user's awareness, without the need for the user to perform additional voiceprint registration operations, thereby improving the voiceprint registration efficiency.

[0043] On the other hand, if the first audio data satisfies the first confidence condition, it is further determined whether the first audio data satisfies the second confidence condition of the second device that has been registered with voiceprints, so as to ensure that the identity information corresponding to the voiceprint information of the second device matches the first audio data. Based on the first audio, the first device is registered with voiceprints, thereby ensuring the accuracy of the identity information of the voiceprint registration of the first device and achieving the purpose of ensuring the security of voiceprint registration.

[0044] Thus, by using the above-mentioned voiceprint registration method, the goal of selecting high-confidence audio on different devices and registering already registered users in the background without their awareness is achieved, thus realizing a technical effect that balances the efficiency and security of voiceprint registration.

[0045] Optionally, in this embodiment, the first device may be, but is not limited to, a smart home device that has not been registered with its voiceprint, and may include, but is not limited to, smart water heaters, smart TVs, smart refrigerators, etc. that have not been registered with their voiceprints; the second device may be, but is not limited to, a client that has been registered with its voiceprint, and may include, but is not limited to, a user's mobile client that has been registered with its voiceprint, and may also include, but is not limited to, a smart home device that has been registered with its voiceprint.

[0046] Optionally, in this embodiment, before obtaining the first audio data collected by the first device that has not performed voiceprint registration, the first device performs clustering processing based on the collected multiple audio data.

[0047] Optionally, in this embodiment, the clustering process may include, but is not limited to, the following: the first device accumulates audio data of various categories of the expected data volume, and obtains models under different audio types based on the clustering algorithm, wherein each model is used to indicate the identity information of a user, and the audio data types corresponding to each model are the same.

[0048] Optionally, in this embodiment, the clustering algorithm may be used, but is not limited to, to aggregate audio data of the same type into the same audio data set. The clustering algorithm may include, but is not limited to, the kmeans algorithm and the meanshift algorithm. The audio data of the same type may be used, but is not limited to, to indicate audio data corresponding to the same user identity information.

[0049] Optionally, in this embodiment, the first audio data may be, but is not limited to, similar audio data obtained by the first device after clustering multiple collected audio data. The first audio data may include, but is not limited to, multiple voiceprint objects, wherein multiple voiceprint objects correspond to the same user's identity information, and each voiceprint object corresponds to the voiceprint instruction of the user with the corresponding identity information. The voiceprint instruction is used to control the first device to perform the corresponding operation.

[0050] Optionally, in this embodiment, satisfying the first confidence condition corresponding to the first device may be, but is not limited to, having a data volume of the voiceprint object used to indicate the first audio data that is greater than or equal to a first preset threshold.

[0051] Optionally, in this embodiment, satisfying the first confidence condition corresponding to the first device may also be used, but is not limited to, to indicate the appearance of new audio data in the first device. The new audio data may be, but is not limited to, audio data used to indicate a new control command corresponding to the first device.

[0052] It should be noted that the control commands corresponding to the new audio data in the first device are appearing for the first time in the first device, but the corresponding identity information may or may not be appearing for the first time in the first device.

[0053] Optionally, in this embodiment, if the first audio satisfies the first confidence condition corresponding to the first device, second audio data corresponding to at least one second device is obtained, wherein the second audio data may be, but is not limited to, audio data used for voiceprint registration of the second device.

[0054] Optionally, in this embodiment, satisfying the second confidence condition corresponding to the second audio can be, but is not limited to, used to indicate that the identity information corresponding to the first audio data is substantially the same as the identity information corresponding to the second audio data.

[0055] Optionally, in this embodiment, registering the voiceprint of the first device based on the first audio data can be, but is not limited to, registering the voiceprint of the first device using the first audio data and the second audio data.

[0056] It should be noted that, under the second confidence condition corresponding to the second audio data, the identity information corresponding to the first audio data is basically the same as the identity information corresponding to the second audio data. Thus, the first audio data and the second audio data correspond to the same user. The second audio data used by this user to register the second device is also applicable to the first device. The first audio data and the second audio data are used to register the voiceprint of the first device.

[0057] To further illustrate, such as Figure 3 As shown, a voiceprint registration method specifically includes the following steps:

[0058] In step S302, the first device 302 performs clustering processing on the collected multiple audio data to obtain multiple audio data sets, wherein each audio data set corresponds to the identity information of a user, and the first device 302 is a device that has not performed voiceprint registration;

[0059] Step S304: Obtain the first audio data acquired by the first device 302;

[0060] Step S306: Determine whether the first audio data satisfies the first confidence condition corresponding to the first device 302. If it satisfies, proceed to step S308. If it does not satisfy, no additional operation is performed. The first confidence condition corresponding to the first device 302 is used to indicate that the number of voiceprint objects in the first audio data is greater than or equal to the first preset threshold, or to indicate that the first audio data includes new audio data from the first device 302.

[0061] Step S308: If the first audio data satisfies the first confidence condition corresponding to the first device 302, obtain the second audio data corresponding to at least one second device, wherein the second device is a device that has been registered with voiceprint through the second audio data.

[0062] Step S310: Determine whether the first audio data satisfies the second confidence condition corresponding to the second audio data. If it satisfies, proceed to step S312. If it does not satisfy, no additional operation is performed. The second confidence condition corresponding to the second audio data is used to indicate that the identity information corresponding to the first audio data is basically consistent with the identity information corresponding to the second audio data.

[0063] Step S312: If the first audio data satisfies the second confidence condition corresponding to the second audio data, the first device is registered with a voiceprint based on the first audio data.

[0064] The embodiments provided in this application obtain first audio data collected by a first device; when the first audio data satisfies a first confidence condition corresponding to the first device, second audio data corresponding to at least one second device is obtained, wherein the second device is a device that has been registered for voiceprint based on the second audio data; when the first audio data satisfies a second confidence condition corresponding to the second audio data, voiceprint registration is performed on the first device based on the first audio data.

[0065] On the one hand, after the first device performs clustering processing, the first audio data of the same type is automatically judged to meet the first confidence condition and the second confidence condition in turn. The voiceprint registration of the first device is completed without the user's awareness, without the need for the user to perform additional voiceprint registration operations, thereby improving the voiceprint registration efficiency.

[0066] On the other hand, if the first audio data satisfies the first confidence condition, it is further determined whether the first audio data satisfies the second confidence condition of the second device that has been registered with voiceprints, so as to ensure that the identity information corresponding to the voiceprint information of the second device matches the first audio data. Based on the first audio, the first device is registered with voiceprints, thereby ensuring the accuracy of the identity information of the voiceprint registration of the first device and achieving the purpose of ensuring the security of voiceprint registration.

[0067] Thus, by using the above-mentioned voiceprint registration method, the goal of selecting high-confidence audio on different devices and registering already registered users in the background without their awareness is achieved, thus realizing a technical effect that balances the efficiency and security of voiceprint registration.

[0068] As an optional approach, before acquiring the second audio data corresponding to at least one second device, the method further includes:

[0069] S1, if the amount of the first audio data is greater than or equal to the first preset threshold, determine that the first audio data satisfies the first confidence condition.

[0070] Optionally, in this embodiment, the amount of data in the first audio data may be, but is not limited to, used to indicate the number of voiceprint objects under the first audio data. The first audio data may include, but is not limited to, multiple voiceprint objects, multiple voiceprint objects corresponding to the same audio data type, multiple voiceprint objects corresponding to the same user's identity information, and each voiceprint object may, but is not limited to, correspond to the user's control command to the first device.

[0071] It should be noted that after the first device continuously accumulates audio data and completes clustering processing to obtain models under multiple audio data corresponding to multiple identity information, when the amount of first audio data used to indicate a certain identity information is greater than or equal to a first preset threshold, it is determined that the first audio data meets the first confidence condition. It automatically verifies whether the identity information corresponding to the first audio data is basically consistent with the identity information corresponding to the second audio data of the registered second device. If so, it automatically registers the voiceprint of the first device based on the first audio data, thus achieving the technical effect of safely and accurately completing the voiceprint registration of the first device without the user's awareness.

[0072] The embodiments provided in this application determine that the first audio data satisfies a first confidence condition when the amount of first audio data is greater than or equal to a first preset threshold. By automatically verifying whether the identity information corresponding to the audio data matches the identity information corresponding to the audio data of a registered second device when the amount of a certain type of audio data collected by the first device is greater than or equal to a certain quantity, and if so, voiceprint registration is performed based on the collected audio data of that type. This achieves the goal of improving the efficiency and security of voiceprint registration without the user's awareness, thus realizing a technical effect that balances the efficiency and security of voiceprint registration.

[0073] As an optional approach, before registering the voiceprint of the first device based on the first audio data, the method further includes:

[0074] S1, obtain the first voiceprint object corresponding to the first audio data, and at least one voiceprint object corresponding to the second audio data;

[0075] S2, if the first voiceprint object satisfies the confidence condition corresponding to at least one voiceprint object, determine that the first audio data satisfies the second confidence condition.

[0076] Optionally, in this embodiment, the first audio data may include, but is not limited to, multiple voiceprint objects, wherein the multiple voiceprint objects include the first voiceprint object.

[0077] Optionally, in this embodiment, the second audio data may include, but is not limited to, at least one voiceprint object, wherein the at least one voiceprint object may be used for voiceprint registration by the second device.

[0078] It should be noted that if the first voiceprint object satisfies at least one confidence condition corresponding to a voiceprint object, the first audio data is determined to satisfy the second confidence condition.

[0079] The embodiments provided in this application obtain a first voiceprint object corresponding to first audio data and at least one voiceprint object corresponding to second audio data. If the first voiceprint object satisfies the confidence condition corresponding to at least one voiceprint object, the first audio data is determined to satisfy a second confidence condition. By judging whether the voiceprint object corresponding to the first audio data and the voiceprint object corresponding to the second audio data satisfy the confidence condition, it is determined whether the first audio data satisfies the second confidence condition, thereby determining whether the identity information corresponding to the first audio data is substantially consistent with the identity information corresponding to the second audio data. This achieves the technical effect of balancing the efficiency and security of voiceprint registration by registering the first device's voiceprint based on the first audio data, provided that the first audio data of the first device is substantially consistent with the identity information of the registered second device.

[0080] As an optional approach, if the first voiceprint object satisfies at least one confidence condition corresponding to the voiceprint object, determining that the first audio data satisfies the second confidence condition includes:

[0081] S1, if a second voiceprint object is identified from at least one voiceprint object whose similarity to the first voiceprint object is greater than a second preset threshold, the first audio data is determined to satisfy the second confidence condition.

[0082] It should be noted that if a second voiceprint object is determined from at least one voiceprint object in the second audio data of the registered second device, and the similarity between the second voiceprint object and the first voiceprint object in the first audio data of the first device is greater than a second preset threshold, then the first audio data is determined to satisfy the second confidence condition.

[0083] Optionally, in this embodiment, the similarity between voiceprint objects can be, but is not limited to, used to indicate the similarity between the identity information indicated by the voiceprint objects.

[0084] Optionally, in this embodiment, the similarity between the first voiceprint object and each voiceprint object in the audio list containing at least one voiceprint object in the second audio data can be determined by mutual mapping comparison between the first voiceprint object and the audio list in the second video data.

[0085] Optionally, in this embodiment, the similarity between the first voiceprint object and each voiceprint object in the second audio data can be determined by, but is not limited to, the voiceprint feature vectors (speaker embeddings) corresponding to each voiceprint object in the first voiceprint object and each voiceprint object in the second audio data.

[0086] To further illustrate, such as Figure 4 As shown, the first voiceprint object 402 corresponds to the first voiceprint feature vector 404, and the second voiceprint object 406 corresponds to the second voiceprint feature vector 408. The lengths of the first voiceprint object 402 and the second voiceprint object 406 can be the same or different; the vector lengths of the first voiceprint feature vector 404 and the second voiceprint feature vector 408 are the same. Furthermore, based on a similarity comparison between the first voiceprint feature vector 404 and the second voiceprint feature vector 408, a similarity result 410 is obtained, which is used to identify the similarity between the first voiceprint object 402 and the second voiceprint object 406.

[0087] It should be noted that this embodiment does not limit the method of comparing the similarity between voiceprint feature vectors. It can be understood that methods such as using Euclidean distance and cosine distance are all optional methods for comparing the similarity of voiceprint feature vectors in this embodiment.

[0088] According to the embodiments provided in this application, when a second voiceprint object is identified from at least one voiceprint object whose similarity to the first voiceprint object is greater than a second preset threshold, the first audio data is determined to satisfy a second confidence condition. If the first audio data satisfies the first confidence condition, the similarity between the first voiceprint object in the first audio data and at least one voiceprint object in the second audio data is further used to determine whether the similarity between the first audio data and the second audio data satisfies the second confidence condition. This allows for the determination of whether the user identity information corresponding to the first audio data and the second audio data is consistent. Thus, while improving the efficiency of voiceprint registration based on the first confidence condition, the second confidence condition ensures the accuracy of the identity information registered through voiceprint registration, thereby achieving security for voiceprint registration.

[0089] As an optional approach, the first audio data acquired by the first device includes:

[0090] S1, when multiple audio data are acquired, the multiple audio data are clustered to obtain at least one audio data set, wherein each data set in the at least one audio data set records audio data of the same type, and the first audio data is the same type of audio data obtained by the first device after clustering the multiple audio data acquired.

[0091] S2, determine a target audio data set from at least one audio data set, and obtain the audio data recorded in the target audio data set, wherein the audio data recorded in the target audio data set includes the first audio data.

[0092] Optionally, in this embodiment, the clustering process may include, but is not limited to, determining at least one audio data set based on the clustering algorithm when the first device acquires multiple audio data sets, wherein the data sets in the same audio data set have the same type of audio data, which may be used, but is not limited to, to indicate the identity information corresponding to the same user.

[0093] It should be noted that the clustering algorithm may include, but is not limited to, the kmeans algorithm and the meanshift algorithm. This embodiment does not impose any additional restrictions on this.

[0094] According to the embodiments provided in this application, when multiple audio data are acquired, the multiple audio data are clustered to obtain at least one audio data set. Each data set in the at least one audio data set records audio data of the same type. The first audio data is the same type of audio data obtained by the first device after clustering the acquired multiple audio data. A target audio data set is determined from the at least one audio data set, and the audio data recorded in the target audio data set is obtained. The audio data recorded in the target audio data set includes the first audio data.

[0095] As an optional approach, after determining the target audio dataset from at least one audio dataset and obtaining the audio data recorded in the target audio dataset, the method further includes:

[0096] S1, if new audio data is recorded in the target audio data set, calculate the audio quality of the new audio data to obtain the quality confidence level corresponding to the new audio data;

[0097] S2, if the quality confidence level is greater than or equal to a preset quality threshold and at least one audio data set contains audio data belonging to the same type of audio data as the new audio data, determine that the first audio data satisfies the first confidence condition.

[0098] Optionally, in this embodiment, the new audio data may be, but is not limited to, audio data corresponding to the new device control command that is not included in the target audio data set.

[0099] Optionally, in this embodiment, when multiple audio data are acquired, the multiple audio data are clustered to obtain at least one audio data set. A target audio data set is determined from the at least one audio data set, and the audio data recorded in the target audio data set is acquired. When new audio data appears in the first device, the model under different audio obtained by the clustering process is used to determine the identity information corresponding to the new audio data.

[0100] It should be noted that when the identity information corresponding to the new audio data is consistent with the identity information corresponding to the first audio data, the new audio data is added to the target audio data set to confirm whether the new audio data is recorded in the target data set, and to confirm that the first audio data meets the first confidence condition.

[0101] Optionally, in this embodiment, after determining the identity information corresponding to the new audio data using the model obtained from different audio data through clustering, it is also possible, but not limited to, obtaining the identity verification score corresponding to the new audio data. If the identity verification score is greater than the preset identity verification threshold, it is determined that the new audio data is of good quality and has high identity verification accuracy.

[0102] Optionally, in this embodiment, if the authentication score is less than or equal to a preset authentication threshold, new audio data may be removed, but is not limited to this.

[0103] Through the embodiments provided in this application, when new audio data is recorded in the target audio data set, the audio quality of the new audio data is calculated to obtain the quality confidence level corresponding to the new audio data; when the quality confidence level is greater than or equal to a preset quality threshold and at least one audio data set contains audio data of the same type as the new audio data, it is determined that the first audio data satisfies the first confidence condition, thereby achieving the goal of improving the efficiency and security of voiceprint registration without the user's awareness, thus realizing the technical effect of balancing the efficiency and security of voiceprint registration.

[0104] As an optional approach, voiceprint registration of the first device based on the first audio data includes:

[0105] S1, register the voiceprint of the first device using the first audio data and the second audio data.

[0106] It should be noted that if the first audio data satisfies the second confidence condition corresponding to the second audio data, the identity information corresponding to the first audio data is basically the same as the identity information corresponding to the second audio data. Thus, voiceprint registration of the first device based on the first audio data can be performed using, but is not limited to, the first audio data and the second audio data.

[0107] The embodiments provided in this application utilize first audio data and second audio data to register the voiceprint of a first device. By combining the voiceprint corresponding to the first audio data of an unregistered first device with the voiceprint corresponding to the second audio data of a registered second device, the aim is to filter high-confidence audio from different devices and perform background registration for already registered users on different devices without the user's awareness, thus achieving a technical effect that balances the efficiency and security of voiceprint registration.

[0108] As an alternative approach, the above voiceprint registration method can be applied to a scenario where a combination of unregistered and registered voiceprints is used for voiceprint registration of smart home devices. The specific steps are as follows:

[0109] Step S502: The smart home devices accumulate the expected amount of audio data;

[0110] Step S504: The smart home device performs a clustering algorithm. Through the clustering algorithm of unregistered voiceprints, models under different audio conditions are obtained, and each model represents a speaker.

[0111] Step S506: Determine whether clustering was successful;

[0112] Step S508: If clustering fails, further accumulate audio data to reach the specified amount of new data (30 for example).

[0113] Step S510: Determine whether a clustered model already exists;

[0114] Step S512: Given an existing clustered model, re-cluster the first 30 audios of each already clustered model with the newly added 30 audios.

[0115] Step S514: In the absence of a clustered model, re-cluster the original audio with the 30 newly added audio tracks;

[0116] Step S516: After successful clustering, new audio is expected to be input;

[0117] Step S518: Use an unregistered voiceprint clustering model to verify which speaker the new audio belongs to and determine whether the model identification was successful.

[0118] Step S520: If the model successfully identifies the speaker to whom the new audio belongs, obtain the score of the audio.

[0119] It should be noted that, based on a pre-defined score threshold, if the audio exceeds the threshold, it is considered to be of good quality.

[0120] It should be noted that if the model fails to identify the speaker to which the new audio belongs, it will proceed to step S508 to continue clustering and training the model.

[0121] Step S522: The user registers their voiceprint by collecting data from the client.

[0122] Step S524: Determine whether the registered voiceprint is from a registered user;

[0123] In step S526, the new audio is then judged by the registered voiceprint to see if it belongs to a speaker who has already registered. At the same time, the confidence level is obtained and it is determined whether the confidence level exceeds a specified threshold. If it exceeds the threshold, the audio quality is considered to be good and it is more in line with the speaker. After the recognition of unregistered voiceprints and the verification of registered voiceprints are satisfied, it is determined whether the speakerid of the unregistered voiceprint pointed to by the audio is the same speaker as the speaker of the registered voiceprint.

[0124] Step S528: If it is the same speaker, record and save the audio, and add the audio to the list of updated audio.

[0125] Step S530: Determine whether the number of audios in the list of updated audios has reached the specified number;

[0126] In step S532, if the number of audios in the updated audio list reaches a specified number, the audios in the updated audio list are concatenated with the original registered audios (the registered audios in step S522) to form new registered audios.

[0127] Step S534: If the number of audios in the updated audio list has not reached the specified number, exit;

[0128] Step S536: Exit if the speakers are not the same.

[0129] Step S538: If the registered voiceprint is not identified as belonging to a registered user, exit the registration process.

[0130] It should be noted that, by using the above-mentioned voiceprint registration method, on the one hand, the problem of cross-channel identification between different devices can be solved by combining unregistered and registered voiceprints; on the other hand, by adding unregistered voiceprints, the amount of speaker audio data from different devices can be increased indefinitely, which helps to improve recognition performance.

[0131] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0132] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0133] According to another aspect of the embodiments of this application, a voiceprint registration apparatus for implementing the above-described voiceprint registration method is also provided. For example... Figure 6 As shown, the device includes:

[0134] The first acquisition unit 602 is used to acquire the first audio data collected by the first device, wherein the first audio data is the same type of audio data obtained by the first device after performing clustering processing on multiple collected audio data;

[0135] The second acquisition unit 604 is used to acquire at least one second audio data corresponding to a second device when the first audio data satisfies the first confidence condition corresponding to the first device, wherein the second device is a device that has been registered with voiceprint through the second audio data.

[0136] The registration unit 606 is used to register the voiceprint of the first device based on the first audio data when the first audio data satisfies the second confidence condition corresponding to the second audio data.

[0137] As an optional solution, the above-mentioned device further includes: a determining module, configured to determine that the first audio data satisfies a first confidence condition if the amount of the first audio data is greater than or equal to a first preset threshold before acquiring the second audio data corresponding to at least one second device.

[0138] As an optional solution, the above-mentioned device further includes: a first acquisition module, configured to acquire a first voiceprint object corresponding to the first audio data and at least one voiceprint object corresponding to the second audio data before registering the voiceprint of the first device based on the first audio data; and a determination module, configured to determine that the first audio data satisfies a second confidence condition if the first voiceprint object satisfies a confidence condition corresponding to at least one voiceprint object before registering the voiceprint of the first device based on the first audio data.

[0139] As an optional solution, the above-mentioned determining module includes: a determining submodule, used to determine that the first audio data satisfies the second confidence condition when a second voiceprint object with a similarity greater than a second preset threshold is determined from at least one voiceprint object.

[0140] As an optional solution, the first acquisition unit 602 includes: a clustering module, used to cluster multiple audio data when multiple audio data are acquired, to obtain at least one audio data set, wherein each data set in the at least one audio data set records audio data of the same type; and a second acquisition module, used to determine a target audio data set from the at least one audio data set and acquire the audio data recorded in the target audio data set, wherein the audio data recorded in the target audio data set includes the first audio data.

[0141] As an optional solution, the above-mentioned apparatus further includes: a third acquisition module, configured to determine that the first audio data satisfies the first confidence condition after determining the target audio data set from at least one audio data set and acquiring the audio data recorded in the target audio data set, and in the case that new audio data is recorded in the target audio data set.

[0142] As an optional solution, the registration unit 606 includes: a registration module, used to register the voiceprint of the first device using the first audio data and the second audio data.

[0143] For specific implementation examples, please refer to the examples shown in the above voiceprint registration method; these will not be repeated here.

[0144] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described voiceprint registration method is also provided, such as... Figure 7 As shown, the electronic device includes a memory 702 and a processor 704. The memory 702 stores a computer program, and the processor 704 is configured to execute the steps of any of the above method embodiments via the computer program.

[0145] Optionally, in this embodiment, the electronic device may be located in at least one of a plurality of network devices in a computer network.

[0146] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0147] S1, acquire the first audio data collected by the first device;

[0148] S2, if the first audio data satisfies the first confidence condition corresponding to the first device, obtain the second audio data corresponding to at least one second device, wherein the second device is a device that has been registered with voiceprint through the second audio data;

[0149] S3, if the first audio data satisfies the second confidence condition corresponding to the second audio data, the first device is registered with a voiceprint based on the first audio data.

[0150] Alternatively, as those skilled in the art will understand, Figure 7 The structure shown is for illustrative purposes only. The electronic device can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, and other terminal devices. Figure 7 This does not limit the structure of the aforementioned electronic device. For example, the electronic device may also include components that are more... Figure 7 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 7 The different configurations shown.

[0151] The memory 702 can be used to store software programs and modules, such as the program instructions / modules corresponding to the voiceprint registration method and apparatus in this embodiment. The processor 704 executes various functional applications and data processing by running the software programs and modules stored in the memory 702, thereby implementing the aforementioned voiceprint registration method. The memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 702 may further include memory remotely located relative to the processor 704, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 702 may be used, but is not limited to, to store information such as a first confidence condition and a second confidence condition. As an example, such as... Figure 7As shown, the memory 702 may include, but is not limited to, the first acquisition unit 602, the second acquisition unit 604, and the registration unit 606 of the voiceprint registration device. Furthermore, it may include, but is not limited to, other module units of the voiceprint registration device, which will not be elaborated upon in this example.

[0152] Optionally, the transmission device 706 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 706 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 706 is a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0153] In addition, the above-mentioned electronic device also includes: a display 708 for displaying information such as the first confidence condition and the second confidence condition; and a connection bus 710 for connecting the various module components in the above-mentioned electronic device.

[0154] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.

[0155] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.

[0156] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0157] It should be noted that the computer system of the electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0158] A computer system includes a Central Processing Unit (CPU), which performs various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) or loaded from RAM. ROM also stores various programs and data required for system operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output interfaces (I / O interfaces) are also connected to the bus.

[0159] The following components are connected to the input / output interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard drives; and communication sections including network interface cards such as LAN cards and modems. The communication section performs communication processing via a network such as the Internet. Drives are also connected to the input / output interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required.

[0160] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions defined in the system of this application.

[0161] According to one aspect of this application, a computer-readable storage medium is provided, wherein a processor of a computer device reads computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.

[0162] Optionally, in this embodiment, the computer-readable storage medium described above may be configured to store a computer program for performing the following steps:

[0163] S1, acquire the first audio data collected by the first device;

[0164] S2, if the first audio data satisfies the first confidence condition corresponding to the first device, obtain the second audio data corresponding to at least one second device, wherein the second device is a device that has been registered with voiceprint through the second audio data;

[0165] S3, if the first audio data satisfies the second confidence condition corresponding to the second audio data, the first device is registered with a voiceprint based on the first audio data.

[0166] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0167] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0168] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0169] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0170] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0171] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0172] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0173] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A voiceprint registration method, characterized in that, include: When a smart home device acquires a cumulative expected number of audio data, the smart home device performs clustering processing on the multiple audio data using a non-registered voiceprint clustering algorithm to obtain at least one audio data set. Each data set in the at least one audio data set records audio data of the same type. Different audio data sets correspond to different models under different audio. Each model represents a speaker. The first audio data is the same type of audio data obtained by the first device after performing the clustering processing on the multiple audio data collected. A target audio data set is determined from the at least one audio data set, and audio data recorded in the target audio data set is obtained, wherein the audio data recorded in the target audio data set includes first audio data; If new audio data is recorded in the target audio data set, the audio quality of the new audio data is calculated to obtain the quality confidence level corresponding to the new audio data. If the quality confidence level is greater than or equal to a preset quality threshold, and the at least one audio data set contains audio data belonging to the same type of audio data as the new audio data, then the first audio data is determined to satisfy the first confidence condition corresponding to the first device. If the first audio data satisfies the first confidence condition, second audio data corresponding to at least one second device is obtained, wherein the second device is a device that has performed the voiceprint registration through the second audio data; If the first audio data satisfies the second confidence condition corresponding to the second audio data, the first device performs the voiceprint registration based on the first audio data; The step of registering the voiceprint of the first device based on the first audio data includes: registering the voiceprint of the first device using the first audio data and the second audio data.

2. The method according to claim 1, characterized in that, Before performing the voiceprint registration on the first device based on the first audio data, the method further includes: Obtain the first voiceprint object corresponding to the first audio data, and at least one voiceprint object corresponding to the second audio data; If the first voiceprint object satisfies the confidence condition corresponding to at least one voiceprint object, the first audio data is determined to satisfy the second confidence condition.

3. The method according to claim 2, characterized in that, The step of determining that the first audio data satisfies the second confidence condition when the first voiceprint object satisfies the confidence condition corresponding to at least one voiceprint object includes: If a second voiceprint object is identified from the at least one voiceprint object whose similarity to the first voiceprint object is greater than a second preset threshold, the first audio data is determined to satisfy the second confidence condition.

4. The method according to any one of claims 1 to 3, characterized in that, Before acquiring the second audio data corresponding to at least one second device, the method further includes: If the amount of the first audio data is greater than or equal to the first preset threshold, it is determined that the first audio data satisfies the first confidence condition.

5. A voiceprint registration device, characterized in that, include: The first acquisition unit is used to acquire the first audio data collected by the first device; The second acquisition unit is configured to acquire at least one second audio data corresponding to a second device when the first audio data satisfies the first confidence condition corresponding to the first device, wherein the second device is a device that has performed the voiceprint registration through the second audio data; A registration unit is configured to register the voiceprint of the first device based on the first audio data when the first audio data satisfies the second confidence condition corresponding to the second audio data; the voiceprint registration of the first device based on the first audio data includes: registering the voiceprint of the first device using the first audio data and the second audio data. The first acquisition unit includes: a clustering module, configured to perform clustering processing on the acquired multiple audio data to obtain at least one audio data set, wherein each data set in the at least one audio data set records audio data of the same type, and the first audio data is the same type of audio data obtained by the first device after performing the clustering processing on the acquired multiple audio data; and a second acquisition module, configured to determine a target audio data set from the at least one audio data set and acquire the audio data recorded in the target audio data set, wherein the audio data recorded in the target audio data set includes the first audio data; The device further includes: a third acquisition module, configured to perform audio quality calculation on the new audio data to obtain a quality confidence level corresponding to the new audio data when the new audio data is recorded in the target audio data set; and a confidence module, configured to determine that the first audio data satisfies the first confidence condition when the quality confidence level is greater than or equal to a preset quality threshold and the at least one audio data set contains audio data belonging to the same type of audio data as the new audio data.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 4.

7. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 4 through the computer program.