Audio registration method and device, storage medium and electronic device
By classifying and clustering unregistered audio data in a smart home system to generate user tags, the problem of accurate audio registration for unregistered users is solved. This enables accurate identification of unregistered users and automatic construction of user profiles, thereby improving the user experience of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
- Filing Date
- 2023-03-31
- Publication Date
- 2026-04-17
AI Technical Summary
Existing smart home systems cannot accurately identify the number of family members and user needs when users are not registered, resulting in poor accuracy in audio registration.
By acquiring initial audio data collected by smart devices, filtering unregistered audio data, and classifying them into different categories based on audio similarity, unregistered user tags are generated, user profiles are registered, and unregistered voiceprint recognition is achieved.
It improves the accuracy of audio registration, enables automatic identification of unregistered users and accurate user profiling, and enhances the user experience of smart home systems.
Smart Images

Figure CN116524935B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home / intelligent home technology, and more specifically, to a method for audio registration, a storage medium, and an electronic device. Background Technology
[0002] With the development of smart homes, the current user profile creation in smart home systems typically relies solely on registered voiceprint recognition. This involves processing the audio data of users after they register to create a user profile. However, for users who haven't registered, the content of their smart home interactions is discarded, making it impossible to accurately describe the specific number of family members and user needs. Therefore, there is a technical problem of poor accuracy in audio registration. Summary of the Invention
[0003] This application provides an audio registration method, apparatus, storage medium, and electronic device to at least solve the technical problem of low audio registration efficiency.
[0004] According to one aspect of the embodiments of this application, an audio registration method is provided, including:
[0005] Acquire initial audio data collected by at least one smart device, and filter out unregistered audio data from the initial audio data;
[0006] Based on the audio similarity between the various audio data in the aforementioned unregistered audio data, the aforementioned unregistered audio data is divided into at least two categories of audio data, wherein each category of audio data in the aforementioned at least two categories of audio data corresponds to a different unregistered first user tag;
[0007] Based on the above at least two types of audio data, the first user tag is registered and a user profile corresponding to the first user tag is generated. When the above at least one smart device receives a control command triggered by the first user tag, it will perform a smart operation to match the user profile corresponding to the first user tag.
[0008] According to another aspect of the embodiments of this application, an audio registration device is also provided, comprising:
[0009] The first acquisition unit is used to acquire initial audio data collected by at least one smart device, and to filter out unregistered audio data from the initial audio data.
[0010] The first processing unit is configured to divide the unregistered audio data into at least two categories of audio data according to the audio similarity between each audio data in the unregistered audio data, wherein each category of audio data in the at least two categories of audio data corresponds to a different unregistered first user tag.
[0011] The first registration unit is used to register the first user tag based on the above-mentioned at least two types of audio data, and generate a user profile corresponding to the first user tag. When the above-mentioned at least one smart device receives a control command triggered by the first user tag, it will perform a smart operation to match the user profile corresponding to the first user tag.
[0012] As an optional solution, the first processing unit mentioned above includes:
[0013] The accumulation module is used to accumulate the aforementioned unregistered audio data;
[0014] The classification module is used to classify the unregistered audio data according to gender and age when the amount of unregistered audio data reaches a preset threshold, thereby obtaining audio sub-data under multiple different categories.
[0015] The input module is used to input the above audio sub-data into the clustering model to obtain at least two classes of audio data.
[0016] As an optional solution, the aforementioned first input module includes:
[0017] The first determination submodule is used to determine K initial cluster centers based on the feature labels corresponding to different categories;
[0018] The second determining submodule is used to determine the minimum Euclidean distance between each of the above audio sub-data and the above initial cluster center;
[0019] The third determining submodule is used to determine the average minimum Euclidean distance value of the above audio data, wherein the average minimum Euclidean distance value of the above audio data is obtained by averaging the minimum Euclidean distance values of a certain number of the above audio subdata.
[0020] The fourth determination submodule is used to determine that the clustering model training is complete when the average minimum Euclidean distance value is less than or equal to a preset threshold.
[0021] As an optional solution, the above-mentioned device further includes:
[0022] The first input submodule is used to determine that after the clustering model training is completed, in response to the first audio data acquisition instruction, the first audio data is input into the trained clustering model to obtain the first audio recognition result. The first audio data is new audio data collected by the device.
[0023] The fifth determination submodule is used to determine whether the above-mentioned first audio data has been completed for audio registration.
[0024] As an optional solution, the fifth determining sub-module mentioned above includes:
[0025] The molecular unit is used to score the sound quality of the first audio data to obtain the first score result;
[0026] The processing subunit is used to add a first identity identifier to the first audio data when the first score result is greater than or equal to a preset threshold.
[0027] As an optional solution, the fifth determining sub-module mentioned above includes:
[0028] Get sub-unit, used to retrieve a pre-saved list of identity identifiers;
[0029] The hit sub-unit is used to match the first identity identifier with the identity identifier list.
[0030] The update subunit is used to update the audio data corresponding to the hit second identity to the audio data corresponding to the first identity when the first identity matches the second identity in the identity list.
[0031] Add a sub-unit to add the first identity to the identity list if the first identity does not match the second identity in the identity list.
[0032] As an optional solution, the above-mentioned device further includes:
[0033] The recognition unit is used to, after acquiring initial audio data collected by at least one smart device, recognize the feature information of the target audio in response to an audio acquisition request, wherein the audio acquisition request is used to request the acquisition of the target audio;
[0034] The second processing unit is used to obtain initial audio data collected by at least one smart device, and then determine the first identity tag and the second identity tag of the target audio based on the feature information of the target audio. The first identity tag is used to indicate the behavioral information corresponding to the target audio, and the second identity tag is used to indicate the registration information corresponding to the target audio.
[0035] The third processing unit is used to acquire initial audio data collected by at least one smart device, and then compare the similarity between the first identity tag and the second identity tag to obtain the comparison result.
[0036] The second registration unit is used to update the registration information of the target audio after acquiring the initial audio data collected by at least one smart device, provided that the comparison result is greater than or equal to a preset threshold.
[0037] The third registration unit is used to obtain the initial audio data collected by at least one smart device, and if the comparison result is less than the preset threshold, obtain the behavioral information corresponding to the first identity tag, and use the behavioral information to complete the registration of the new user.
[0038] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program.
[0039] According to another aspect of the embodiments of this application, an electronic device is provided, including a memory and a processor, characterized in that the memory stores a computer program.
[0040] In this embodiment, initial audio data collected by at least one smart device is acquired, and unregistered audio data is filtered out from the initial audio data. Based on the audio similarity between the various audio data in the unregistered audio data, the unregistered audio data is divided into at least two categories of audio data, where each category of audio data corresponds to a different unregistered first user tag. Based on the at least two categories of audio data, the first user tag is registered, and a user profile corresponding to the first user tag is generated. When the at least one smart device receives a control command triggered by the first user tag, it executes an intelligent operation matching the user profile corresponding to the first user tag. By using the audio similarity between the unregistered audio data for clustering and registering user profiles based on different categories, the purpose of automatic unregistered voiceprint recognition is achieved, thereby improving the accuracy of audio registration and solving the technical problem of poor audio registration accuracy. Attached Figure Description
[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0042] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This is a schematic diagram of the hardware environment for an interaction method of a smart device according to an embodiment of this application;
[0044] Figure 2 This is a flowchart illustrating an optional audio registration method according to an embodiment of this application;
[0045] Figure 3 This is a schematic diagram illustrating an example of another optional audio registration method according to an embodiment of this application;
[0046] Figure 4 This is a schematic diagram illustrating an example of another optional audio registration method according to an embodiment of this application;
[0047] Figure 5 This is a schematic diagram illustrating an example of another optional audio registration method according to an embodiment of this application;
[0048] Figure 6 This is a schematic diagram of an optional audio registration method apparatus according to an embodiment of this application;
[0049] Figure 7 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application. Detailed Implementation
[0050] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0051] It should be noted that the terms "first," "second," etc., used in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0052] According to one aspect of the embodiments of this application, an interaction method for smart home devices is provided. This interaction method for smart home devices is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligencehouse ecosystems. Optionally, in this embodiment, the interaction method for smart home devices can be applied to, for example... Figure 1The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0053] The network may include, but is not limited to, at least one of the following: wired network, wireless network. A wired network may include, but is not limited to, at least one of the following: wide area network (WAN), metropolitan area network (MAN), local area network (LAN). A wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PCs, mobile phones, tablets, smart air conditioners, smart range hoods, smart refrigerators, smart ovens, smart stoves, smart washing machines, smart water heaters, smart washing equipment, smart dishwashers, smart projectors, smart TVs, smart clothes racks, smart curtains, smart audio-visual equipment, smart sockets, smart speakers, smart speakers, smart fresh air systems, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaners, smart window cleaning robots, smart mopping robots, smart air purifiers, smart steam ovens, smart microwave ovens, smart water heaters, smart air purifiers, smart water dispensers, smart door locks, etc.
[0054] Alternatively, as an alternative implementation method, such as Figure 2 As shown, the audio registration method includes:
[0055] S202, acquire initial audio data collected by at least one smart device, and filter out unregistered audio data from the initial audio data;
[0056] S204, according to the audio similarity between each audio data in the unregistered audio data, divide the unregistered audio data into at least two categories of audio data, wherein each category of audio data in the at least two categories of audio data corresponds to a different unregistered first user tag;
[0057] S206, based on the at least two types of audio data, register the first user tag and generate a user profile corresponding to the first user tag, wherein when the at least one smart device receives a control command triggered by the first user tag, it will perform a smart operation to match the user profile corresponding to the first user tag.
[0058] Optionally, in this embodiment, with the development of technology and the improvement of people's living standards, more and more families are installing smart home devices, such as air conditioners, televisions, refrigerators, washing machines, etc., which can be voice-interacted and voice-announced. Different family members have different ways of using smart devices and their frequency of use. Creating a user profile based on user preferences helps to optimize product design and improve user experience. However, the formation of a user profile requires various types of data from family members, including the distinction of each family member's voice. Only by determining which content was sent by which user can a user profile be formed. In smart homes, user profiles are usually formed using only registered voiceprints, because only registered audio can identify which speaker performed which function interaction. Knowing the user's specific needs allows for the classification of the user into different types. This embodiment uses clustered voiceprints to help the big data platform perform user-insensitive clustering of users with unregistered voiceprints, obtaining the speaker ID of the unregistered voiceprint. By combining clustered voiceprints with registered and unregistered voiceprints, a user profile is created, including: speaker ID, speaker's gender, age, and specific interaction content.
[0059] Optionally, in this embodiment, the initial audio data can be, but is not limited to, audio data collected when smart home devices interact with users via audio. Dividing the unregistered audio data into at least two categories based on the audio similarity between the various audio data points can be understood as, after accumulating a certain amount of unregistered audio data, performing feature recognition on the unregistered audio data and classifying it according to various categories, such as age, gender, etc., grouping similar data together. The specific classification does not concern itself with the category label; the goal is to aggregate similar data together to form clusters. Clustering is an unsupervised learning method, while classification separates different data. The process involves obtaining a classifier through a training dataset and then using the classifier to predict unknown data. Classification is a supervised learning method.
[0060] Optionally, in this embodiment, at least two types of audio data are obtained corresponding to different unregistered first user tags. Unregistered first user tags can be understood as, but are not limited to, first user tags corresponding to audio data that the user has not pre-registered. User profiles can be understood as target user models built on real data. They are an effective tool for outlining target users and connecting user needs. The establishment of user profiles can include, but is not limited to: collecting user behavior data, serializing and storing the data in the original database for aggregation and cleaning, then analyzing and judging the data based on classification and clustering, and finally visually constructing the user profile model to understand and grasp the user's personal preferences and behavioral habits. When a control command triggered by the user is received, the intelligent operation of matching the user profile corresponding to the user is executed.
[0061] To illustrate further, for example... Figure 3 As shown, the smart air conditioner collects multiple initial audio data 302, which are "Turn on the air conditioner" and "I want to turn on the air conditioner". The initial audio data 302 is clustered and classified according to age to obtain two types of audio data: children and adults. These two types of audio data correspond to different unregistered first user tags. Based on at least two types of audio data, the first user tag is registered to generate a user profile 304 corresponding to the first user tag. When a control command triggered by the first user tag is received, the smart operation matched by the user profile 304 corresponding to the first user tag will be executed.
[0062] In this embodiment, initial audio data collected by at least one smart device is acquired, and unregistered audio data is filtered out from the initial audio data. Based on the audio similarity between the various audio data in the unregistered audio data, the unregistered audio data is divided into at least two categories of audio data, where each category of audio data corresponds to a different unregistered first user tag. Based on the at least two categories of audio data, the first user tag is registered, and a user profile corresponding to the first user tag is generated. When the at least one smart device receives a control command triggered by the first user tag, it executes an intelligent operation matching the user profile corresponding to the first user tag. By using the audio similarity between the unregistered audio data for clustering and registering user profiles based on different categories, the purpose of automatic unregistered voiceprint recognition is achieved, thereby improving the accuracy of audio registration and solving the technical problem of poor audio registration accuracy.
[0063] As an optional approach, the unregistered audio data is divided into at least two categories based on the audio similarity between the various audio data points, including:
[0064] The unregistered audio data is accumulated;
[0065] When the amount of unregistered audio data reaches a preset threshold, the unregistered audio data is classified according to gender and age to obtain multiple audio sub-data under different categories;
[0066] The audio sub-data is input into a clustering model to obtain at least two classes of audio data.
[0067] Optionally, in this embodiment, it can be understood, but is not limited to, accumulating the number of unregistered audio data. When the accumulated unregistered audio data reaches a certain preset threshold, the unregistered audio data is classified according to gender and age to obtain multiple audio sub-data under different categories. The classification criteria are not limited to age and gender, and may include, but are not limited to, classification based on different languages, which can further improve the accuracy and flexibility of clustering model generation.
[0068] The embodiments provided in this application accumulate the unregistered audio data; when the unregistered audio data reaches a preset threshold, the unregistered audio data is classified according to gender and age to obtain multiple audio sub-data under different categories; the audio sub-data is input into a clustering model to obtain at least two types of audio data, thereby improving the accuracy of timbre selection.
[0069] As an optional approach, inputting the audio sub-data into a clustering model to obtain at least two classes of audio data includes:
[0070] Based on the feature labels corresponding to different categories, determine K initial cluster centers;
[0071] Determine the minimum Euclidean distance between each audio sub-data and the initial cluster center;
[0072] The average minimum Euclidean distance value of the audio data is determined, wherein the average minimum Euclidean distance value of the audio data is obtained by averaging the minimum Euclidean distance values of a certain number of audio sub-data.
[0073] If the average minimum Euclidean distance value is less than or equal to a preset threshold, the clustering model is considered to have completed training.
[0074] Through the embodiments provided in this application, K initial cluster centers are determined based on the feature labels corresponding to different categories; the minimum Euclidean distance between each audio sub-data and the initial cluster centers is determined; the average minimum Euclidean distance value of the audio data is determined, wherein the average minimum Euclidean distance value of the audio data is obtained by averaging the minimum Euclidean distance values of a certain number of audio sub-data; when the average minimum Euclidean distance value is less than or equal to a preset threshold, the clustering model training is determined to be complete, thus achieving the accuracy of clustering model generation.
[0075] As an optional approach, after the clustering model training is completed, the method further includes:
[0076] In response to a first audio data acquisition instruction, the first audio data is input into the trained clustering model to obtain a first audio recognition result, wherein the first audio data is new audio data collected by the device;
[0077] The first audio data is confirmed to be successfully registered.
[0078] Optionally, in this embodiment, it can be understood, but is not limited to, that after the clustering model is completed, when a new user audio is received, in response to the first audio data acquisition instruction, feature extraction is performed on a user audio, which is extracted as first audio data. The first audio data is then input into the trained clustering model to obtain the first audio recognition result. Based on the indication of the first audio recognition result, the user profile corresponding to the first audio data is determined to complete the audio registration.
[0079] Through the embodiments provided in this application, in response to a first audio data acquisition instruction, the first audio data is input into the trained clustering model to obtain a first audio recognition result. The first audio data is new audio data collected by the device. Determining that the first audio data has completed audio registration achieves the technical effect of improving the accuracy of timbre selection.
[0080] As an optional approach, determining that the first audio data has completed audio registration includes:
[0081] The first audio data is scored in terms of sound quality to obtain a first score result;
[0082] If the first score result is greater than or equal to a preset threshold, the first audio data is assigned a first identity identifier.
[0083] To illustrate further, for example... Figure 4As shown, audio evaluation and verification can be performed on the collected audio, including but not limited to scoring the sound quality of the first audio data and judging the sound quality by using a pre-given threshold. If the obtained first score is greater than or equal to the preset threshold, the audio quality is considered to be good. The confidence level corresponding to the first audio data can be obtained, and it can be judged whether the confidence level exceeds the specified threshold. If it exceeds the specified threshold, the audio quality is considered to be good and more in line with the speaker.
[0084] Optionally, in this embodiment, the first identity identifier can be understood as, but is not limited to, a speaker's ID information. After obtaining the first audio data and performing audio verification, if it is determined that the quality of the first audio data is good, the first audio data can be assigned ID information to facilitate subsequent comparison and registration-based voiceprint recognition.
[0085] Through the embodiments provided in this application, the first audio data is scored in terms of sound quality to obtain a first score result; when the first score result is greater than or equal to a preset threshold, the first audio data is assigned a first identity identifier, thereby improving the accuracy of audio recognition.
[0086] As an optional approach, determining that the first audio data has completed audio registration includes:
[0087] Retrieve a pre-saved list of identities;
[0088] Match the first identity identifier with the identity identifier list;
[0089] If the first identity identifier matches the second identity identifier in the identity identifier list, the audio data corresponding to the matched second identity identifier will be updated to the audio data corresponding to the first identity identifier;
[0090] If the first identity does not match the second identity in the identity list, the first identity is added to the identity list.
[0091] Optionally, in this embodiment, the identity list may include, but is not limited to, information such as the speaker's gender and age corresponding to the audio content. The identity list may include, but is not limited to, all registered audio data. By organizing the identity list, the complete structure of the family can be obtained, such as the specific number of family members, the male-to-female ratio, whether there are elderly people and children, etc. Through home environment settings, such as whether a children's mode or an elderly mode is needed, more humanized home services can be provided.
[0092] Optionally, in this embodiment, the first identity identifier is matched and matched against the identity identifier list. The second identity identifier can be understood as, but is not limited to, the registered audio. If the first identity identifier matches the registered audio data, it is determined that the audio data and the matched second audio data belong to the same speaker. Therefore, the audio data corresponding to the second identity identifier is updated to the audio data corresponding to the first identity identifier. If the first identity identifier does not match the second identity identifier in the identity identifier list, it is determined that the first audio data corresponding to the first identity identifier is unregistered audio data. Therefore, the first identity identifier is added to the identity identifier list.
[0093] The embodiments provided in this application are used to obtain a pre-saved list of identity identifiers;
[0094] Match the first identity identifier with the identity identifier list;
[0095] If the first identity identifier matches the second identity identifier in the identity identifier list, the audio data corresponding to the matched second identity identifier will be updated to the audio data corresponding to the first identity identifier;
[0096] If the first identity does not match the second identity in the identity list, the first identity is added to the identity list, which has the beneficial effect of improving the accuracy of audio data registration.
[0097] As an optional approach, the acquisition of initial audio data collected by at least one smart device further includes:
[0098] In response to an audio acquisition request, feature information of the target audio is identified, wherein the audio acquisition request is used to request the acquisition of the target audio;
[0099] Based on the feature information of the target audio, a first identity tag and a second identity tag of the target audio are determined, wherein the first identity tag is used to indicate the behavioral information corresponding to the target audio, and the second identity tag is used to indicate the registration information corresponding to the target audio;
[0100] The similarity between the first identity tag and the second identity tag is compared to obtain the comparison result;
[0101] If the comparison result is greater than or equal to a preset threshold, the registration information of the target audio is updated;
[0102] If the comparison result is less than the preset threshold, obtain the behavioral information corresponding to the first identity tag, and use the behavioral information to complete the registration of the new user.
[0103] Through the embodiments provided in this application, in response to an audio acquisition request, feature information of the target audio is identified, wherein the audio acquisition request is used to request the acquisition of the target audio;
[0104] Based on the feature information of the target audio, a first identity tag and a second identity tag of the target audio are determined, wherein the first identity tag is used to indicate the behavioral information corresponding to the target audio, and the second identity tag is used to indicate the registration information corresponding to the target audio;
[0105] The similarity between the first identity tag and the second identity tag is compared to obtain the comparison result;
[0106] If the comparison result is greater than or equal to a preset threshold, the registration information of the target audio is updated;
[0107] If the comparison result is less than the preset threshold, the behavioral information corresponding to the first identity tag is obtained, and the new user is registered using the behavioral information, thereby improving the accuracy of audio registration.
[0108] To facilitate understanding, the audio registration method will be applied to a specific audio registration scenario:
[0109] Optionally, in this embodiment, for example Figure 5As shown, a certain device (let's say a water heater) accumulates a expected amount of audio data. Based on the recognition results (gender result) of unregistered voiceprints, the data is classified. Each category of data is clustered into voiceprints, and the success of the clustering is determined. If clustering is successful, different clustering models are obtained, each representing a speaker. When new audio is injected after successful clustering, the unregistered voiceprint clustering model is used to verify which speaker the new audio belongs to. Through verification, a score is obtained for the audio. If the score exceeds a pre-defined threshold, the audio quality is considered good. The new audio is then used to determine if it belongs to a registered speaker, and a confidence score is obtained. If the confidence score exceeds a specified threshold, the audio quality is also considered good, closely matching the speaker. This satisfies both unregistered and registered voiceprint recognition. After verification, it is determined whether the speakerid of the unregistered voiceprint pointed to by the audio is the same speaker as the speaker of the registered voiceprint. If they are the same speakerid, the audio needs to be added to the list of updated audio. If the unregistered voice clustering model adds a new speakerid, it is used as a new unregistered speaker. By generating a new list of audio corresponding to each speakerid, which contains information such as the speaker's gender, age, and corresponding audio content, the user profile is completed. Through sorting, the overall structure of the family can also be obtained, such as the number of family members, the male-to-female ratio, and whether there are elderly people or children. In this way, by setting the home environment, whether a children's mode or an elderly mode is needed, more humanized home services can be provided. This embodiment performs clustering voiceprint recognition on the voice data of the home environment to help complete the user profile.
[0110] Optionally, in this embodiment, to determine whether registered and unregistered voiceprints belong to the same speaker, a comparison can be made by mapping existing audio lists to determine similarity. If the similarity is greater than a specified threshold, they are considered to be the same speaker.
[0111] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0112] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0113] According to another aspect of the embodiments of this application, an apparatus for implementing audio registration is also provided. For example... Figure 6 As shown, the device includes:
[0114] The first acquisition unit 602 is used to acquire initial audio data collected by at least one smart device, and to filter out unregistered audio data from the initial audio data;
[0115] The first processing unit 604 is configured to divide the unregistered audio data into at least two categories of audio data according to the audio similarity between each audio data in the unregistered audio data, wherein each category of audio data in the at least two categories of audio data corresponds to a different unregistered first user tag;
[0116] The first registration unit 606 is used to register the first user tag based on the at least two types of audio data and generate a user profile corresponding to the first user tag. When the at least one smart device receives a control command triggered by the first user tag, it will perform a smart operation to match the user profile corresponding to the first user tag.
[0117] For specific implementation examples, please refer to the example shown in the audio registration method; these examples will not be repeated here.
[0118] As an optional solution, the first processing unit includes:
[0119] An accumulation module is used to accumulate the unregistered audio data;
[0120] The classification module is used to classify the unregistered audio data according to gender and age when the amount of unregistered audio data reaches a preset threshold, thereby obtaining audio sub-data under multiple different categories.
[0121] The input module is used to input the audio sub-data into the clustering model to obtain at least two types of audio data.
[0122] For specific implementation examples, please refer to the example shown in the audio registration method; these examples will not be repeated here.
[0123] As an optional solution, the first input module includes:
[0124] The first determination submodule is used to determine K initial cluster centers based on the feature labels corresponding to different categories;
[0125] The second determining submodule is used to determine the minimum Euclidean distance between each audio subdata and the initial cluster center;
[0126] The third determining submodule is used to determine the average minimum Euclidean distance value of the audio data, wherein the average minimum Euclidean distance value of the audio data is obtained by averaging the minimum Euclidean distance values of a certain number of audio subdata.
[0127] The fourth determination submodule is used to determine that the clustering model training is complete when the average minimum Euclidean distance value is less than or equal to a preset threshold.
[0128] For specific implementation examples, please refer to the example shown in the audio registration method; these examples will not be repeated here.
[0129] As an optional solution, the device further includes:
[0130] The first input submodule is used to determine that after the clustering model training is completed, in response to the first audio data acquisition instruction, input the first audio data into the trained clustering model to obtain the first audio recognition result, wherein the first audio data is new audio data collected by the device;
[0131] The fifth determining submodule is used to determine whether the first audio data has completed audio registration.
[0132] For specific implementation examples, please refer to the example shown in the audio registration method; these examples will not be repeated here.
[0133] As an optional solution, the fifth determining submodule includes:
[0134] A molecular unit is used to score the sound quality of the first audio data to obtain a first score result;
[0135] The processing subunit is used to add a first identity identifier to the first audio data when the first score result is greater than or equal to a preset threshold.
[0136] For specific implementation examples, please refer to the example shown in the audio registration method; these examples will not be repeated here.
[0137] As an optional solution, the fifth determining submodule includes:
[0138] Get sub-unit, used to retrieve a pre-saved list of identity identifiers;
[0139] The hit sub-unit is used to match the first identity identifier with the identity identifier list;
[0140] The update subunit is used to update the audio data corresponding to the hit second identity to the audio data corresponding to the first identity when the first identity matches the second identity in the identity list;
[0141] An add subunit is configured to add the first identity to the identity list if the first identity does not match the second identity in the identity list.
[0142] For specific implementation examples, please refer to the example shown in the audio registration method; these examples will not be repeated here.
[0143] As an optional solution, the device further includes:
[0144] The recognition unit is configured to, after acquiring initial audio data collected by at least one smart device, recognize the feature information of the target audio in response to an audio acquisition request, wherein the audio acquisition request is used to request the acquisition of the target audio;
[0145] The second unit processing unit is used to obtain initial audio data collected by at least one smart device, and then determine a first identity tag and a second identity tag of the target audio based on the feature information of the target audio, wherein the first identity tag is used to indicate the behavioral information corresponding to the target audio, and the second identity tag is used to indicate the registration information corresponding to the target audio.
[0146] The third processing unit is used to acquire initial audio data collected by at least one smart device, and then compare the similarity between the first identity tag and the second identity tag to obtain the comparison result.
[0147] The second registration unit is used to update the registration information of the target audio after acquiring initial audio data collected by at least one smart device, provided that the comparison result is greater than or equal to a preset threshold.
[0148] The third registration unit is used to obtain the initial audio data collected by at least one smart device, and if the comparison result is less than the preset threshold, obtain the behavioral information corresponding to the first identity tag, and use the behavioral information to complete the registration of the new user.
[0149] For specific implementation examples, please refer to the example shown in the audio registration method; these examples will not be repeated here.
[0150] According to another aspect of the embodiments of this application, an electronic device for implementing an audio registration method is also provided, such as... Figure 7 As shown, the electronic device includes a memory 702 and a processor 704. The memory 702 stores a computer program, and the processor 704 is configured to execute the steps of any of the method embodiments via the computer program.
[0151] Optionally, in this embodiment, the electronic device may be located in at least one of a plurality of network devices in a computer network.
[0152] Optionally, in this embodiment, the processor may be configured to perform the following steps via a computer program:
[0153] S1, acquire initial audio data collected by at least one smart device, and filter out unregistered audio data from the initial audio data;
[0154] S2, according to the audio similarity between each audio data in the unregistered audio data, the unregistered audio data is divided into at least two categories of audio data, wherein each category of audio data in the at least two categories of audio data corresponds to a different unregistered first user tag;
[0155] S3, based on the at least two types of audio data, register the first user tag and generate a user profile corresponding to the first user tag, wherein when the at least one smart device receives a control command triggered by the first user tag, it will execute a smart operation to match the user profile corresponding to the first user tag.
[0156] Alternatively, as those skilled in the art will understand, Figure 7 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 7 It does not impose limitations on the structure of electronic devices. For example, electronic devices may also include components that are more... Figure 7 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 7 The different configurations shown.
[0157] The memory 702 can be used to store software programs and modules, such as the program instructions / modules corresponding to the audio registration method and apparatus in this embodiment. The processor 704 executes various functional applications and data processing by running the software programs and modules stored in the memory 702, thereby implementing the audio registration method. The memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 702 may further include memory remotely located relative to the processor 704, and these remote memories can be connected to the terminal via a network. Examples of networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 702 may be used, but is not limited to, to store information such as initial audio data. As an example, such as... Figure 7 As shown, the memory 702 may include, but is not limited to, the first acquisition unit 602, the first processing unit 604, and the first registration unit 606 of the audio registration device. Furthermore, it may include, but is not limited to, other module units in the audio registration device of the virtual model, which will not be elaborated upon in this example.
[0158] Optionally, the specific examples of the network through which the transmission device 706 receives or sends data may include wired and wireless networks. In one example, the transmission device 706 includes a Network Interface Controller (NIC), which can be connected to other network devices and routers via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 706 is a Radio Frequency (RF) module used for wireless communication with the Internet.
[0159] In addition, the electronic device also includes: a display 708 for displaying information such as tone indication information; and a connection bus 710 for connecting various module components in the electronic device.
[0160] In other embodiments, the terminal device or server can be a node in a distributed system, which can be a blockchain system. This blockchain system is formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any type of computing device, such as a server or terminal, can become a node in the blockchain system by joining this P2P network.
[0161] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.
[0162] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0163] It should be noted that the computer system of the electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0164] A computer system includes a Central Processing Unit (CPU), which performs various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) or loaded from RAM. ROM also stores various programs and data required for system operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output interfaces (I / O interfaces) are also connected to the bus.
[0165] The following components are connected to the input / output interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard drives; and communication sections including network interface cards such as LAN cards and modems. The communication section performs communication processing via a network such as the Internet. Drives are also connected to the input / output interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required.
[0166] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions defined in the system of this application.
[0167] According to one aspect of this application, a computer-readable storage medium is provided, from which a processor of a computer device reads computer instructions, and the processor executes the computer instructions, causing the computer device to perform methods provided in various alternative implementations.
[0168] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0169] S1, acquire initial audio data collected by at least one smart device, and filter out unregistered audio data from the initial audio data;
[0170] S2, according to the audio similarity between each audio data in the unregistered audio data, the unregistered audio data is divided into at least two categories of audio data, wherein each category of audio data in the at least two categories of audio data corresponds to a different unregistered first user tag;
[0171] S3, based on the at least two types of audio data, register the first user tag and generate a user profile corresponding to the first user tag, wherein when the at least one smart device receives a control command triggered by the first user tag, it will execute a smart operation to match the user profile corresponding to the first user tag.
[0172] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the embodiment can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0173] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0174] If the integrated units in the embodiments are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods of the various embodiments of this application.
[0175] In the embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0176] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or the indirect coupling or communication connection of units or modules may be electrical or other forms.
[0177] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0178] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0179] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
[0180] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An audio registration method, characterized in that, include: Acquire initial audio data collected by at least one smart device, and filter out unregistered audio data from the initial audio data; Based on the audio similarity between the various audio data in the unregistered audio data, the unregistered audio data is divided into at least two categories of audio data, wherein each category of audio data in the at least two categories of audio data corresponds to a different unregistered first user tag; Based on the at least two types of audio data, the first user tag is registered and a user profile corresponding to the first user tag is generated. When the at least one smart device receives a control command triggered by the first user tag, it will perform a smart operation to match the user profile corresponding to the first user tag. In response to an audio acquisition request, feature information of the target audio is identified, wherein the audio acquisition request is used to request the acquisition of the target audio; Based on the feature information of the target audio, a first identity tag and a second identity tag of the target audio are determined, wherein the first identity tag is used to indicate the behavioral information corresponding to the target audio, and the second identity tag is used to indicate the registration information corresponding to the target audio; The similarity between the first identity tag and the second identity tag is compared to obtain the comparison result; If the comparison result is greater than or equal to a preset threshold, the registration information of the target audio is updated; if the comparison result is less than the preset threshold, the behavioral information corresponding to the first identity tag is obtained, and the behavioral information is used to complete the registration of the new user.
2. The method according to claim 1, characterized in that, The unregistered audio data is divided into at least two categories based on the audio similarity between the individual audio data points, including: The unregistered audio data is accumulated; When the amount of unregistered audio data reaches a preset threshold, the unregistered audio data is classified according to gender and age to obtain multiple audio sub-data under different categories; The audio sub-data is input into a clustering model to obtain at least two classes of audio data.
3. The method according to claim 2, characterized in that, The step of inputting the audio sub-data into the clustering model to obtain at least two types of audio data includes: Based on the feature labels corresponding to different categories, determine K initial cluster centers; Determine the minimum Euclidean distance between each audio sub-data and the initial cluster center; The average minimum Euclidean distance value of the audio data is determined, wherein the average minimum Euclidean distance value of the audio data is obtained by averaging the minimum Euclidean distance values of a certain number of audio sub-data. If the average minimum Euclidean distance value is less than or equal to a preset threshold, the clustering model is considered to have completed training.
4. The method according to claim 3, characterized in that, After the clustering model training is completed, the method further includes: In response to a first audio data acquisition instruction, the first audio data is input into the trained clustering model to obtain a first audio recognition result, wherein the first audio data is new audio data collected by the device; The first audio data has been confirmed to be successfully registered.
5. The method according to claim 4, characterized in that, The step of determining that the first audio data has completed audio registration includes: The first audio data is scored in terms of sound quality to obtain a first score result; If the first score result is greater than or equal to a preset threshold, the first audio data is assigned a first identity identifier.
6. The method according to claim 5, characterized in that, The step of determining that the first audio data has completed audio registration includes: Retrieve a pre-saved list of identities; Match the first identity identifier with the identity identifier list; If the first identity identifier matches the second identity identifier in the identity identifier list, the audio data corresponding to the matched second identity identifier will be updated to the audio data corresponding to the first identity identifier; If the first identity does not match the second identity in the identity list, the first identity is added to the identity list.
7. An audio registration device, characterized in that, include: The first acquisition unit is used to acquire initial audio data collected by at least one smart device, and to filter out unregistered audio data from the initial audio data; The first processing unit is configured to divide the unregistered audio data into at least two categories of audio data according to the audio similarity between each audio data in the unregistered audio data, wherein each category of audio data in the at least two categories of audio data corresponds to a different unregistered first user tag; The first registration unit is used to register the first user tag based on the at least two types of audio data and generate a user profile corresponding to the first user tag. When the at least one smart device receives a control command triggered by the first user tag, it will perform a smart operation to match the user profile corresponding to the first user tag. The device is further configured to, in response to an audio acquisition request, identify feature information of a target audio, wherein the audio acquisition request is used to request acquisition of the target audio; determine a first identity tag and a second identity tag of the target audio based on the feature information of the target audio, wherein the first identity tag is used to indicate behavioral information corresponding to the target audio, and the second identity tag is used to indicate registration information corresponding to the target audio; compare the similarity of the first identity tag and the second identity tag to obtain a comparison result; update the registration information of the target audio if the comparison result is greater than or equal to a preset threshold; and acquire the behavioral information corresponding to the first identity tag and use the behavioral information to complete the registration of a new user if the comparison result is less than the preset threshold.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 6.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 6 through the computer program.
Citation Information
Patent Citations
Automatic registration voiceprint recognition method and device
CN113241080A
Identity determination method and device of target object, storage medium and electronic device
CN114121020A
Household appliance and voiceprint control method thereof, server and readable storage medium
CN115482813A