Speech user recognition method and device, electronic equipment and storage medium

CN115662443BActive Publication Date: 2026-09-15IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211183402.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2026-09-15
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

[0003]但是,当用户的身体情况出现变化导致声纹特征改变时,或者当环境干扰较严重时,声纹识别会出现声纹偏移现象,无法准确识别用户身份

Benefits of technology

[0017] The voice user recognition method proposed in this application determines a first voiceprint identifier corresponding to the user's voice by extracting the voiceprint features of the user's voice; the first voiceprint identifier is compared with voiceprint identifiers in a pre-set voiceprint identifier lookup table to determine a first primary voiceprint identifier corresponding to the first voiceprint identifier; the voiceprint identifier lookup table contains primary and secondary voiceprint identifiers for each user, wherein each user's primary voiceprint identifier is a voiceprint identifier whose frequency is recognized within a preset time period is greater than a set frequency threshold, and the user's secondary voiceprint identifiers are voiceprint identifiers whose voiceprint similarity to the user's primary voiceprint identifier is greater than a set similarity threshold. Using the technical solution of this application, the voiceprint identifier lookup table can be used to associate all secondary voiceprint identifiers of the same user with the primary voiceprint identifier. When the voiceprint corresponding to the user's voice shifts, the primary voiceprint identifier of the user can be accurately retrieved from the voiceprint identifier lookup table using the secondary voiceprint identifier after the voiceprint shift, thus improving the accuracy of user information determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115662443B_ABST
    Figure CN115662443B_ABST
Patent Text Reader

Abstract

The application provides a voice user identification method and device, electronic equipment and a storage medium. The method comprises the following steps: determining a first voiceprint identifier corresponding to a user voice by extracting a voiceprint feature of the user voice; comparing the first voiceprint identifier with a voiceprint identifier in a voiceprint identifier comparison table, and determining a first main voiceprint identifier corresponding to the first voiceprint identifier; and the voiceprint identifier comparison table comprises main voiceprint identifiers and secondary voiceprint identifiers of each user. According to the technical scheme, all secondary voiceprint identifiers of the same user are associated with the main voiceprint identifier by using the voiceprint identifier comparison table. When the voiceprint of the user voice is offset, the main voiceprint identifier of the user can be accurately queried from the voiceprint identifier comparison table through the secondary voiceprint identifier after the voiceprint offset, and the accuracy of user information determination is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voiceprint recognition technology, and in particular to a voice user recognition method, device, electronic device and storage medium. Background Technology

[0002] With the continuous development of artificial intelligence technology, voice interaction has gradually become widespread in people's lives. For example, operating a TV with a voice remote has become a common practice, making operation more efficient and convenient. A voice analysis engine can recognize the voiceprint of a user based on their input, thereby analyzing user information such as interests and preferences. This allows the set-top box to provide personalized content recommendations based on the user's interests.

[0003] However, when a user's physical condition changes, causing changes in voiceprint characteristics, or when there is significant environmental interference, voiceprint recognition may exhibit voiceprint shift, making it impossible to accurately identify the user's identity. Summary of the Invention

[0004] Based on the deficiencies and shortcomings of the prior art, this application proposes a voice user recognition method, device, electronic device and storage medium, which can improve the accuracy of the acquired voice user information.

[0005] The first aspect of this application provides a voice user recognition method, including:

[0006] The first voiceprint identifier corresponding to the user's voice is determined by extracting the voiceprint features of the user's voice.

[0007] The first voiceprint identifier is compared with the voiceprint identifier in the pre-set voiceprint identifier lookup table to determine the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0008] The voiceprint identifier lookup table includes a primary voiceprint identifier and a secondary voiceprint identifier for each user. The primary voiceprint identifier for each user is a voiceprint identifier whose frequency is greater than a set frequency threshold within a preset time period. The secondary voiceprint identifier for the user is a voiceprint identifier whose voiceprint similarity to the user's primary voiceprint identifier is greater than a set similarity threshold.

[0009] A second aspect of this application provides a voice user recognition device, comprising:

[0010] The voiceprint identifier determination module is used to determine a first voiceprint identifier corresponding to the user's voice by extracting the voiceprint features of the user's voice;

[0011] The voiceprint identifier comparison module is used to compare the first voiceprint identifier with the voiceprint identifiers in a pre-set voiceprint identifier lookup table to determine the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0012] The voiceprint identifier lookup table includes a primary voiceprint identifier and a secondary voiceprint identifier for each user. The primary voiceprint identifier for each user is a voiceprint identifier whose frequency is greater than a set frequency threshold within a preset time period. The secondary voiceprint identifier for the user is a voiceprint identifier whose voiceprint similarity to the user's primary voiceprint identifier is greater than a set similarity threshold.

[0013] A third aspect of this application provides an electronic device, including: a memory and a processor;

[0014] The memory is connected to the processor and is used to store programs;

[0015] The processor is used to implement the above-described voice user recognition method by running the program in the memory.

[0016] A fourth aspect of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the above-described voice user recognition method.

[0017] The voice user recognition method proposed in this application determines a first voiceprint identifier corresponding to the user's voice by extracting the voiceprint features of the user's voice; the first voiceprint identifier is compared with voiceprint identifiers in a pre-set voiceprint identifier lookup table to determine a first primary voiceprint identifier corresponding to the first voiceprint identifier; the voiceprint identifier lookup table contains primary and secondary voiceprint identifiers for each user, wherein each user's primary voiceprint identifier is a voiceprint identifier whose frequency is recognized within a preset time period is greater than a set frequency threshold, and the user's secondary voiceprint identifiers are voiceprint identifiers whose voiceprint similarity to the user's primary voiceprint identifier is greater than a set similarity threshold. Using the technical solution of this application, the voiceprint identifier lookup table can be used to associate all secondary voiceprint identifiers of the same user with the primary voiceprint identifier. When the voiceprint corresponding to the user's voice shifts, the primary voiceprint identifier of the user can be accurately retrieved from the voiceprint identifier lookup table using the secondary voiceprint identifier after the voiceprint shift, thus improving the accuracy of user information determination. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating a voice user recognition method provided in an embodiment of this application;

[0020] Figure 2 This is a schematic diagram of the processing flow for determining the voiceprint identifier lookup table provided in the embodiments of this application;

[0021] Figure 3 This is a schematic diagram of the processing flow for determining the secondary voiceprint identifier corresponding to the primary voiceprint identifier, provided in an embodiment of this application.

[0022] Figure 4 This is a schematic diagram of the processing flow for updating preference data corresponding to voiceprint identifiers provided in an embodiment of this application;

[0023] Figure 5 This is a schematic diagram of the structure of a voice user recognition device provided in an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] The technical solution of this application is applicable to voiceprint recognition scenarios. Using the technical solution of this application can improve the accuracy of user information determination.

[0026] With the continuous development of artificial intelligence technology, voice interaction has gradually become widespread in people's lives, becoming one of our mainstream interaction methods. For example, people operate televisions using voice remote controls, making operation more efficient and convenient. Voice interaction mainly involves receiving user input, analyzing and recognizing the voice, and then enabling the terminal device to provide feedback based on the analysis and recognition results. For example, in voice interaction between a user and a set-top box, the user can input "open the TV series list" using the remote control. The set-top box can then respond to the user's voice by outputting and displaying the TV series list on the connected television. Since different users have different interests, personalized content can be provided based on user preferences to meet individual needs.

[0027] In existing technologies, user interest preferences and other information are stored based on the user's voiceprint identifier. Typically, a voice analysis engine is used to analyze the user's input voiceprint. Based on the corresponding voiceprint identifier, the interest preferences and other information are extracted from pre-stored information so that the terminal device (such as a set-top box) can provide personalized feedback based on this information. However, when a user's physical condition changes (e.g., during puberty or vocal cord damage), causing changes in voiceprint characteristics, or when environmental interference is severe, the voiceprint analysis engine may experience voiceprint shift, resulting in a different voiceprint identifier than the one corresponding to the user's interest preferences and other information. This makes it difficult to accurately identify the user, affecting the accuracy of voice user information and ultimately preventing the acquisition of accurate interest preferences and other information.

[0028] In view of the shortcomings of the existing technology and the low accuracy of recognizing voice user information in reality, the inventors of this application have proposed a voice user recognition method after research and experimentation. This method can improve the accuracy of user information determination.

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] This application provides a method for voice user recognition. See [link to relevant documentation] Figure 1 As shown, the method includes:

[0031] S101. Determine the first voiceprint identifier corresponding to the user's voice by extracting the voiceprint features of the user's voice.

[0032] Specifically, after receiving user voice, this embodiment can use a voiceprint analysis engine to perform preprocessing such as noise removal on the user voice, then extract the voiceprint features of the user voice from the preprocessed user voice, and finally perform real-time clustering model prediction based on the voiceprint features to determine the first voiceprint identifier corresponding to the user voice.

[0033] Clustering models utilize deep learning to pre-cluster speaker voiceprint features, grouping the voiceprint features of the same person into a single cluster and assigning each cluster a unique voiceprint identifier. In practical applications, when the same speaker inputs speech and the voiceprint features are extracted, the clustering model can determine the speaker's voiceprint identifier based on these features.

[0034] S102. Compare the first voiceprint identifier with the voiceprint identifier in the pre-set voiceprint identifier lookup table to determine the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0035] Specifically, in this embodiment, a voiceprint identifier lookup table is pre-set. This table includes the primary and secondary voiceprint identifiers of each user. For the same user, all of the user's secondary voiceprint identifiers are associated with the user's primary voiceprint identifier. The user's primary voiceprint identifier is a voiceprint identifier whose frequency is greater than a set frequency threshold within a preset time period; the user's secondary voiceprint identifiers are voiceprint identifiers whose voiceprint similarity to the user's primary voiceprint identifier is greater than a set similarity threshold.

[0036] In this embodiment, after determining the first voiceprint identifier corresponding to the user's voice, the first voiceprint identifier is compared with the voiceprint identifier in the voiceprint identifier lookup table. The voiceprint identifier that is the same as the first voiceprint identifier is found from the voiceprint identifier lookup table, and it is determined whether the found voiceprint identifier is the primary voiceprint identifier or the secondary voiceprint identifier, thereby determining the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0037] In this embodiment, the pre-set voiceprint identifier lookup table can be either a primary voiceprint-secondary voiceprint identifier lookup table or a secondary voiceprint-primary voiceprint identifier lookup table. Using the primary voiceprint-secondary voiceprint identifier lookup table, one can look up all secondary voiceprints corresponding to a primary voiceprint identifier, and vice versa. The voiceprint identifier lookup table not only records a user's primary voiceprint identifier but also captures secondary voiceprint identifiers when the user's voiceprint shifts, and associates these secondary voiceprint identifiers with the primary voiceprint identifier. Thus, when a user's voiceprint shifts, the primary voiceprint identifier corresponding to the secondary voiceprint identifier at the time of the shift can be retrieved from the voiceprint identifier lookup table. Even if the user's voice shifts, the user's identity can be accurately determined, thereby accurately identifying the user's information.

[0038] The specific execution method for this step can be described as follows:

[0039] The first voiceprint identifier is compared with the primary voiceprint identifier in the pre-set voiceprint identifier lookup table (primary voiceprint-secondary voiceprint identifier lookup table). If a primary voiceprint identifier in the lookup table (primary voiceprint-secondary voiceprint identifier lookup table) is identical to the first voiceprint identifier, then the first voiceprint identifier is the primary voiceprint identifier. In this case, the primary voiceprint identifier in the lookup table (primary voiceprint-secondary voiceprint identifier lookup table) that is identical to the first voiceprint identifier (also the first voiceprint identifier) ​​is taken as the first primary voiceprint identifier corresponding to the first voiceprint identifier. If no primary voiceprint identifier in the lookup table (primary voiceprint-secondary voiceprint identifier lookup table) is identical to the first voiceprint identifier, then the primary voiceprint identifier corresponding to the secondary voiceprint identifier that is identical to the first voiceprint identifier is searched from the secondary voiceprint identifier lookup table (secondary voiceprint-primary voiceprint identifier lookup table), and this primary voiceprint identifier is taken as the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0040] The specific execution method for this step can also be described as follows:

[0041] The first voiceprint identifier is compared with the secondary voiceprint identifiers in a pre-set voiceprint identifier lookup table (secondary voiceprint-primary voiceprint identifier lookup table). If a secondary voiceprint identifier identical to the first voiceprint identifier exists in the secondary voiceprint identifier lookup table, then the first voiceprint identifier is the secondary voiceprint identifier. In this case, the primary voiceprint identifier corresponding to the secondary voiceprint identifier identical to the first voiceprint identifier is searched in the voiceprint identifier lookup table (secondary voiceprint-primary voiceprint identifier lookup table) and used as the first primary voiceprint identifier corresponding to the first voiceprint identifier. If no secondary voiceprint identifier identical to the first voiceprint identifier exists in the secondary voiceprint identifier lookup table (primary voiceprint-secondary voiceprint identifier lookup table), then the primary voiceprint identifier identical to the first voiceprint identifier is searched in the primary voiceprint identifier lookup table (primary voiceprint-secondary voiceprint identifier lookup table) and used as the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0042] As described above, the voice user recognition method proposed in this application extracts the voiceprint features of the user's voice to determine a first voiceprint identifier corresponding to the user's voice; the first voiceprint identifier is compared with voiceprint identifiers in a pre-set voiceprint identifier lookup table to determine a first primary voiceprint identifier corresponding to the first voiceprint identifier; the voiceprint identifier lookup table contains the primary voiceprint identifier and secondary voiceprint identifiers of each user, wherein the primary voiceprint identifier of each user is the voiceprint identifier whose frequency is greater than a set frequency threshold within a preset time period, and the secondary voiceprint identifier of the user is the voiceprint identifier whose voiceprint similarity to the user's primary voiceprint identifier is greater than a set similarity threshold. Using the technical solution of this application, the voiceprint identifier lookup table can be used to associate all secondary voiceprint identifiers of the same user with the primary voiceprint identifier. When the voiceprint corresponding to the user's voice shifts, the primary voiceprint identifier of the user can be accurately retrieved from the voiceprint identifier lookup table using the secondary voiceprint identifier after the voiceprint shift, thus improving the accuracy of user information determination.

[0043] Furthermore, each time a user inputs voice, this embodiment needs to use a voice analysis engine to analyze the voice interaction information carrying the first voiceprint identifier corresponding to the user's voice. After determining the first primary voiceprint identifier corresponding to the first voiceprint identifier using the above embodiment, the voice interaction information corresponding to the user's voice needs to be merged and stored in the voice interaction information corresponding to the first primary voiceprint identifier, so that all voice interaction information of the same user can be stored and queried based on the user's primary voiceprint identifier.

[0044] As an optional implementation, another embodiment of this application discloses that, after step S102, the following steps can also be performed:

[0045] Based on a pre-set voiceprint identifier and preference data mapping table, the preference data corresponding to the first primary voiceprint identifier is obtained and used as the preference data of the user corresponding to the first primary voiceprint identifier.

[0046] Specifically, in this embodiment, speech recognition is performed on each received user voice message. The speech interaction information corresponding to all recognized user voice messages is stored according to the voiceprint identifier of each user voice message; that is, speech interaction information belonging to the same voiceprint identifier is merged and stored. This speech interaction information includes user operation data, speech text, and text intent. This embodiment can perform behavioral analysis based on relevant information in all speech interaction information corresponding to a voiceprint identifier, and then determine the preference data corresponding to that voiceprint identifier based on different weights for different behaviors.

[0047] This embodiment requires merging and storing the voice interaction information corresponding to all user voices of the secondary voiceprint identifiers corresponding to the primary voiceprint identifier into the voice interaction information corresponding to the primary voiceprint identifier, based on a voiceprint identifier lookup table, thereby achieving the fusion of voice interaction behaviors of the same user. Similarly, it merges and stores the preference data corresponding to the secondary voiceprint identifiers corresponding to the primary voiceprint identifier into the preference data corresponding to the primary voiceprint identifier, thereby achieving the fusion of preference data of the same user. A voiceprint identifier-preference data lookup table can be pre-set based on each voiceprint identifier and its corresponding preference data.

[0048] When, based on the voiceprint identifier lookup table, the first voiceprint identifier corresponding to the user's voice is found to be either the primary or secondary voiceprint identifier in the table, the corresponding preference data is retrieved from a pre-set voiceprint identifier and preference data lookup table according to the primary voiceprint identifier. The corresponding terminal device can then provide personalized feedback based on the user's voice interaction information and the retrieved preference data, aligning with the user's preferences.

[0049] As an optional implementation, see [link to implementation details]. Figure 2 As shown in another embodiment of this application, the processing steps for determining the voiceprint identifier lookup table are as follows:

[0050] S201. Obtain the voiceprint identifiers identified within a preset time period.

[0051] Specifically, for user-inputted voice, a voice analysis engine can be used to recognize all voice inputs, obtaining voice interaction information carrying voiceprint identifiers for each voice input. This voice interaction information includes the voice input time and the device identifier of the terminal device to which the voice input needs to be sent. In this embodiment, to construct a voiceprint identifier lookup table for a specific terminal device, it is first necessary to filter out voice inputs from all voice inputs whose device identifiers in the voice interaction information match the terminal device. Then, based on the voice input time of these voice inputs, all voice inputs within a preset time period are collected, and the corresponding voiceprint identifiers are obtained.

[0052] Furthermore, in this embodiment, a voiceprint identifier lookup table is only needed for terminal devices with low user mobility. For terminal devices with high user mobility, the user variation is too great; most users may only perform a single voice interaction with that terminal device, and the same user will not generate multiple voiceprint identifiers. Therefore, constructing a voiceprint identifier lookup table is not very useful. For example, for set-top boxes in home scenarios, user mobility is very low, and family members will repeatedly perform voice interactions with the set-top box. Therefore, it is necessary to construct a corresponding voiceprint identifier lookup table.

[0053] Therefore, this embodiment also needs to determine whether the terminal device is one that requires the construction of a voiceprint identifier lookup table. By judging whether the number of voiceprint identifiers obtained within a preset time period exceeds a preset number after deduplication, if the number of deduplicated voiceprint identifiers exceeds the preset number, it indicates that the user mobility of the terminal device is high, and no voiceprint identifier lookup table will be constructed for the terminal device. If the number of deduplicated voiceprint identifiers does not exceed the preset number, it indicates that the user mobility of the terminal device is low, and a voiceprint identifier lookup table needs to be constructed for the terminal device.

[0054] S202. Based on the frequency of voiceprint identification within a preset time period, determine the primary voiceprint identifier and secondary voiceprint identifier among all voiceprint identifiers.

[0055] Specifically, in this embodiment, for all voiceprint identifiers obtained in the above steps within a preset time period, the system determines whether each voiceprint identifier is a primary or secondary voiceprint identifier based on the frequency at which each voiceprint identifier is identified within the preset time period. The specific steps are as follows:

[0056] First, select the voiceprint identifiers that have been identified for more than the preset number of days within the preset time period and have been identified more than the preset number of times per day on average, and designate them as the main voiceprint identifiers.

[0057] Based on the voice interaction information corresponding to all voiceprint identifiers, the number of days each voiceprint identifier was identified within a preset time period, the total number of times each voiceprint identifier was identified within the preset time period, and the number of days the terminal device operated within the preset time period are determined. The preset number of days is obtained by multiplying the number of days the terminal device operated within the preset time period by a pre-set ratio coefficient. The average number of times a voiceprint identifier was identified per day is obtained by dividing the total number of times the voiceprint identifier was identified within the preset time period by the number of days it was identified within the preset time period. In this embodiment, voiceprint identifiers that were identified for more than the preset number of days within the preset time period and whose average number of identifications per day was greater than the preset number of identifications are selected as the primary voiceprint identifiers. The preset time period can be set to 30 days, the pre-set ratio coefficient can be 1 / 3, and the preset number of identifications is preferably set to 1. This embodiment can adjust the preset time period, ratio coefficient, and preset number of identifications according to actual needs.

[0058] For example, when the terminal device is a set-top box, the preset time period is 30 days, the preset ratio coefficient is 1 / 3, and the preset number of times is 1, the following statistics are calculated: the number of days each voiceprint identifier appearing under the set-top box was identified in the past 30 days (i.e., the number of active days for each voiceprint identifier in the past 30 days), the total number of times each voiceprint identifier was identified in the past 30 days (i.e., the number of active times for each voiceprint identifier in the past 30 days), and the number of days the set-top box was operational in the past 30 days (i.e., the number of active days for the set-top box in the past 30 days). One-third of the number of days the set-top box was operational in the past 30 days is taken as the preset number of days. The value obtained by dividing the number of active days of the voiceprint identifier in the past 30 days by the number of active days of the voiceprint identifier in the past 30 days is taken as the average number of times the voiceprint identifier was identified per day (i.e., the average daily number of active times for the voiceprint identifier). From all the voiceprint identifiers appearing under the set-top box, the voiceprint identifiers with an active number greater than one-third of the set-top box's active days and an average daily number of active times greater than 1 are selected as the primary voiceprint identifier.

[0059] Second, all voiceprint identifiers other than the primary voiceprint identifier are designated as secondary voiceprint identifiers.

[0060] After determining the primary voiceprint identifier among all the obtained voiceprint identifiers through the above steps, all other voiceprint identifiers besides the primary voiceprint identifier are used as secondary voiceprint identifiers.

[0061] S203. Based on the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, determine the secondary voiceprint identifier corresponding to each main voiceprint identifier from all secondary voiceprint identifiers, and obtain the voiceprint identifier lookup table.

[0062] Specifically, after determining whether all voiceprint identifiers are primary or secondary voiceprint identifiers, a matching analysis needs to be performed on each primary voiceprint identifier and each secondary voiceprint identifier. From all secondary voiceprint identifiers, the secondary voiceprint identifiers that match each primary voiceprint identifier are identified, and each primary voiceprint identifier is associated with all secondary voiceprint identifiers that match it. The matching analysis between primary and secondary voiceprint identifiers requires first obtaining the voice interaction information corresponding to all user voices of the primary voiceprint identifier and the voice interaction information corresponding to all user voices of the secondary voiceprint identifiers. The behavioral feature similarity between the voice interaction information corresponding to the primary and secondary voiceprint identifiers is calculated. Based on this behavioral feature similarity, a match between the primary and secondary voiceprint identifiers is determined. A higher behavioral feature similarity indicates more similar behavioral features between the users of the primary and secondary voiceprint identifiers, increasing the likelihood that the users of the primary and secondary voiceprint identifiers are the same user, and thus increasing the probability of association between the primary and secondary voiceprint identifiers. Therefore, in this embodiment, a corresponding threshold can be preset, and the behavioral feature similarity between the voice interaction information corresponding to two voiceprint identifiers can be compared with the preset threshold. Two voiceprint identifiers whose behavioral feature similarity exceeds the threshold are judged as a match, and two voiceprint identifiers whose behavioral feature similarity does not exceed the threshold are judged as a mismatch.

[0063] In this embodiment, a voiceprint identifier lookup table is constructed based on the determined matching relationship between each primary voiceprint identifier and each secondary voiceprint identifier. Thus, the secondary voiceprint identifier corresponding to each primary voiceprint identifier can be queried according to the voiceprint identifier lookup table, and both the primary voiceprint identifier and the secondary voiceprint identifier corresponding to the primary voiceprint identifier belong to the same user's voiceprint identifier.

[0064] Furthermore, every so often, it is necessary to obtain the new voiceprint identifiers on the terminal device compared to the last time. Then, the new voiceprint identifiers are divided into primary and secondary voiceprint identifiers according to the above method. The new secondary voiceprint identifiers are then matched and analyzed with both the new primary voiceprint identifiers and the previous primary voiceprint identifiers. This results in the determination of a new voiceprint identifier lookup table for the new voiceprint identifiers. Finally, the new voiceprint identifier lookup table is merged with the previous voiceprint identifier lookup table to obtain the latest voiceprint identifier lookup table.

[0065] As an optional implementation, another embodiment of this application discloses that, after determining the voiceprint identifier lookup table, the following steps are also included:

[0066] First, the voice interaction information of the secondary voiceprint identifiers corresponding to the primary voiceprint identifier is merged and stored in the voice interaction information corresponding to the primary voiceprint identifier.

[0067] Specifically, after determining the voiceprint identifier lookup table in this embodiment, the voice interaction information of all secondary voiceprint identifiers corresponding to the main voiceprint identifier needs to be merged and stored in the voice interaction information of the main voiceprint identifier. In this way, the voice interaction information corresponding to all user voices of each user can be stored based on the main voiceprint identifier of that user, and all voice interaction information corresponding to that user can be queried directly through the main voiceprint identifier of that user.

[0068] Second, the preference data corresponding to the secondary voiceprint identifier corresponding to the primary voiceprint identifier is merged and stored into the preference data corresponding to the primary voiceprint identifier.

[0069] Specifically, each voiceprint identifier corresponds to identified voice interaction information, and each voiceprint identifier needs to undergo behavioral analysis based on its corresponding voice interaction information to obtain the behavioral features corresponding to that voiceprint identifier. This behavioral analysis involves analyzing data such as user actions, voice-text, and text intent contained in the voice interaction information. Since the voice interaction information corresponding to a voiceprint identifier includes all voice interaction information recognized by that voiceprint identifier, the behavioral analysis involves analyzing several sets of user actions, voice-text, and text intent data to obtain the behavioral features corresponding to that voiceprint identifier. Different behaviors have different levels of importance for analyzing user preferences. For example, actions like turning on or off a device are not important for determining user preferences, while the user's control over playing a TV series is more important. Therefore, different weights need to be pre-set for different behaviors. In this embodiment, based on the pre-set weights and the behavioral features corresponding to each voiceprint identifier, the preference data corresponding to each voiceprint identifier is determined, thereby identifying the preference data corresponding to each primary voiceprint identifier and each secondary voiceprint identifier.

[0070] Based on the constructed voiceprint identifier lookup table, the preference data corresponding to the secondary voiceprint identifier corresponding to the primary voiceprint identifier needs to be merged and stored in the preference data corresponding to the primary voiceprint identifier. In this way, all preference data corresponding to each user can be stored based on the primary voiceprint identifier corresponding to that user, and the preference data corresponding to that user can be queried directly through the primary voiceprint identifier corresponding to that user.

[0071] As an optional implementation, see [link to implementation details]. Figure 3 As shown, another embodiment of this application discloses that, in step S203 above, determining the secondary voiceprint identifier corresponding to each primary voiceprint identifier from all secondary voiceprint identifiers based on the voice interaction information corresponding to the primary voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier includes:

[0072] S301. Based on the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, calculate the voiceprint similarity between the main voiceprint identifier and the secondary voiceprint identifier.

[0073] Specifically, this embodiment can utilize the voice interaction information corresponding to the primary voiceprint identifier and the secondary voiceprint identifier to calculate the behavioral feature similarity between the users corresponding to the primary and secondary voiceprint identifiers. The higher the behavioral feature similarity between the users corresponding to the two voiceprint identifiers, the greater the likelihood that the users corresponding to the two voiceprint identifiers are the same user. Therefore, the behavioral feature similarity between the users corresponding to the two voiceprint identifiers can be used to evaluate the matching possibility between the two voiceprint identifiers (i.e., the voiceprint similarity between the two voiceprint identifiers). Thus, this embodiment can use the behavioral feature similarity between the users corresponding to the primary and secondary voiceprint identifiers as the voiceprint similarity between the primary and secondary voiceprint identifiers.

[0074] This step specifically includes:

[0075] First, using the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, calculate the single-attribute similarity between the main voiceprint identifier and the secondary voiceprint identifier.

[0076] In this embodiment, the voice interaction information corresponding to each voiceprint identifier includes: voice text, the age range of the voice speaker (e.g., children, adults, and the elderly), the gender of the voice speaker, and the voice generation time. All content in the voice interaction information can be used to analyze the behavioral feature similarity between the primary and secondary voiceprint identifiers from multiple perspectives. For example, the voice text in the voice interaction information can be used to analyze the semantic similarity between the primary and secondary voiceprint identifiers; the age range of the voice speaker in the voice interaction information can be used to analyze the age similarity between the primary and secondary voiceprint identifiers; the gender of the voice speaker in the voice interaction information can be used to analyze the gender similarity between the primary and secondary voiceprint identifiers; or the voice generation time in the voice interaction information can be used to analyze the active period similarity between the primary and secondary voiceprint identifiers. In this embodiment, semantic similarity, age similarity, gender similarity, and active period similarity can all be used as single-attribute similarities. That is, the single-attribute similarity between the primary and secondary voiceprint identifiers includes at least one of semantic similarity, age similarity, gender similarity, and active period similarity.

[0077] Specifically, when single-attribute similarity includes semantic similarity, this embodiment can take all the speech text in the speech interaction information corresponding to the main voiceprint identifier as the speech text set corresponding to the main voiceprint identifier, and take all the speech text in the speech interaction information corresponding to the secondary voiceprint identifier as the speech text set corresponding to the secondary voiceprint identifier. Then, the overlap between the speech text set corresponding to the main voiceprint identifier and the speech text set corresponding to the secondary voiceprint identifier is calculated, and the overlap is taken as the semantic similarity between the main voiceprint identifier and the secondary voiceprint identifier.

[0078] When single-attribute similarity includes age similarity, the frequency distribution of the age group corresponding to the main voiceprint identifier is determined based on the age group of the speaker in the voice interaction information corresponding to the main voiceprint identifier. That is, the voice interaction information corresponding to the main voiceprint identifier includes voice interaction information corresponding to several users' voices, and each user's voice interaction information contains the age group of the speaker corresponding to that user's voice. Based on the age groups of the speakers corresponding to all users' voices of the main voiceprint identifier, the frequency of each speaker's age group can be determined. For example, if the voice interaction information corresponding to the main voiceprint identifier includes data on the age groups of 10 users' voices, of which 0 are children, 9 are adults, and 1 is an elderly person, then the frequency distribution of the age group corresponding to the main voiceprint identifier can be represented as [0, 9, 1]. The frequency distribution of the age group corresponding to the secondary voiceprint identifier is also determined using the above method based on the age group of the speaker in the voice interaction information corresponding to the secondary voiceprint identifier. Then, the similarity between the age group frequency distribution corresponding to the main voiceprint identifier and the age group frequency distribution corresponding to the secondary voiceprint identifier is calculated, which is used as the age similarity between the main voiceprint identifier and the secondary voiceprint identifier.

[0079] When single-attribute similarity includes gender similarity, the gender frequency distribution corresponding to the main voiceprint identifier is determined based on the gender data of all speakers included in the voice interaction information corresponding to the main voiceprint identifier. For example, if the voice interaction information corresponding to the main voiceprint identifier includes data on the gender of speakers for 10 user voices, with 8 male speakers and 2 female speakers, then the gender frequency distribution corresponding to the main voiceprint identifier can be represented as [8, 2]. The gender frequency distribution corresponding to the secondary voiceprint identifier is also determined using the same method based on the gender data of all speakers in the voice interaction information corresponding to the secondary voiceprint identifier. Then, the similarity between the gender frequency distributions corresponding to the main and secondary voiceprint identifiers is calculated and used as the gender similarity between the main and secondary voiceprint identifiers.

[0080] When single-attribute similarity includes active time period similarity, time period segmentation is first required. This segmentation can be based on a 24-hour day, dividing it into 24 time periods, or it can follow other segmentation methods, such as dividing two hours into one time period. This embodiment does not impose specific restrictions on time period segmentation; the segmentation can be tailored to the actual situation. In this embodiment, the time periods containing the voice generation times of each voice interaction information corresponding to the main voiceprint identifier are all considered active time periods. Then, the number of active times (i.e., the number of voice generation times in the voice interaction information corresponding to the main voiceprint identifier within each active time period) is analyzed. Simultaneously, based on all the voice generation times contained in the voice interaction information corresponding to the secondary voiceprint identifier, the active time periods corresponding to that secondary voiceprint identifier and the number of active times in each active time period are determined.

[0081] Active periods appearing in both the active periods corresponding to the main voiceprint identifier and the active periods corresponding to the secondary voiceprint identifier are defined as the common active periods for both the main and secondary voiceprint identifiers (i.e., periods in which both the main and secondary voiceprint identifiers are identified). Then, based on the number of active periods corresponding to the main voiceprint identifier, the number of active periods of the main voiceprint identifier within the common active periods is extracted, thus obtaining the common active period frequency distribution of the main voiceprint identifier. Similarly, based on the number of active periods corresponding to the secondary voiceprint identifier, the number of active periods of the secondary voiceprint identifier within the common active periods is extracted, thus obtaining the common active period frequency distribution of the secondary voiceprint identifier. The ratio between the number of common active periods and the total number of periods is used as a weight to calculate the similarity between the common active period frequency distributions of the main and secondary voiceprint identifiers. Finally, the product of this similarity and the weight is used as the active period similarity between the main and secondary voiceprint identifiers.

[0082] The similarity calculation can be performed using cosine similarity or Euclidean distance similarity, etc. This embodiment does not limit the method of similarity calculation.

[0083] Second, based on all single-attribute similarities and the pre-set weights corresponding to each single-attribute similarity, calculate the voiceprint similarity between the main voiceprint identifier and the secondary voiceprint identifier.

[0084] In this embodiment, the weights corresponding to each single attribute similarity are pre-set. The product between the single attribute similarity and the weight of the single attribute similarity is calculated as the similarity score corresponding to the single attribute similarity. Then, the total score between the similarity scores corresponding to each single attribute similarity is calculated as the behavioral feature similarity between the user corresponding to the main voiceprint identifier and the user corresponding to the secondary voiceprint identifier, that is, the voiceprint similarity between the main voiceprint identifier and the secondary voiceprint identifier.

[0085] When single-attribute similarity includes four types: semantic similarity, age similarity, gender similarity, and active time period similarity, the formula for calculating voiceprint similarity is as follows:

[0086] When the semantic similarity is not 0, the voiceprint similarity = active period similarity × 0.1 + semantic similarity × 0.3 + 1.0 + age similarity × 0.3 + gender similarity × 0.3;

[0087] When the semantic similarity is 0, the voiceprint similarity = active time period similarity × 0.1 + age similarity × 0.3 + gender similarity × 0.3.

[0088] S302. Based on the voiceprint similarity between the main voiceprint identifier and each secondary voiceprint identifier and a set similarity threshold, select the secondary voiceprint identifier corresponding to the main voiceprint identifier from all secondary voiceprint identifiers.

[0089] Specifically, in this embodiment, the voiceprint similarity between the main voiceprint identifier and each of the secondary voiceprint identifiers is compared with a set similarity threshold. The secondary voiceprint identifiers with a similarity greater than the set similarity threshold are identified as the secondary voiceprint identifiers corresponding to the main voiceprint identifier. In this way, the secondary voiceprint identifiers corresponding to each main voiceprint identifier can be determined. Each main voiceprint identifier can correspond to zero or one or more secondary voiceprint identifiers.

[0090] This step can also be done in the following way:

[0091] First, from all secondary voiceprint identifiers, select those with a voiceprint similarity greater than a set similarity threshold with the primary voiceprint identifier as candidate secondary voiceprint identifiers corresponding to the primary voiceprint identifier.

[0092] The voiceprint similarity between the main voiceprint identifier and each secondary voiceprint identifier is compared with a set similarity threshold. From all secondary voiceprint identifiers, the secondary voiceprint identifiers with a voiceprint similarity greater than the set similarity threshold are selected as candidate secondary voiceprint identifiers corresponding to the main voiceprint identifier.

[0093] Alternatively, this embodiment can also select candidate secondary voiceprint identifiers by sorting voiceprint similarity. That is, the voiceprint similarity between the main voiceprint identifier and each secondary voiceprint identifier is sorted according to its magnitude, and based on a preset number, the secondary voiceprint identifier with the highest voiceprint similarity is selected as the candidate secondary voiceprint identifier corresponding to the main voiceprint identifier. For example, if the preset number is 5, then the 5 secondary voiceprint identifiers with the highest voiceprint similarity are selected as the candidate secondary voiceprint identifiers corresponding to the main voiceprint identifier.

[0094] Second, based on the single-attribute similarity between the candidate secondary voiceprint identifier and the primary voiceprint identifier, the secondary voiceprint identifier corresponding to the primary voiceprint identifier is selected from the candidate secondary voiceprint identifiers.

[0095] After selecting candidate secondary voiceprint identifiers corresponding to the primary voiceprint identifier from all secondary voiceprint identifiers, it is necessary to select the secondary voiceprint identifier corresponding to the primary voiceprint identifier from the candidate secondary voiceprint identifiers based on the single attribute similarity between each candidate secondary voiceprint identifier and the primary voiceprint identifier.

[0096] For example, a secondary voiceprint identifier that conforms to the following rules is selected from the candidate secondary voiceprint identifiers as the secondary voiceprint identifier corresponding to the primary voiceprint identifier:

[0097] When both age similarity and gender similarity are not empty, semantic similarity is not 0 and age similarity is greater than 0.5 and gender similarity is greater than 0.5; or, when age similarity is empty, semantic similarity is not 0 and gender similarity is greater than 0.5; or, when gender similarity is empty, semantic similarity is not 0 and age similarity is greater than 0.5.

[0098] This embodiment first selects candidate secondary voiceprint identifiers corresponding to the primary voiceprint identifier, and then selects the secondary voiceprint identifier corresponding to the primary voiceprint identifier from the candidate secondary voiceprint identifiers. This two-step conditional judgment is used to determine the secondary voiceprint identifier corresponding to the primary voiceprint identifier, improving the accuracy of the matching analysis between the primary and secondary voiceprint identifiers. This embodiment utilizes the single-attribute similarity between the candidate secondary voiceprint identifiers and the primary voiceprint identifier to select the secondary voiceprint identifier corresponding to the primary voiceprint identifier from the candidate secondary voiceprint identifiers. Other single-attribute similarity selection rules can also be used for selection; this embodiment does not limit the choice.

[0099] As an optional implementation, see [link to implementation details]. Figure 4 As shown, another embodiment of this application discloses a voice user recognition method, which further includes:

[0100] S401. Perform user preference analysis on the voice interaction information corresponding to the user's voice of the first voiceprint identifier within the preset collection time period, and determine the current preference data corresponding to the first voiceprint identifier.

[0101] Specifically, the preference data corresponding to each voiceprint identifier needs to be continuously updated based on the user voice input corresponding to each voiceprint identifier. Therefore, this embodiment requires collecting and summarizing the voice interaction information corresponding to the user voice of the first voiceprint identifier at each preset collection time. Then, user preference analysis (i.e., behavior analysis) is performed on all voice interaction information corresponding to the first voiceprint identifier within the preset collection time to analyze the current preference data corresponding to the first voiceprint identifier within the preset collection time. The specific method of analyzing preference data using voice interaction information has been described in detail in the above embodiments and will not be repeated in this embodiment.

[0102] In this embodiment, the preset collection duration is preferably set to one day, so the voice interaction information corresponding to the user's voice with the first voiceprint identifier needs to be collected and summarized every day.

[0103] S402. Update the voiceprint identifier and preference data correspondence table using the preference data obtained by fusing the current preference data corresponding to the first voiceprint identifier and the preference data corresponding to the first main voiceprint identifier.

[0104] Once the current preference data corresponding to the first voiceprint identifier is determined, it is also necessary to look up the preference data corresponding to the first primary voiceprint identifier from the voiceprint identifier and preference data correspondence table based on the pre-determined first primary voiceprint identifier. Then, the current preference data and the preference data corresponding to the first primary voiceprint identifier are merged. Finally, the preference data corresponding to the first primary voiceprint identifier in the voiceprint identifier and preference data correspondence table is updated with the merged preference data. This realizes the updating of user preference data based on the user's daily voice interaction operations, improves the accuracy of user preference data, and thus ensures the accuracy of personalized feedback operations on user voice.

[0105] The fusion of current preference data and preference data corresponding to the pre-recorded first primary voiceprint identifier requires first attenuating the pre-recorded preference data corresponding to the first primary voiceprint identifier, and then merging the attenuated preference data with the current preference data to obtain the fused preference data. Because previously determined preference data may change over time, the acceptability of previous preference data is lower than that of current preference data. Therefore, it is necessary to attenuate previous preference data to improve the accuracy of the preference data. In this embodiment, an attenuation coefficient can be set to attenuate the preference data.

[0106] Corresponding to the above-described voice user recognition method, this application also proposes a voice user recognition device, see [link to relevant documentation]. Figure 5 As shown, the device includes:

[0107] The voiceprint identifier determination module 100 is used to determine the first voiceprint identifier corresponding to the user's voice by extracting the voiceprint features of the user's voice.

[0108] The voiceprint identifier comparison module 110 is used to compare the first voiceprint identifier with the voiceprint identifier in the pre-set voiceprint identifier lookup table to determine the first main voiceprint identifier corresponding to the first voiceprint identifier.

[0109] The voiceprint identifier lookup table contains the primary voiceprint identifier and secondary voiceprint identifier for each user. The primary voiceprint identifier for each user is the voiceprint identifier that is identified at a frequency greater than a set frequency threshold within a preset time period. The secondary voiceprint identifier for the user is the voiceprint identifier whose voiceprint similarity to the user's primary voiceprint identifier is greater than a set similarity threshold.

[0110] The voice user recognition device proposed in this application includes a voiceprint identification module 100 that extracts the voiceprint features of the user's voice to determine a first voiceprint identifier corresponding to the user's voice; and a voiceprint identification comparison module 110 that compares the first voiceprint identifier with voiceprint identifiers in a pre-set voiceprint identifier lookup table to determine a first primary voiceprint identifier corresponding to the first voiceprint identifier. The voiceprint identifier lookup table includes primary and secondary voiceprint identifiers for each user. Each user's primary voiceprint identifier is a voiceprint identifier whose frequency of recognition within a preset time period exceeds a set frequency threshold, and the user's secondary voiceprint identifiers are voiceprint identifiers whose voiceprint similarity to the user's primary voiceprint identifier exceeds a set similarity threshold. By using the technical solution of this embodiment, the voiceprint identifier lookup table can be used to associate all secondary voiceprint identifiers of the same user with the primary voiceprint identifier. When the voiceprint corresponding to the user's voice shifts, the primary voiceprint identifier of the user can be accurately retrieved from the voiceprint identifier lookup table using the secondary voiceprint identifier after the shift, thus improving the accuracy of user information determination.

[0111] As an optional implementation, another embodiment of this application also discloses a voiceprint identification comparison module 110, specifically used for:

[0112] Compare the first voiceprint identifier with the secondary voiceprint identifier in the pre-set voiceprint identifier lookup table;

[0113] If there is a secondary voiceprint identifier in the voiceprint identifier lookup table that is the same as the first voiceprint identifier, then look up the primary voiceprint identifier corresponding to the secondary voiceprint identifier that is the same as the first voiceprint identifier in the voiceprint identifier lookup table and use it as the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0114] If there is no secondary voiceprint identifier in the voiceprint identifier lookup table that is the same as the first voiceprint identifier, then the first voiceprint identifier will be used as the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0115] As an optional implementation, another embodiment of this application also discloses a voiceprint identification comparison module 110, which is specifically used for:

[0116] Compare the first voiceprint identifier with the main voiceprint identifier in the pre-set voiceprint identifier lookup table;

[0117] If there is a primary voiceprint identifier in the voiceprint identifier lookup table that is the same as the first voiceprint identifier, then the primary voiceprint identifier that is the same as the first voiceprint identifier will be used as the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0118] If there is no primary voiceprint identifier in the voiceprint identifier lookup table that is the same as the first voiceprint identifier, then look up the primary voiceprint identifier corresponding to the secondary voiceprint identifier that is the same as the first voiceprint identifier in the voiceprint identifier lookup table, and use it as the first primary voiceprint identifier corresponding to the first voiceprint identifier.

[0119] As an optional implementation, another embodiment of this application also discloses a voice user recognition device, which further includes: a voiceprint acquisition module, a voiceprint segmentation module, and a voiceprint correspondence module.

[0120] The voiceprint acquisition module is used to acquire voiceprint identifiers recognized within a preset time period;

[0121] The voiceprint segmentation module is used to determine the primary and secondary voiceprint identifiers among all voiceprint identifiers based on the frequency of voiceprint identifiers being identified within a preset time period.

[0122] The voiceprint mapping module is used to determine the secondary voiceprint identifier corresponding to each primary voiceprint identifier from all secondary voiceprint identifiers based on the voice interaction information corresponding to the primary voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, and to obtain a voiceprint identifier lookup table.

[0123] As an optional implementation, another embodiment of this application also discloses a voice user recognition device, which further includes a merging storage module.

[0124] The merging and storage module is used to merge and store the voice interaction information of the secondary voiceprint identifier corresponding to the main voiceprint identifier into the voice interaction information corresponding to the main voiceprint identifier; and to merge and store the preference data corresponding to the secondary voiceprint identifier corresponding to the main voiceprint identifier into the preference data corresponding to the main voiceprint identifier.

[0125] As an optional implementation, another embodiment of this application also discloses a voiceprint segmentation module, specifically used for:

[0126] Voiceprint identifiers that are identified for more than a preset number of days within a preset time period and are identified more than a preset number of times per day on average are selected as primary voiceprint identifiers.

[0127] All voiceprint identifiers other than the primary voiceprint identifier are designated as secondary voiceprint identifiers.

[0128] As an optional implementation, another embodiment of this application also discloses a voiceprint correspondence module, including: a similarity calculation unit and a voiceprint selection unit.

[0129] The similarity calculation unit is used to calculate the voiceprint similarity between the main voiceprint identifier and the secondary voiceprint identifier based on the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier.

[0130] The voiceprint selection unit is used to select the secondary voiceprint identifier corresponding to the main voiceprint identifier from all secondary voiceprint identifiers based on the voiceprint similarity between the main voiceprint identifier and each secondary voiceprint identifier and a set similarity threshold.

[0131] As an optional implementation, another embodiment of this application also discloses a similarity calculation unit, specifically used for:

[0132] Using the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, calculate the single-attribute similarity between the main voiceprint identifier and the secondary voiceprint identifier; wherein, the voice interaction information includes: voice text, voice speaker age group, voice speaker gender, and voice generation time; the single-attribute similarity includes at least one of: semantic similarity, age similarity, gender similarity, and active time period similarity.

[0133] The voiceprint similarity between the main voiceprint identifier and the secondary voiceprint identifier is calculated based on all single-attribute similarities and the pre-set weights corresponding to each single-attribute similarity.

[0134] As an optional implementation, another embodiment of this application also discloses that, if single-attribute similarity includes semantic similarity, the similarity calculation unit is specifically used for:

[0135] Extract the voice text set of the voice interaction information corresponding to the main voiceprint identifier, and extract the voice text set of the voice interaction information corresponding to the secondary voiceprint identifier;

[0136] The overlap between the set of speech texts corresponding to the main voiceprint identifier and the set of speech texts corresponding to the secondary voiceprint identifier is calculated as the semantic similarity between the main voiceprint identifier and the secondary voiceprint identifier.

[0137] As an optional implementation, another embodiment of this application also discloses that, if the single-attribute similarity includes age similarity, the similarity calculation unit is specifically used for:

[0138] Based on the age range of the speaker in the voice interaction information corresponding to the main voiceprint identifier and the age range of the speaker in the voice interaction information corresponding to the secondary voiceprint identifier, determine the frequency distribution of the age range corresponding to the main voiceprint identifier and the frequency distribution of the age range corresponding to the secondary voiceprint identifier.

[0139] Calculate the similarity between the frequency distribution of the age group corresponding to the main voiceprint identifier and the frequency distribution of the age group corresponding to the secondary voiceprint identifier, and use it as the age similarity between the main voiceprint identifier and the secondary voiceprint identifier;

[0140] If single-attribute similarity includes gender similarity, the similarity calculation unit is specifically used for:

[0141] Based on the gender of the speaker in the voice interaction information corresponding to the main voiceprint identifier and the gender of the speaker in the voice interaction information corresponding to the secondary voiceprint identifier, determine the gender frequency distribution corresponding to the main voiceprint identifier and the gender frequency distribution corresponding to the secondary voiceprint identifier.

[0142] Calculate the similarity between the gender frequency distribution corresponding to the main voiceprint identifier and the gender frequency distribution corresponding to the secondary voiceprint identifier, and use it as the gender similarity between the main voiceprint identifier and the secondary voiceprint identifier.

[0143] As an optional implementation, another embodiment of this application also discloses that, if the single-attribute similarity includes active time period similarity, the similarity calculation unit is specifically used for:

[0144] Based on the speech generation time of the speech interaction information corresponding to the main voiceprint identifier and the speech generation time of the speech interaction information corresponding to the secondary voiceprint identifier, the common active period between the main voiceprint identifier and the secondary voiceprint identifier is determined; the common active period is the period when both the main voiceprint identifier and the secondary voiceprint identifier are identified.

[0145] Based on the speech generation time of the speech interaction information corresponding to the main voiceprint identifier, the speech generation time of the speech interaction information corresponding to the secondary voiceprint identifier, and the common active period, determine the distribution of the number of common active periods of the main voiceprint identifier, the distribution of the number of common active periods of the secondary voiceprint identifier, and the weight corresponding to the number of common active periods.

[0146] Calculate the similarity between the common active period frequency distribution of the main voiceprint identifier and the common active period frequency distribution of the secondary voiceprint identifier, and use the product of the similarity and the weight as the active period similarity between the main voiceprint identifier and the secondary voiceprint identifier.

[0147] As an optional implementation, another embodiment of this application also discloses a voiceprint selection unit, specifically used for:

[0148] From all secondary voiceprint identifiers, select those with a voiceprint similarity greater than a set similarity threshold with the primary voiceprint identifier, and use them as candidate secondary voiceprint identifiers corresponding to the primary voiceprint identifier.

[0149] Based on the single-attribute similarity between the candidate secondary voiceprint identifier and the primary voiceprint identifier, the secondary voiceprint identifier corresponding to the primary voiceprint identifier is selected from the candidate secondary voiceprint identifiers.

[0150] As an optional implementation, another embodiment of this application also discloses a voice user recognition device, which further includes: a preference data acquisition module, used to acquire preference data corresponding to a first primary voiceprint identifier based on a pre-set voiceprint identifier and preference data correspondence table, as the preference data of the user corresponding to the first primary voiceprint identifier.

[0151] As an optional implementation, another embodiment of this application also discloses a voice user recognition device, which further includes: a preference analysis module and a fusion module;

[0152] The preference analysis module is used to perform user preference analysis on the voice interaction information corresponding to the user's voice of the first voiceprint identifier within a preset collection time, and to determine the current preference data corresponding to the first voiceprint identifier.

[0153] The fusion module is used to update the voiceprint identifier and preference data correspondence table by using the preference data obtained by fusing the current preference data corresponding to the first voiceprint identifier and the preference data corresponding to the first main voiceprint identifier.

[0154] As an optional implementation, another embodiment of this application also discloses a voice user recognition device, which further includes: an information merging module, used to merge the voice interaction information corresponding to the user's voice into the voice interaction information corresponding to the first main voiceprint identifier.

[0155] The voice user recognition device provided in this embodiment belongs to the same application concept as the voice user recognition method provided in the above embodiments of this application. It can execute the voice user recognition method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects of executing the voice user recognition method. Technical details not described in detail in this embodiment can be found in the specific processing content of the voice user recognition method provided in the above embodiments of this application, and will not be repeated here.

[0156] Another embodiment of this application discloses an electronic device, see [link to relevant documentation] Figure 6 As shown, the device includes:

[0157] Memory 200 and processor 210;

[0158] The memory 200 is connected to the processor 210 and is used to store programs;

[0159] The processor 210 is configured to implement the voice user recognition method disclosed in any of the above embodiments by running the program stored in the memory 200.

[0160] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0161] The processor 210, memory 200, communication interface 220, input device 230, and output device 240 are interconnected via a bus. Among them:

[0162] A bus can include a pathway for transmitting information between various components of a computer system.

[0163] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0164] Processor 210 may include a main processor, as well as a baseband chip, modem, etc.

[0165] The memory 200 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0166] Input device 230 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0167] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0168] The communication interface 220 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0169] The processor 2102 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of the voice user recognition method provided in the embodiments of this application.

[0170] Another embodiment of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the various steps of the voice user recognition method provided in any of the above embodiments.

[0171] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0172] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0173] The steps in the methods of the various embodiments of this application can be adjusted, combined, or deleted according to actual needs.

[0174] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.

[0175] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0176] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0177] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0178] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0179] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0180] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0181] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for voice user recognition, characterized in that, include: The first voiceprint identifier corresponding to the user's voice is determined by extracting the voiceprint features of the user's voice. The first voiceprint identifier is compared with the voiceprint identifier in the pre-set voiceprint identifier lookup table. The voiceprint identifier that is the same as the first voiceprint identifier is found from the voiceprint identifier lookup table. Based on the voiceprint identifier that is the same as the first voiceprint identifier, the first main voiceprint identifier corresponding to the first voiceprint identifier is determined. The voiceprint identifier lookup table contains the primary voiceprint identifier and secondary voiceprint identifier for each user. All secondary voiceprint identifiers of the same user are associated with the primary voiceprint identifier. The primary voiceprint identifier of each user is the voiceprint identifier that is identified at a frequency greater than a set frequency threshold within a preset time period. The secondary voiceprint identifier of the user is the voiceprint identifier that has a voiceprint similarity greater than a set similarity threshold with the primary voiceprint identifier of the user. If the voiceprint identifier that is the same as the first voiceprint identifier is the primary voiceprint identifier, then the voiceprint identifier that is the same as the first voiceprint identifier is used as the first primary voiceprint identifier. If the voiceprint identifier that is the same as the first voiceprint identifier is the secondary voiceprint identifier, then the primary voiceprint identifier associated with the voiceprint identifier that is the same as the first voiceprint identifier is used as the first primary voiceprint identifier.

2. The method according to claim 1, characterized in that, The first voiceprint identifier is compared with voiceprint identifiers in a pre-set voiceprint identifier lookup table. A voiceprint identifier matching the first voiceprint identifier is found in the lookup table. Based on this matching voiceprint identifier, a first primary voiceprint identifier corresponding to the first voiceprint identifier is determined, including: Compare the first voiceprint identifier with the secondary voiceprint identifier in the pre-set voiceprint identifier lookup table; If there is a secondary voiceprint identifier in the voiceprint identifier lookup table that is the same as the first voiceprint identifier, then the primary voiceprint identifier corresponding to the secondary voiceprint identifier that is the same as the first voiceprint identifier is retrieved from the voiceprint identifier lookup table and used as the first primary voiceprint identifier corresponding to the first voiceprint identifier. If there is no secondary voiceprint identifier in the voiceprint identifier lookup table that is the same as the first voiceprint identifier, then the first voiceprint identifier is used as the first primary voiceprint identifier corresponding to the first voiceprint identifier.

3. The method according to claim 1, characterized in that, The first voiceprint identifier is compared with voiceprint identifiers in a pre-set voiceprint identifier lookup table. A voiceprint identifier matching the first voiceprint identifier is found in the lookup table. Based on this matching voiceprint identifier, a first primary voiceprint identifier corresponding to the first voiceprint identifier is determined, including: Compare the first voiceprint identifier with the main voiceprint identifier in the pre-set voiceprint identifier lookup table; If the voiceprint identifier lookup table contains a main voiceprint identifier that is the same as the first voiceprint identifier, then the main voiceprint identifier that is the same as the first voiceprint identifier shall be used as the first main voiceprint identifier corresponding to the first voiceprint identifier. If there is no primary voiceprint identifier in the voiceprint identifier lookup table that is the same as the first voiceprint identifier, then the primary voiceprint identifier corresponding to the secondary voiceprint identifier that is the same as the first voiceprint identifier is retrieved from the voiceprint identifier lookup table and used as the first primary voiceprint identifier corresponding to the first voiceprint identifier.

4. The method according to claim 1, characterized in that, The voiceprint identifier lookup table is obtained through the following processing: Obtain voiceprint identifiers recognized within a preset time period; Based on the frequency of voiceprint identification within a preset time period, determine the primary and secondary voiceprint identifiers among all voiceprint identifiers. Based on the voice interaction information corresponding to the primary voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, the secondary voiceprint identifier corresponding to each primary voiceprint identifier is determined from all secondary voiceprint identifiers, thus obtaining the voiceprint identifier lookup table.

5. The method according to claim 4, characterized in that, Also includes: The voice interaction information of the secondary voiceprint identifier corresponding to the primary voiceprint identifier is merged and stored into the voice interaction information corresponding to the primary voiceprint identifier. The preference data corresponding to the secondary voiceprint identifier corresponding to the primary voiceprint identifier are merged and stored into the preference data corresponding to the primary voiceprint identifier.

6. The method according to claim 4, characterized in that, Based on the frequency of voiceprint identification within a preset time period, the primary and secondary voiceprint identifiers among all voiceprint identifiers are determined, including: Voiceprint identifiers that are identified for more than a preset number of days within a preset time period and are identified more than a preset number of times per day on average are selected as primary voiceprint identifiers. All voiceprint identifiers other than the primary voiceprint identifier are designated as secondary voiceprint identifiers.

7. The method according to claim 4, characterized in that, Based on the voice interaction information corresponding to the primary voiceprint identifier and the secondary voiceprint identifier, the secondary voiceprint identifier corresponding to each primary voiceprint identifier is determined from all secondary voiceprint identifiers, including: Based on the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, calculate the voiceprint similarity between the main voiceprint identifier and the secondary voiceprint identifier. Based on the voiceprint similarity between the main voiceprint identifier and each secondary voiceprint identifier and a set similarity threshold, the secondary voiceprint identifier corresponding to the main voiceprint identifier is selected from all secondary voiceprint identifiers.

8. The method according to claim 7, characterized in that, Based on the voice interaction information corresponding to the primary voiceprint identifier and the secondary voiceprint identifier, the voiceprint similarity between the primary and secondary voiceprint identifiers is calculated, including: Using the voice interaction information corresponding to the primary voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, a single-attribute similarity is calculated between the primary voiceprint identifier and the secondary voiceprint identifier; wherein, the voice interaction information includes: voice text, age group of the voice speaker, gender of the voice speaker, and voice generation time; the single-attribute similarity includes at least one of: semantic similarity, age similarity, gender similarity, and active time period similarity; The voiceprint similarity between the main voiceprint identifier and the secondary voiceprint identifier is calculated based on all single-attribute similarities and the pre-set weights corresponding to each single-attribute similarity.

9. The method according to claim 8, characterized in that, If the single-attribute similarity includes semantic similarity, the semantic similarity between the main voiceprint identifier and the secondary voiceprint identifier is calculated using the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, including: Extract the voice text set of the voice interaction information corresponding to the main voiceprint identifier, and extract the voice text set of the voice interaction information corresponding to the secondary voiceprint identifier; The overlap between the speech text set corresponding to the main voiceprint identifier and the speech text set corresponding to the secondary voiceprint identifier is calculated, and used as the semantic similarity between the main voiceprint identifier and the secondary voiceprint identifier.

10. The method according to claim 8, characterized in that, If the single-attribute similarity includes age similarity, the age similarity between the main voiceprint identifier and the secondary voiceprint identifier is calculated using the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, including: Based on the age range of the speaker in the voice interaction information corresponding to the main voiceprint identifier and the age range of the speaker in the voice interaction information corresponding to the secondary voiceprint identifier, determine the frequency distribution of the age range corresponding to the main voiceprint identifier and the frequency distribution of the age range corresponding to the secondary voiceprint identifier. Calculate the similarity between the age group frequency distribution corresponding to the main voiceprint identifier and the age group frequency distribution corresponding to the secondary voiceprint identifier, and use it as the age similarity between the main voiceprint identifier and the secondary voiceprint identifier; If the single-attribute similarity includes gender similarity, the gender similarity between the main voiceprint identifier and the secondary voiceprint identifier is calculated using the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, including: Based on the gender of the speaker in the voice interaction information corresponding to the main voiceprint identifier and the gender of the speaker in the voice interaction information corresponding to the secondary voiceprint identifier, determine the gender frequency distribution corresponding to the main voiceprint identifier and the gender frequency distribution corresponding to the secondary voiceprint identifier. The similarity between the gender frequency distribution corresponding to the main voiceprint identifier and the gender frequency distribution corresponding to the secondary voiceprint identifier is calculated and used as the gender similarity between the main voiceprint identifier and the secondary voiceprint identifier.

11. The method according to claim 8, characterized in that, If the single-attribute similarity includes active time period similarity, the active time period similarity between the main voiceprint identifier and the secondary voiceprint identifier is calculated using the voice interaction information corresponding to the main voiceprint identifier and the voice interaction information corresponding to the secondary voiceprint identifier, including: Based on the speech generation time of the speech interaction information corresponding to the main voiceprint identifier and the speech generation time of the speech interaction information corresponding to the secondary voiceprint identifier, the common active period between the main voiceprint identifier and the secondary voiceprint identifier is determined; the common active period is the period during which both the main voiceprint identifier and the secondary voiceprint identifier are identified. Based on the voice generation time of the voice interaction information corresponding to the main voiceprint identifier, the voice generation time of the voice interaction information corresponding to the secondary voiceprint identifier, and the common active period, determine the distribution of the number of common active periods of the main voiceprint identifier, the distribution of the number of common active periods of the secondary voiceprint identifier, and the weight corresponding to the number of common active periods. Calculate the similarity between the common active period frequency distribution of the main voiceprint identifier and the common active period frequency distribution of the secondary voiceprint identifier, and use the product of the similarity and the weight as the active period similarity between the main voiceprint identifier and the secondary voiceprint identifier.

12. The method according to claim 7, characterized in that, Based on the voiceprint similarity between the main voiceprint identifier and each secondary voiceprint identifier, and a set similarity threshold, the secondary voiceprint identifier corresponding to the main voiceprint identifier is selected from all secondary voiceprint identifiers, including: From all secondary voiceprint identifiers, select those with a voiceprint similarity greater than a set similarity threshold with the primary voiceprint identifier, and use them as candidate secondary voiceprint identifiers corresponding to the primary voiceprint identifier. Based on the single-attribute similarity between the candidate secondary voiceprint identifier and the primary voiceprint identifier, the secondary voiceprint identifier corresponding to the primary voiceprint identifier is selected from the candidate secondary voiceprint identifiers.

13. The method according to claim 1, characterized in that, Also includes: Based on a pre-set voiceprint identifier and preference data correspondence table, the preference data corresponding to the first primary voiceprint identifier is obtained and used as the preference data of the user corresponding to the first primary voiceprint identifier.

14. The method according to claim 13, characterized in that, Also includes: User preference analysis is performed on the voice interaction information corresponding to the user's voice of the first voiceprint identifier within a preset collection time to determine the current preference data corresponding to the first voiceprint identifier; The voiceprint identifier and preference data correspondence table is updated using the preference data obtained by fusing the current preference data corresponding to the first voiceprint identifier and the preference data corresponding to the first main voiceprint identifier.

15. The method according to claim 1, characterized in that, Also includes: The voice interaction information corresponding to the user's voice is merged into the voice interaction information corresponding to the first primary voiceprint identifier.

16. A voice user recognition device, characterized in that, include: The voiceprint identifier determination module is used to determine a first voiceprint identifier corresponding to the user's voice by extracting the voiceprint features of the user's voice; The voiceprint identifier comparison module is used to compare the first voiceprint identifier with the voiceprint identifier in a pre-set voiceprint identifier lookup table, find the voiceprint identifier that is the same as the first voiceprint identifier from the voiceprint identifier lookup table, and determine the first main voiceprint identifier corresponding to the first voiceprint identifier based on the voiceprint identifier that is the same as the first voiceprint identifier. The voiceprint identifier lookup table contains the primary voiceprint identifier and secondary voiceprint identifier for each user. All secondary voiceprint identifiers of the same user are associated with the primary voiceprint identifier. The primary voiceprint identifier of each user is the voiceprint identifier that is identified at a frequency greater than a set frequency threshold within a preset time period. The secondary voiceprint identifier of the user is the voiceprint identifier that has a voiceprint similarity greater than a set similarity threshold with the primary voiceprint identifier of the user. If the voiceprint identifier that is the same as the first voiceprint identifier is the primary voiceprint identifier, then the voiceprint identifier that is the same as the first voiceprint identifier is used as the first primary voiceprint identifier. If the voiceprint identifier that is the same as the first voiceprint identifier is the secondary voiceprint identifier, then the primary voiceprint identifier associated with the voiceprint identifier that is the same as the first voiceprint identifier is used as the first primary voiceprint identifier.

17. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the voice user recognition method as described in any one of claims 1 to 15 by running a program in the memory.

18. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the voice user recognition method as described in any one of claims 1 to 15.

Citation Information

Patent Citations

  • Identity authentication method, apparatus and device based on voice-print identification, and storage medium

    CN109584886A

  • User distinguishing method, user behavior library determination method, device and equipment

    CN111429920A