Voiceprint template updating method and related device
By dynamically updating the voiceprint template and utilizing the user's voice and image information, the update ratio of voiceprint features is optimized, which solves the problem of low recognition accuracy caused by changes in the user's voice, improves the accuracy and recall rate of voiceprint recognition, and reduces the false recognition rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2021-11-25
- Publication Date
- 2026-04-17
AI Technical Summary
During voiceprint recognition, changes in a user's voice can lead to low recognition accuracy, and existing technologies struggle to effectively update voiceprint templates to adapt to these changes.
This paper provides a method for updating voiceprint templates. By acquiring the user's voice and image information, voiceprint features are extracted. Based on the comparison of the features with the registered voiceprint templates, the voiceprint templates are dynamically updated or re-registered. The update ratio is optimized by utilizing the confidence and influence weight of biometric information to improve the recognition accuracy.
By dynamically updating the voiceprint template, the accuracy of voiceprint recognition is improved, adapting to changes in the user's voice, enhancing the accuracy and recall rate of recognition, and reducing the false recognition rate.
Smart Images

Figure CN116168708B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a method for updating voiceprint templates and related equipment. Background Technology
[0002] Voiceprint recognition technology is a technique that uses a speaker's voiceprint to identify the speaker's identity. It can be effectively applied in fields such as smart homes, smart buildings, and financial security. The voiceprint recognition process generally includes two parts: voiceprint registration and voiceprint verification.
[0003] Voiceprint registration may include: acquiring the speaker's voice information, extracting corresponding voiceprint features based on the speaker's voice information, and generating a voiceprint template corresponding to the speaker based on the extracted voiceprint features. Voiceprint verification may include: when the voice information of a certain speaker is detected, extracting corresponding voiceprint features based on the speaker's voice information, and comparing the extracted voiceprint features with the voiceprint template generated in voiceprint registration to determine whether the current speaker is the same person as the speaker at the time of voiceprint registration. Summary of the Invention
[0004] In some possible scenarios, the speaker's voice during voiceprint registration may change and differ from the voice at the time of registration, which may lead to low accuracy in voiceprint recognition.
[0005] To address the aforementioned technical problems, this application provides a voiceprint template updating method and related equipment. The technical solution provided by this application can dynamically update a user's voiceprint template based on the degree of change in the user's voice, effectively improving the accuracy of voiceprint recognition.
[0006] In a first aspect, this application provides a voiceprint template updating method, the method comprising: acquiring voice information of a first user; the first user being a user who has registered a voiceprint template; extracting voiceprint features corresponding to the voice information; obtaining a first scoring result based on the voiceprint features and the voiceprint template already registered by the first user; updating the voiceprint template already registered by the first user based on the voiceprint features when the first scoring result is greater than a first voiceprint threshold; or, re-registering the voiceprint template of the first user based on the voiceprint features when the first scoring result is less than a second voiceprint threshold. The second voiceprint threshold is less than or equal to the first voiceprint threshold.
[0007] For a first user who has already registered a voiceprint template, this voiceprint template update method obtains a first score result based on the voiceprint features and the first user's registered voiceprint template. When the first score result is greater than a first voiceprint threshold, the first user's registered voiceprint template is updated based on the voiceprint features. Alternatively, when the first score result is less than a second voiceprint threshold, the first user's voiceprint template is re-registered based on the voiceprint features. This allows for dynamic updates to the first user's voiceprint template according to the degree of change in the first user's voice, resulting in higher accuracy when using the first user's voiceprint template for voiceprint recognition.
[0008] When the first voiceprint threshold and the second voiceprint threshold are the same (i.e., equal), there is only one voiceprint threshold (which is both the first and second voiceprint threshold). In this method, when the first score result is greater than the voiceprint threshold, the voiceprint template registered by the first user is updated according to the voiceprint features; or, when the first score result is less than the voiceprint threshold, the voiceprint template of the first user is re-registered according to the voiceprint features.
[0009] Optionally, when the first score is equal to the voiceprint threshold, the voiceprint template registered by the first user can be updated based on the voiceprint features, or the voiceprint template of the first user can be re-registered based on the voiceprint features, without restriction.
[0010] When the first voiceprint threshold and the second voiceprint threshold are different (i.e., the second voiceprint threshold is less than the first voiceprint threshold), no processing is required for the first score result that is greater than the second voiceprint threshold but less than the first voiceprint threshold, such as keeping the first user's voiceprint template unchanged.
[0011] Optionally, when the second voiceprint threshold is less than the first voiceprint threshold, if the first score equals the second voiceprint threshold, it can be considered that the first user's voice has undergone a sudden change, and the voiceprint template corresponding to the first user is re-registered. Alternatively, it can be considered that the sudden change or gradual change in the first user's voice cannot be distinguished, and the voiceprint template corresponding to the first user remains unchanged. If the first score equals the first voiceprint threshold, it can be considered that the first user's voice has undergone a gradual change, and the voiceprint template corresponding to the first user is updated according to the voiceprint features corresponding to the first user's voice information. Alternatively, it can be considered that the sudden change or gradual change in the first user's voice cannot be distinguished, and the voiceprint template corresponding to the first user remains unchanged.
[0012] According to the first aspect, updating the voiceprint template registered by the first user based on voiceprint features includes: obtaining the product between the voiceprint template registered by the first user and a preset first value to obtain a first product result; obtaining the product between the voiceprint feature and the corresponding second value to obtain a second product result; summing the first product result and the second product result to obtain a first summation result; summing the first value and the second value to obtain a second summation result; obtaining the ratio of the first summation result and the second summation result to obtain a new voiceprint template; and replacing the voiceprint template registered by the first user with the new voiceprint template.
[0013] This application does not impose any limitation on the size of the first value. When the first value is larger, the proportion of the voiceprint features corresponding to the first user's voice information in the new voiceprint template is smaller; when the first value is smaller, the proportion of the voiceprint features corresponding to the first user's voice information in the new voiceprint template is larger. In practical applications, the first value can be set to a larger or smaller value according to specific needs, and no special restrictions are imposed here.
[0014] For example, in one possible implementation, the process of generating a new voiceprint template when updating the voiceprint template registered by the first user based on the voiceprint features can be implemented according to the following formula (4). That is, the voiceprint features corresponding to the voice information of the first user can be proportionally superimposed onto the voiceprint template corresponding to the first user according to the following formula (4).
[0015]
[0016] In formula (4), μ′ k This represents a new voiceprint template; t represents the sum of the number of voiceprint features corresponding to the first user's voice information and the number of voiceprint templates corresponding to the first user, where the first user's voiceprint template is generally one, and the voiceprint features corresponding to the first user's voice information can be one or more; x t This refers to a specific voiceprint feature or voiceprint template (i.e., the voiceprint template already registered by the first user); when x t When representing voiceprint features, γ kt The value of x can be 1; when x t When representing a voiceprint template, γ kt The value can take any value within the range of 6 to 20, such as 10, 12, etc.; ∑ t γ kt x t This means taking γ k1 x1 to γ kt x t The sum; ∑ t γ kt This means taking γ k1 To γ ktThe sum of.
[0017] Optionally, the second value corresponding to the voiceprint feature is related to the confidence level of the first user's biometric information in at least two dimensions, and the influence weight of the biometric information in each of the at least two dimensions.
[0018] For example, before obtaining the product between the voiceprint feature and the second value corresponding to the voiceprint feature, the method further includes: obtaining the second value corresponding to the voiceprint feature based on the confidence level corresponding to the first user's biometric information in at least two dimensions, and the influence weight corresponding to the biometric information in each of the at least two dimensions.
[0019] In one implementation, obtaining the second value corresponding to the voiceprint feature based on the confidence levels corresponding to the biometric information of the first user in at least two dimensions and the influence weights corresponding to the biometric information in each of the at least two dimensions includes: summing the influence weights corresponding to the biometric information of the first user in at least two dimensions to obtain a third summation result; obtaining the ratio between the influence weights corresponding to the biometric information in each of the at least two dimensions and the confidence levels to obtain the ratios corresponding to the biometric information in each dimension; summing the ratios corresponding to the biometric information in each of the at least two dimensions to obtain a fourth summation result; and obtaining the ratio between the third summation result and the fourth summation result to obtain the second value corresponding to the voiceprint feature.
[0020] For example, taking the calculation of the second value based on the confidence level corresponding to the biological information of k dimensions (k is an integer greater than 1) and the influence weight corresponding to the biological information of each of the k dimensions as an example, the second value can be calculated by the following formula (5).
[0021]
[0022] In formula (5), a k ∑a represents the influence weight corresponding to the biological information in the k-th dimension; k This represents the summation of the influence weights corresponding to the k dimensions of biological information; This represents the confidence level corresponding to the biological information in the k-th dimension; This means that the ratio between the influence weight and the confidence level corresponding to the biological information of each dimension is calculated first, and then the ratios corresponding to the biological information of each of the k dimensions are summed; λ represents the second value.
[0023] In this implementation, the second value is calculated based on the confidence level corresponding to multiple (e.g., at least two) dimensions of biological information, as well as the influence weight corresponding to each dimension of biological information in the aforementioned multiple dimensions. This can better optimize the proportion of voiceprint template updates, so as to achieve the purpose of updating the voiceprint template based on multi-dimensional fusion information, more objectively reflect the accuracy of each dimension, and further improve the accuracy of voiceprint recognition.
[0024] Optionally, the above-mentioned at least two dimensions of biometric information may include: visual biometric information, voiceprint biometric information, lip movement biometric information, etc., without limitation.
[0025] According to the first aspect, or any implementation of the first aspect above, the first voiceprint threshold and the second voiceprint threshold satisfy the following: precision is greater than or equal to the first precision threshold, recall is greater than or equal to the first recall threshold, and false entry rate is less than or equal to the first false entry rate threshold.
[0026] For example, the first precision threshold could be 95%, the first recall threshold could be 95%, and the first false entry threshold could be 5%.
[0027] According to the first aspect, or any implementation of the first aspect above, the acquisition of the first user's voice information includes: when someone is detected speaking, acquiring image information and the voice information emitted by the speaker; the image information is related to the speaker; based on the image information and the voice information emitted by the speaker, filtering out the speaker's image from the image information; based on the speaker's image and a preset user image, identifying whether the speaker is a first user who has completed voiceprint registration; when the speaker is a first user who has completed voiceprint registration, acquiring the first user's voice information.
[0028] In one possible implementation, the step of filtering out the speaker's image from the image information based on the image information and the speaker's voice information includes: performing lip movement recognition based on the image information to identify the person who has made lip movements from the people included in the image information; performing sound source localization based on the voice information to obtain the first location of the speaker; and identifying the speaker from the people who have made lip movements based on the first location and obtaining the speaker's image.
[0029] In another possible implementation, the step of filtering the speaker's image from the image information based on the image information and the speaker's voice information includes: performing lip reading recognition based on the image information to identify the persons who have lip movements from the persons included in the image information, and the natural language statement corresponding to the lip movements of each person who has lip movements; performing speech recognition based on the voice information to obtain the natural language statement corresponding to the voice information; and identifying the speaker from the persons who have lip movements based on the natural language statement corresponding to the voice information and the natural language statement corresponding to the lip movements of each person who has lip movements, and obtaining the speaker's image.
[0030] In another possible implementation, the step of filtering out the speaker's image from the image information based on the image information and the speaker's voice information includes: performing lip movement recognition and lip reading recognition based on the image information, and performing sound source localization and speech recognition based on the voice information, determining the person who meets the following conditions 1 and / or 2 from the persons included in the image information as the speaker, and obtaining the speaker's image.
[0031] Condition 1: Lip movement occurred, and its location was consistent with the sound source localization result.
[0032] Condition 2: Lip movement occurred, and the natural language statement corresponding to the lip movement was the same as the natural language statement corresponding to the sound information.
[0033] In another possible implementation, the step of filtering out the speaker's image from the image information based on the image information and the speaker's voice information includes: performing lip movement recognition and lip reading recognition based on the image information, performing sound source localization and speech recognition based on the voice information, and performing interaction intention recognition based on the image information, determining the person who meets one or more of the following conditions 1, 2, and 3 from the persons included in the image information as the speaker, and obtaining the speaker's image.
[0034] Condition 1: Lip movement occurred, and its location is consistent with the sound source localization result; Condition 2: Lip movement occurred, and the natural language sentence corresponding to the lip movement is the same as the natural language sentence corresponding to the sound information; Condition 3: The interaction willingness value meets the preset requirements.
[0035] In one possible implementation, the interaction willingness value in condition 3 above must meet a preset requirement, which may include: the interaction willingness value is the highest among all the interaction willingness values of the people included in the image. For example, if the image includes 3 people, and user 1's interaction willingness value is the highest among the 3 people, then user 1 meets condition 3.
[0036] In another possible implementation, the interaction willingness value in condition 3 above must meet a preset requirement, which may include: the interaction willingness value is greater than or equal to a preset interaction willingness threshold. For example, the interaction willingness threshold may be 0.7, 0.8, etc. When a person's interaction willingness value is greater than or equal to the interaction willingness threshold, it can be determined that the person meets condition 3.
[0037] Optionally, when there are multiple people whose interaction willingness value is greater than or equal to the interaction willingness threshold, the person with the highest interaction willingness value can be determined from among the multiple people as the person who meets condition 3.
[0038] In this implementation, when a person in the image simultaneously meets conditions 1, 2, and 3, that person is identified as the speaker emitting the sound information, ensuring that only that person is engaging in dialogue or interaction with the intelligent robot.
[0039] Secondly, this application provides a terminal device. This terminal device can be used to implement the voiceprint template updating method described in the first aspect. The terminal device can be a smart robot, mobile phone, tablet computer, smart TV, router, in-vehicle system, watch, desktop computer, laptop computer, handheld computer, notebook computer, super mobile personal computer, netbook, as well as cellular phones, personal digital assistants, augmented reality / virtual reality devices, etc. This application does not limit the specific type of terminal device.
[0040] The functions of this terminal device can be implemented through hardware or through hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the aforementioned functions, such as a transceiver unit and a processing unit.
[0041] The transceiver unit can be used to send and receive information or data, or to communicate with other devices. The processing unit can be used to process the data. The transceiver unit and the processing unit can cooperate to implement the voiceprint template update method as described in the first aspect and any implementation thereof.
[0042] Thirdly, this application provides a terminal device. The terminal device includes: a processor, a camera, and a voice acquisition device; the processor is connected to the camera and the voice acquisition device respectively; the processor is used to invoke the camera and the voice acquisition device to implement the voiceprint template update method as described in the first aspect and any implementation thereof.
[0043] Fourthly, this application provides a terminal device. The terminal device includes: a processor; a memory for storing processor-executable instructions; the processor is configured to, when executing the instructions, cause the terminal device to implement the voiceprint template update method as described in the first aspect and any implementation thereof.
[0044] Fifthly, this application provides a computer-readable storage medium storing computer program instructions thereon; when the computer program instructions are executed by a terminal device, the terminal device implements the voiceprint template update method as described in the first aspect and any implementation thereof.
[0045] Sixthly, this application provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a terminal device, the processor in the terminal device implements the voiceprint template update method as described in the first aspect and any implementation thereof.
[0046] The technical effects corresponding to the second aspect and any implementation of the second aspect, the third aspect and any implementation of the third aspect, the fourth aspect and any implementation of the fourth aspect, the fifth aspect and any implementation of the fifth aspect, and the sixth aspect and any implementation of the sixth aspect can be referred to the technical effects corresponding to the first aspect and any implementation of the first aspect mentioned above, and will not be repeated here.
[0047] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0048] Figure 1 A schematic diagram illustrating an application scenario provided in an embodiment of this application;
[0049] Figure 2 This is a schematic diagram of the structure of an intelligent robot provided in an embodiment of this application;
[0050] Figure 3 A schematic diagram illustrating the process of updating the voiceprint template of the first user as provided in an embodiment of this application;
[0051] Figure 4 Another schematic diagram illustrating the process of updating the voiceprint template of the first user as provided in an embodiment of this application;
[0052] Figure 5 A schematic diagram illustrating the implementation logic of voiceprint registration provided in an embodiment of this application;
[0053] Figure 6 A schematic diagram illustrating the implementation logic of voiceprint verification provided in an embodiment of this application;
[0054] Figure 7 A schematic diagram illustrating the implementation logic of voiceprint template updating provided in this application embodiment;
[0055] Figure 8 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0056] The terminology used in the following embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to also include expressions such as “one or more,” unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, “at least one” and “one or more” refer to one or more (including two). The character “ / ” generally indicates that the preceding and following objects are in an “or” relationship.
[0057] References to "one embodiment" or "some embodiments" as used in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes both direct and indirect connections, unless otherwise stated.
[0058] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0059] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0060] In real life, everyone's voice has its own characteristics. Because the size, shape, and function of the human vocal organs differ, different people produce voices with different features. A voiceprint is one form of representation of these characteristics. For example, sound can be analyzed using electroacoustic instruments to display the sound wave spectrum carrying speech information. The acoustic characteristics (such as wavelength, frequency, intensity, and rhythm) of the sound wave spectrum corresponding to different sounds may differ; these acoustic characteristics of the sound wave spectrum can be called the voiceprint.
[0061] Voiceprint recognition technology is a technique that uses a speaker's voiceprint to identify the speaker's identity. It can be effectively applied in fields such as smart homes, smart buildings, and financial security.
[0062] Voiceprint recognition generally includes two parts: voiceprint registration and voiceprint verification. Voiceprint registration includes: acquiring the speaker's voice information, extracting corresponding voiceprint features based on the speaker's voice information, and generating (or registering) a voiceprint template corresponding to the speaker based on the extracted voiceprint features. Voiceprint verification includes: when a speaker's voice information is detected, extracting corresponding voiceprint features based on that speaker's voice information, and comparing the extracted voiceprint features with the voiceprint template generated during voiceprint registration to determine if the current speaker is the same person as the speaker at the time of voiceprint registration. For example, if the similarity between the extracted voiceprint features and the voiceprint template generated during voiceprint registration is greater than or equal to a preset threshold, it can be determined that the current speaker and the speaker at the time of voiceprint registration are the same person; otherwise, it can be determined that the current speaker and the speaker at the time of voiceprint registration are different.
[0063] However, in some possible scenarios, the speaker's voice during voiceprint registration may change and differ from the voice at the time of registration, which may lead to low accuracy in voiceprint recognition. For example, if the speaker's voice changes during voiceprint registration, voiceprint verification may not be able to accurately identify whether the current speaker is the same person as during voiceprint registration.
[0064] For example, scenarios where a speaker's voice may change include: the speaker catching a cold, or the speaker's voice changing with age. For instance, when a speaker catches a cold, the physiological characteristics of their oral cavity may change, leading to a hoarse or lowered voice. Similarly, when a speaker is a teenager, the development of their vocal organs as they age may also cause their voice to change. No limitations are placed on the scenarios in which a speaker's voice may change.
[0065] This application provides a voiceprint template update method, which may include: when someone is detected speaking, identifying whether the speaker is a user who has completed voiceprint registration (i.e., the user corresponding to an existing voiceprint template); when the speaker is a user who has completed voiceprint registration, updating the voiceprint template corresponding to the user according to the degree of change in the user's voice.
[0066] This method improves the accuracy of voiceprint recognition by updating the user's corresponding voiceprint template.
[0067] In some embodiments, the voiceprint template updating method provided in this application can be applied to a first device, which may have voiceprint recognition functionality. For example, the first device may include a voiceprint registration function and a voiceprint verification function. Based on the voiceprint registration function, the first device can receive a user's voice information, extract corresponding voiceprint features based on the user's voice information, and generate a voiceprint template corresponding to the user based on the extracted voiceprint features. That is, based on the voiceprint registration function, a user can register their own voiceprint template with the first device.
[0068] Based on the voiceprint verification function, the first device can extract the corresponding voiceprint features based on the voice information of a speaker when it detects the speaker's voice information, and compare the extracted voiceprint features with the voiceprint template generated in the voiceprint registration to determine whether the current speaker is the user corresponding to the voiceprint template in the voiceprint registration. In this application, the first device can also identify whether the speaker has completed voiceprint registration when it detects someone speaking; if the speaker has completed voiceprint registration, the device updates the voiceprint template corresponding to the user according to the degree of change in the user's voice.
[0069] For example, the first device may be a smart robot, mobile phone, tablet computer, smart TV, router, in-vehicle system, watch, desktop computer, laptop computer, handheld computer, laptop, super mobile personal computer, netbook, as well as cellular phone, personal digital assistant, augmented reality / virtual reality device and other terminal devices. This application does not limit the specific type of the first device.
[0070] For example, Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of this application. For example... Figure 1 As shown, a possible application scenario may include user A and intelligent robot 101. Intelligent robot 101 may include voiceprint registration function and voiceprint verification function.
[0071] Based on the voiceprint registration function, the intelligent robot 101 can receive the voice information of user A, extract the corresponding voiceprint features based on the voice information of user A, and generate the voiceprint template corresponding to user A based on the extracted voiceprint features.
[0072] Based on the voiceprint verification function, the intelligent robot 101 can extract the corresponding voiceprint features when it detects the voice information of a speaker, and compare the extracted voiceprint features with the voiceprint template generated in the voiceprint registration to determine if the current speaker is user A corresponding to the voiceprint template in the voiceprint registration. For example, when the intelligent robot 101 detects that a speaker is asking "How's the weather today?", it can identify whether the speaker is user A in the aforementioned manner. If the speaker is identified as user A, the intelligent robot 101 can reply "Today's weather is cloudy, 18 to 25 degrees Celsius". If the speaker is not identified as user A, the intelligent robot 101 can remain silent.
[0073] In this application, the intelligent robot 101 can also identify whether the speaker is a user A who has completed voiceprint registration when it detects someone speaking; when the speaker is a user A who has completed voiceprint registration, the robot updates the voiceprint template corresponding to user A according to the degree of change in user A's voice.
[0074] The following example uses the application of this application embodiment on an intelligent robot. First, the process of identifying whether the speaker is a user who has completed voiceprint registration when someone is detected speaking in the voiceprint template update method provided in this application embodiment will be described by way of example.
[0075] In this embodiment, when a user registers their voiceprint template with the intelligent robot, the intelligent robot can obtain and store the user's corresponding image (such as a facial image) as a preset user image. When someone is detected speaking, the intelligent robot's steps to identify whether the speaker is a user who has already completed voiceprint registration may include: when someone is detected speaking, acquiring surrounding image information (i.e., image information related to the speaker) and the speaker's voice information; determining the speaker based on the image information and voice information; and determining whether the speaker who emitted the voice information is a user who has already completed voiceprint registration based on the image corresponding to the speaker in the image information and the preset user image recognition (such as facial recognition). That is, the preset user image is the image corresponding to a user who has already completed voiceprint registration.
[0076] For example, Figure 2 This is a structural schematic diagram of an intelligent robot provided in an embodiment of this application. Figure 2 As shown, the intelligent robot may include at least: a chip 210, a camera 220, and a voice acquisition device 230, with the chip 210 connected to the camera 220 and the voice acquisition device 230 respectively.
[0077] The voice acquisition device 230, also known as a voice front-end device, may include audio devices with voice acquisition functions, such as sound sensors and microphones. The intelligent robot detecting someone speaking can mean that the voice acquisition device 230 has received sound data. The steps for the intelligent robot to acquire the speaker's voice information may include: the intelligent robot calling the voice acquisition device 230 to acquire the speaker's voice information.
[0078] The steps for the intelligent robot to acquire surrounding image information may include: the intelligent robot calling camera 220 to take pictures of the surrounding environment, obtaining multiple images or video stream data. For example, the intelligent robot may call camera 220 to take multiple pictures of the surrounding environment to obtain multiple images, or the intelligent robot may call camera 220 to continuously take pictures of the surrounding environment for a first duration to obtain video stream data.
[0079] Optionally, the first duration can be a preset duration, such as 2 minutes, 5 minutes, etc. Alternatively, the first duration can also be related to the duration of the speaker's voice information. For example, when someone is detected speaking, the intelligent robot can call the camera 220 to start recording, and when the speaker's voice stops, the intelligent robot controls the camera 220 to stop recording. This application does not limit the first duration.
[0080] Chip 210 can be used in intelligent robots to determine the speaker of the voice message based on image and sound information, and to identify whether the speaker of the voice message is a user who has completed voiceprint registration based on image information and a preset user image. For example, chip 210 can be pre-programmed with algorithms or programs that can achieve the aforementioned functions.
[0081] It should be noted that the above Figure 2 This is merely an exemplary illustration of the structure of an intelligent robot. In other examples, the intelligent robot may include... Figure 2 This may involve more or fewer components, or combining certain components, or splitting certain components, or different component arrangements. Figure 2 The components shown can be implemented in hardware, software, or a combination of both. This application does not limit the specific structure of the intelligent robot.
[0082] Additionally, when the first device is a mobile phone, tablet, smart TV, router, in-vehicle system, watch, desktop computer, laptop computer, handheld computer, notebook computer, super mobile personal computer, netbook, as well as other terminal devices such as cellular phones, personal digital assistants, augmented reality / virtual reality devices, etc., its structure can also refer to the above. Figure 2 As shown, this application also does not impose any limitations.
[0083] In some embodiments, the step of determining the speaker who issued the sound information based on image information and sound information may include: performing lip movement recognition based on image information, locating the sound source based on sound information, and determining the speaker who issued the sound information from the people included in the image.
[0084] Taking an image information system comprising multiple images as an example, where each image may contain one or more people, lip movement recognition on these images can identify which person's face showed lip movement. Sound source localization based on audio information can determine the location of the speaker. Combining the speaker's location with the individuals showing lip movement, the speaker can be identified. For example, if the speaker's location is position one, and user 1 at position one shows lip movement, then user 1 can be identified as the speaker.
[0085] Optionally, in one implementation, lip movement recognition can be performed based on image information to filter out people whose faces show lip movements from all people included in the image. Then, sound source localization can be performed based on sound information, and people whose lip movements match the sound source localization results can be selected as the speakers who give the sound information.
[0086] In another implementation, the sound source can be located first based on the sound information to determine the location of the speaker. Then, based on the image information, lip movement recognition can be performed on the people in the location of the speaker, and the people whose faces show lip movement can be identified as the speakers.
[0087] In other embodiments, the step of determining the speaker who made the sound information based on image information and sound information may also include: performing lip reading based on image information, performing speech recognition based on sound information, and determining the speaker who made the sound information from the people included in the image.
[0088] Taking multiple images as an example, each image may contain one or more people. Lip reading of these images can identify which person's face showed lip movement and the corresponding natural language statement. Speech recognition based on sound information can then determine the corresponding natural language statement. By comparing the natural language statement corresponding to the lip movement of the person with the corresponding natural language statement in the sound information, it can be determined which person's lip movement matches the natural language statement in the sound information; this person is then identified as the speaker who uttered the sound information.
[0089] For example, an intelligent robot may include a lip-reading model and a speech recognition model. When performing lip-reading on multiple images, the lip shape change features of each person who moved their lips can be extracted from the images first. Then, the lip shape change features of each person who moved their lips are input into the lip-reading model, which can output the sound data corresponding to the lip shape of each person who moved their lips. After inputting the sound data corresponding to the lip shape of each person who moved their lips into the speech recognition model, the speech recognition model can output the natural language sentence corresponding to the lip movements of each person who moved their lips. When performing speech recognition based on sound information, the sound information can be input into the speech recognition model, which can output the natural language sentence corresponding to the sound information.
[0090] In some embodiments, the step of determining the speaker who issued the sound information based on image information and sound information may further include: performing lip movement recognition and lip reading recognition based on image information, performing sound source localization and speech recognition based on sound information, and determining that the person included in the image is the speaker who issued the sound information when the person meets the following conditions 1 and / or conditions 2.
[0091] Condition 1: Lip movement occurred, and its location was consistent with the sound source localization result.
[0092] Condition 2: Lip movement occurred, and the natural language statement corresponding to the lip movement was the same as the natural language statement corresponding to the sound information.
[0093] The process of determining the speaker who emitted the sound information based on image and sound information can also be called the process of judging the consistency between image and sound information. When the image corresponding to a person is consistent with the sound information emitted by the speaker, the person is determined to be the speaker who emitted the sound information.
[0094] In some embodiments, the step of determining the speaker who issued the sound information based on image information and sound information may further include: performing lip movement recognition and lip reading recognition based on image information, performing sound source localization and speech recognition based on sound information, and performing interaction intention recognition based on image information. When a person included in the image meets one or more of the following conditions 1, 2, and 3, the person is determined to be the speaker who issued the sound information.
[0095] Conditions 1 and 2 are as described in the above embodiments and will not be repeated here.
[0096] Condition 3: The willingness to interact meets the preset requirements.
[0097] The step of recognizing interaction intention based on image information may include: determining a sequence of behavioral parameters for each person in the image; the sequence of behavioral parameters for each person may include one or more of the following: the person's facial angle, the distance between the person and the intelligent robot, and the person's actions; the person's actions may include facial movements, body movements, etc. For each person in the image, an interaction intention value is obtained based on the person's behavioral parameter sequence using a preset interaction intention value model. The interaction intention value can be used to represent the person's willingness to interact with the intelligent robot, and can also be referred to as the person's focus level when interacting with the intelligent robot. It should be understood that a larger interaction intention value indicates a stronger willingness to interact with the intelligent robot; a smaller interaction intention value indicates a weaker willingness to interact with the intelligent robot.
[0098] Optionally, the aforementioned interaction willingness value model can be obtained by training a neural network using a training dataset through machine learning. The training dataset can include one or more training data sets. Each training data set can include a sequence of behavioral parameters for a user and annotation information for that training data set. The annotation information can characterize the strength of the user's interaction willingness with the intelligent robot corresponding to that training data set. Different training data sets can correspond to the same or different users.
[0099] For example, in one possible implementation, the annotation information of the training data can be 0 or 1. When the annotation information is 0, it indicates that the user corresponding to the training data has the strongest willingness to interact with the intelligent robot. When the annotation information is 1, it indicates that the user corresponding to the training data has the weakest willingness to interact with the intelligent robot, such as no willingness to interact.
[0100] When training the model to obtain the interaction willingness value mentioned above, each piece of training data can be used as the input of the neural network, and the annotation information of each piece of training data can be used as the output of the neural network, so that the neural network can learn the mapping relationship between the user's behavior parameter sequence and the annotation information.
[0101] The above-mentioned step of obtaining the interaction willingness value of each person in the image based on the sequence of behavioral parameters of that person through a preset interaction willingness value model may include: for each person in the image, inputting the sequence of behavioral parameters of that person into a trained interaction willingness value model to obtain the interaction willingness value output by the interaction willingness value model.
[0102] For example, the interaction willingness value is related to the annotation information. For instance, the interaction willingness value can be a value between 0 and 1, such as 0.5, 0.8, etc.
[0103] In one possible implementation, the interaction willingness value in condition 3 above must meet a preset requirement, which may include: the interaction willingness value is the highest among all the interaction willingness values of the people included in the image. For example, if the image includes 3 people, and user 1's interaction willingness value is the highest among the 3 people, then user 1 meets condition 3.
[0104] In another possible implementation, the interaction willingness value in condition 3 above must meet a preset requirement, which may include: the interaction willingness value is greater than or equal to a preset interaction willingness threshold. For example, the interaction willingness threshold may be 0.7, 0.8, etc. When a person's interaction willingness value is greater than or equal to the interaction willingness threshold, it can be determined that the person meets condition 3.
[0105] Optionally, when there are multiple people whose interaction willingness value is greater than or equal to the interaction willingness threshold, the person with the highest interaction willingness value can be determined from among the multiple people as the person who meets condition 3.
[0106] In this embodiment, when a person included in the image simultaneously meets conditions 1, 2, and 3, that person is identified as the speaker emitting the sound information, ensuring that only that person is having a dialogue or interaction with the intelligent robot.
[0107] Although the above example illustrates the steps for determining the speaker of sound information based on image and sound information, it should be understood that when the image information is video stream data, the video stream data is also actually composed of multiple frames of images, and the principle is the same as the example above where the image information includes multiple images.
[0108] After determining the speaker by using the image and sound information in the above manner, the intelligent robot can identify whether the speaker is a user who has completed voiceprint registration based on the image of the speaker in the image information and the preset user image.
[0109] That is, in this embodiment of the application, the speaker's image can be filtered from the image information based on the image information and the speaker's voice information; and the speaker's image and the preset user image can be used to identify whether the speaker is a user who has completed voiceprint registration.
[0110] For example, after identifying the speaker, an image corresponding to the speaker can be selected or cropped from the image information. This image is then compared with a preset user image. If the similarity between the speaker's image and the preset user image is greater than a certain preset threshold, the speaker is determined to be a user who has already completed voiceprint registration. If the similarity between the speaker's image and the preset user image is less than the preset threshold, the speaker is determined not to be a user who has already completed voiceprint registration.
[0111] Optionally, when the similarity between the image corresponding to the speaker who sent the voice information and the preset user image is equal to the preset threshold, it can be determined that the speaker who sent the voice information is a user who has completed voiceprint registration, or it can be determined that the speaker who sent the voice information is not a user who has completed voiceprint registration; no restriction is imposed here.
[0112] The above embodiments use the application of the present application embodiments on intelligent robots as examples to illustrate the process of identifying whether the speaker is a user who has completed voiceprint registration when someone is detected speaking in the voiceprint template update method provided in the present application embodiments.
[0113] The following example uses the application of this application embodiment on an intelligent robot to continue explaining the process of updating the voiceprint template of the user according to the degree of change in the user's voice when the speaker is a user who has completed voiceprint registration in the voiceprint template update method provided in this application embodiment.
[0114] For example, the structure of an intelligent robot is as described above. Figure 2 As shown in the example, chip 210 can also be used in intelligent robots to update the voiceprint template corresponding to the user according to the degree of change in the user's voice.
[0115] Assuming the user who has already completed voiceprint registration is the first user, the steps for updating the voiceprint template corresponding to the first user based on the degree of change in the first user's voice when the speaker is the first user can include: determining whether the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user is greater than a voiceprint threshold; when the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user is greater than the voiceprint threshold, it can be considered that the first user's voice has undergone a gradual change, and the voiceprint template corresponding to the first user can be updated based on the voiceprint features corresponding to the first user's voice information; when the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user is less than or equal to the voiceprint threshold, it can be considered that the first user's voice has undergone a sudden change, and the voiceprint template corresponding to the first user can be re-registered.
[0116] For example, Figure 3 This is a schematic diagram illustrating the process of updating the voiceprint template of the first user as provided in an embodiment of this application.
[0117] like Figure 3 As shown, when the speaker is the first user, the step of updating the voiceprint template corresponding to the first user according to the degree of change in the first user's voice may include: S301-S305.
[0118] S301, Obtain the voice information of the first user.
[0119] For example, if the speaker is the first user, then the first user's voice information can be the voice information emitted by the speaker as described in the above embodiments. Alternatively, when the first user speaks again, the first user's voice information described in S301 can also be the re-acquired first user's voice information.
[0120] S302. Extract the voiceprint features corresponding to the voice information of the first user.
[0121] S303. Determine whether the similarity between the voiceprint feature corresponding to the voice information of the first user and the voiceprint template corresponding to the first user is greater than the voiceprint threshold.
[0122] If yes, that is, when the similarity between the voiceprint feature corresponding to the first user's voice information and the voiceprint template corresponding to the first user is greater than the voiceprint threshold, S304 can be executed; if no, that is, when the similarity between the voiceprint feature corresponding to the first user's voice information and the voiceprint template corresponding to the first user is less than or equal to the voiceprint threshold, S305 can be executed.
[0123] The voiceprint threshold is a preset value, such as one preset in the intelligent robot. Understandably, before S303, the intelligent robot can first calculate the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user.
[0124] For example, calculating the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user can include: obtaining a first score result based on the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user. The first score result can characterize or indicate the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user. For example, the higher the first score result, the higher the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user; the lower the first score result, the lower the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user.
[0125] In some embodiments, the step of obtaining a first scoring result based on the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user may include: according to the following Figure 6 or Figure 7 As shown, the voiceprint features corresponding to the voice information of the first user are input into the backend scoring module 404 to obtain the scoring result of the backend scoring module 404 scoring the voiceprint features corresponding to the voice information of the first user and the voiceprint template corresponding to the first user. This scoring result is the first score result.
[0126] In other embodiments, the step of obtaining a first score result based on the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user may include: calculating the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user using a preset similarity algorithm to obtain the first score result. The similarity algorithm may include distance metric algorithms (such as Euclidean distance, Minkowski distance, Mahalanobis distance, Manhattan distance, etc.) and similarity metric algorithms (such as vector space cosine similarity, Pearson correlation, etc.), and is not limited thereto.
[0127] In some possible implementations, the voiceprint threshold can satisfy the following conditions: precision is greater than or equal to a first precision threshold, recall is greater than or equal to a first recall threshold, and false entry rate is less than or equal to a first false entry rate threshold. For example, the first precision threshold can be 95%, the first recall threshold can be 95%, and the first false entry rate threshold can be 5%.
[0128] Precision, also known as accuracy, is the proportion of correctly predicted positive results out of all correctly predicted positive results. Precision can represent the accuracy of predictions for positive samples. For example, in this application, precision can refer to the probability that the intelligent robot can correctly identify the first user as a registrant (i.e., a user who has registered their voiceprint template).
[0129] For example, the accuracy can be calculated using the following formula (1).
[0130]
[0131] In formula (1), precision represents the accuracy rate, TP represents the number of times a user is correctly predicted to be a registrant, and FP represents the number of times an unregistered user is incorrectly identified as a registrant.
[0132] Recall, in relation to the original sample, represents the probability that a sample that is actually positive will be predicted as positive. In other words, recall is the proportion of correctly predicted positive samples out of all actually positive samples. A high recall means there may be more false positives, but it will strive to find every object that should be found. For example, in this application, recall could refer to the probability that the AI can recognize the first user as the registrant.
[0133] For example, recall rate can be calculated using the following formula (2).
[0134]
[0135] In formula (2), recall represents the recall rate, TP represents the number of times the user is correctly predicted to be a registrant, and FN represents the number of times the user is actually a registrant but is incorrectly predicted not to be a registrant.
[0136] The false entry rate refers to the probability of predicting a negative sample as a positive sample. For example, in this application, the false entry rate can refer to the probability that the intelligent robot incorrectly predicts that the first user, who is not a registrant (i.e., an unregistered user), is a registrant.
[0137] For example, the false entry rate can be calculated using the following formula (3).
[0138]
[0139] In formula (3), β represents the false entry rate, FP represents the number of times an unregistered user is mistakenly identified as a registered user, and TN represents the number of times a user is not a registered user and is predicted not to be a registered user.
[0140] Optionally, the aforementioned voiceprint threshold can be obtained by training with large amounts of data samples.
[0141] For example, the voiceprint threshold can be obtained by conducting experiments based on a large amount of experimental data.
[0142] For example, Table 1 below uses experimental data including 103 voice data from registered users and 95 voice data from unregistered users as an example to show 11 sets of experimental results corresponding to different voiceprint thresholds (only 11 sets are used as examples here), as well as the precision, recall, and false entry rate corresponding to each set of experimental results.
[0143] Table 1
[0144]
[0145] As shown in Table 1, in the first group of experiments, when the voiceprint threshold was 15.0, TP equaled 103, FP equaled 14, FN equaled 0, and TN equaled 81. Based on the results of the first group of experiments, the precision rate was 88.03%, the recall rate was 100.00%, and the false entry rate was 14.74%. In the second group of experiments, when the voiceprint threshold was 15.5, TP equaled 103, FP equaled 12, FN equaled 0, and TN equaled 83. Based on the results of the second group of experiments, the precision rate was 89.57%, the recall rate was 100.00%, and the false entry rate was 12.63%. ... and so on. In the eleventh group of experiments, when the voiceprint threshold was 20.0, TP equaled 100, FP equaled 4, FN equaled 3, and TN equaled 91. Based on the results of the eleventh group of experiments, the precision rate was 96.15%, the recall rate was 97.09%, and the false entry rate was 4.21%.
[0146] As shown in Table 1, different voiceprint thresholds affect accuracy, precision, and recall. In this embodiment, values such as 19.5 and 20.0, which satisfy precision greater than or equal to 95%, recall greater than or equal to 95%, and false entry rate less than or equal to 5%, can be selected as voiceprint thresholds. This application does not limit the specific size of the voiceprint threshold.
[0147] S304. Update the voiceprint template corresponding to the first user based on the voiceprint features corresponding to the first user's voice information.
[0148] In some embodiments, S304 may include: fusing the voiceprint features corresponding to the voice information of the first user and the voiceprint template corresponding to the first user to obtain a new voiceprint template; replacing the voiceprint template corresponding to the first user with the new voiceprint template, thereby updating the voiceprint template corresponding to the first user. That is, the updated voiceprint template corresponding to the first user will be changed to the aforementioned new voiceprint template.
[0149] Optionally, fusing the voiceprint features corresponding to the voice information of the first user and the voiceprint template corresponding to the first user to obtain a new voiceprint template may include: proportionally superimposing the voiceprint features corresponding to the voice information of the first user into the voiceprint template corresponding to the first user to obtain a new voiceprint template.
[0150] For example, in one possible implementation, the step of proportionally superimposing the voiceprint features corresponding to the voice information of the first user onto the voiceprint template corresponding to the first user may include: according to the following formula (4), proportionally superimposing the voiceprint features corresponding to the voice information of the first user onto the voiceprint template corresponding to the first user.
[0151]
[0152] In formula (4), μ′ k This represents a new voiceprint template; t represents the sum of the number of voiceprint features corresponding to the first user's voice information and the number of voiceprint templates corresponding to the first user, where the first user's voiceprint template is generally one, and the voiceprint features corresponding to the first user's voice information can be one or more; x t This refers to a specific voiceprint feature or voiceprint template (i.e., the voiceprint template already registered by the first user); when x t When representing voiceprint features, γ kt The value of x can be 1; when x t When representing a voiceprint template, γ kt The value can take any value within the range of 6 to 20, such as 10, 12, etc.; ∑ t γ kt x t This means taking γ k1 x1 to γ kt x t The sum; ∑ t γ kt This means taking γ k1 To γ kt The sum of.
[0153] That is, in this implementation, updating the voiceprint template corresponding to the first user based on the voiceprint features corresponding to the first user's voice information can include: obtaining the voiceprint template already registered by the first user and a preset first value (such as when x...). t When representing a voiceprint template, γkt The product of the values (which is the first value) yields the first product result (e.g., when x is the first value). t When representing a voiceprint template, γ kt x t That is, the result of the first product); obtain the voiceprint feature and the corresponding second value (when x t When representing voiceprint features, γ kt The product of the values of x and x is the second value, yielding the second product result (e.g., when x is the second value). t When representing voiceprint features, γ kt x t That is, the second product result); sum the first product result and the second product result to obtain the first summation result (as shown in formula (4) ∑ t γ kt x t Summing the first and second values yields the second summation result (as shown in formula (4)). t γ kt ); Obtain the ratio of the first summation result and the second summation result to obtain a new voiceprint template (i.e., μ′ in formula (4)). k Replace the voiceprint template already registered by the first user with a new voiceprint template.
[0154] For example, taking the voiceprint template corresponding to the first user as x1, the voiceprint features corresponding to the first user's voice information include x2. x1 and x2 can be the X vector output by the 6th layer of the 7-layer TDNN, as shown below. k1 and γ k2 Substituting into the above formula (4), the following equation (1) can be generated.
[0155]
[0156] In the above equation (1), μ′ k Z1 is an unknown, while all others are known. For example, z1 and z2 are known, and γ is an unknown. k1 It can be 12 (taking 12 as an example), γ k2 It can be 1. Solving equation (1) above will yield μ′. k This means that the voiceprint features corresponding to the voice information of the first user are superimposed on the voiceprint template corresponding to the first user in a proportional manner to obtain a new voiceprint template.
[0157] It should be noted that when x t When representing a voiceprint template, this application uses γ kt The value (i.e., the γ value corresponding to the voiceprint template) kt The size of ) is not limited, when the γ corresponding to the voiceprint template ktThe larger the γ value, the smaller the proportion of the voiceprint features corresponding to the first user's voice information in the new voiceprint template; when the γ value corresponding to the voiceprint template is larger... kt The smaller the value, the greater the proportion of the voiceprint features corresponding to the first user's voice information in the new voiceprint template. In practical applications, the γ value corresponding to the voiceprint template can be adjusted according to specific needs. kt Set to a larger or smaller value; no special restrictions are imposed here.
[0158] Optionally, this application specifies the γ corresponding to the voiceprint template. kt The range of values for is also not limited. For example, in some other possible implementations, the γ corresponding to the voiceprint template... kt The value range of γ may not be limited to the above 6 to 20, such as: the γ corresponding to the voiceprint template kt The value can also be any value in the range of 5 to 19, or any value in the range of 10 to 20, etc., and there is no restriction here.
[0159] S305. Based on the voiceprint features corresponding to the first user's voice information, re-register the voiceprint template corresponding to the first user.
[0160] For example, the step of re-registering the voiceprint template corresponding to the first user based on the voiceprint features corresponding to the first user's voice information can refer to the voiceprint registration process in voiceprint recognition. For example, S305 may include: generating a voiceprint template corresponding to the first user based on the voiceprint features corresponding to the first user's voice information, and deleting the voiceprint template already registered by the first user.
[0161] Optionally, the above Figure 3 In the illustrated embodiment, the first user's voice is considered to have undergone a sudden change when the similarity between the voiceprint feature corresponding to the first user's voice information and the voiceprint template corresponding to the first user is equal to the voiceprint threshold. In other embodiments, the first user's voice can also be considered to have undergone a gradual change when the similarity between the voiceprint feature corresponding to the first user's voice information and the voiceprint template corresponding to the first user is equal to the voiceprint threshold. The voiceprint template corresponding to the first user can be updated according to the voiceprint feature corresponding to the first user's voice information. No limitation is imposed here.
[0162] The voiceprint template updating method provided in this application update, when someone is detected speaking, identifies whether the speaker is a user who has completed voiceprint registration, and updates the voiceprint template corresponding to the user according to the degree of change in the user's voice when the speaker is a user who has completed voiceprint registration. This allows the voiceprint template corresponding to the user to be updated in a timely manner as the user's voice changes, and the accuracy of subsequent voiceprint recognition of the user using the user's voiceprint template can be higher.
[0163] In other embodiments, in formula (4) above, when x t When representing voiceprint features, γ kt The value can also be a value in the range of 0 to 1, such as 0.8. When x t When representing voiceprint features, γ kt The value can be specifically calculated based on the confidence levels corresponding to multiple (e.g., at least two) dimensions of biological information, as well as the influence weight of each of the aforementioned multiple dimensions of biological information. For example, multiple dimensions may include: visual dimension, voiceprint dimension, lip movement dimension, etc.
[0164] That is, when x t When representing voiceprint features, γ kt The value can be related to the confidence level of the first user's biometric information in at least two dimensions, and the influence weight of the biometric information in each of the at least two dimensions.
[0165] For example, before calculating the new voiceprint template according to formula (4), or, in formula (4), obtaining the voiceprint features and the corresponding γ of the voiceprint features. xt Before the product, the method further includes: obtaining the γ corresponding to the voiceprint feature based on the confidence level corresponding to the biometric information of the first user in at least two dimensions, and the influence weight corresponding to the biometric information in each of the at least two dimensions. kt .
[0166] Confidence level, also known as reliability, confidence level, or confidence coefficient, refers to the probability that the population parameter value falls within a certain range of the sample statistics.
[0167] Optionally, in the embodiments of this application, the confidence level corresponding to each dimension of biological information can be obtained by the corresponding artificial intelligence (AI) recognition model or AI recognition algorithm.
[0168] For example, the step described in the above embodiments of identifying whether the speaker emitting the voice information is a user who has completed voiceprint registration based on the image corresponding to the speaker emitting the voice information and the preset user image can include: using a facial recognition AI model to identify the image corresponding to the speaker emitting the voice information and the preset user image to determine whether the speaker emitting the voice information is a user who has completed voiceprint registration. When performing the identification, the facial recognition AI model can simultaneously output the confidence level corresponding to the visual dimension of biometric information, such as 95% or 96%.
[0169] Similarly, the confidence levels corresponding to other dimensions of biometric information such as voiceprints and lip movements can be obtained by the corresponding AI recognition models.
[0170] The influence weight of biometric information for each dimension can be the reciprocal of the error rate corresponding to that dimension's biometric information. For example, taking the visual dimension, the influence weight of biometric information in the visual dimension can be called the visual image weight, which is the reciprocal of the visual error rate. The visual error rate is the ratio of the number of incorrectly identified images to the total number of identified images. For example, assuming there are 1000 images corresponding to the speaker emitting the sound information, and when identifying the image corresponding to the speaker emitting the sound information and a preset user image, the number of incorrectly identified images is 100, then the visual error rate is 10%, and the reciprocal of the visual error rate is 10.
[0171] For example, when x t When representing voiceprint features, γ kt The value of λ is calculated by the following formula (5), which is to calculate λ based on the confidence level corresponding to the biological information of k dimensions (k is an integer greater than 1) and the influence weight corresponding to the biological information of each dimension in the k dimensions.
[0172]
[0173] In formula (5), a k ∑a represents the influence weight corresponding to the biological information in the k-th dimension; k This represents the summation of the influence weights corresponding to the k dimensions of biological information; This represents the confidence level corresponding to the biological information in the k-th dimension; This means first calculating the ratio between the influence weight and the confidence level corresponding to the biological information of each dimension, obtaining the ratio corresponding to the biological information of each dimension, and then summing the ratios corresponding to the biological information of k dimensions respectively.
[0174] Based on the derivation of formula (5), the value of λ can be expressed by the following formula (6).
[0175]
[0176] That is, the step of obtaining λ based on the confidence levels of the first user's biometric information in at least two dimensions and the influence weights of the biometric information in each of the at least two dimensions may include: summing the influence weights of the first user's biometric information in at least two dimensions to obtain a third summation result (such as ∑a in formula (6)). k ); Obtain the ratio between the influence weight and the confidence level of the biological information corresponding to each of the at least two dimensions, and obtain the ratio corresponding to the biological information of each dimension; Sum the ratios corresponding to the biological information of each of the at least two dimensions to obtain the fourth summation result (as shown in formula (6)). ); Obtain the ratio of the third summation result to the fourth summation result to get λ.
[0177] For example, in one possible implementation, λ can be calculated based on the confidence levels corresponding to the three dimensions of biological information, namely visual, voiceprint, and lip movement, and the influence weights corresponding to the biological information of each of the aforementioned three dimensions. That is, k can be equal to 3.
[0178] For example, suppose the influence weight corresponding to the visual dimension is a1, and the confidence level is CV. * The influence weight corresponding to the voiceprint dimension is a2, and the confidence level is VP. * The influence weight corresponding to the lip movement dimension is a3, and the confidence level is LP. * Then a1, CV * a2, VP * a3, and LP * Substituting into the above formula (5), we can obtain the following formula (7).
[0179]
[0180] Where a1 can be the reciprocal of the error rate in the visual dimension; a2 can be the reciprocal of the error rate in the voiceprint dimension; and a3 can be the reciprocal of the error rate in the lip movement dimension.
[0181] Alternatively, λ can be calculated based on the confidence levels of any two of the three dimensions of biometric information (visual, voiceprint, and lip movement) and the influence weights of each of the aforementioned two dimensions, i.e., k can be equal to 2. For example, assuming the influence weight of the visual dimension is 10 and the confidence level is 95%, and the influence weight of the voiceprint dimension is 5 and the confidence level is 85%, then λ can be calculated as 0.914 according to formula (5) or (6).
[0182] After calculating λ according to the above formula (5) or (6), λ can be used as the γ corresponding to the voiceprint feature. kt This allows the voiceprint features corresponding to the first user's voice information to be proportionally superimposed onto the voiceprint template corresponding to the first user. That is, when calculating the new voiceprint template according to formula (4), when x... t When representing voiceprint features, γ kt The value of can be λ. λ can be a value in the range of 0 to 1, that is, λ is greater than or equal to 0 and less than or equal to 1.
[0183] Taking the voiceprint template corresponding to the first user as x1, the voiceprint features corresponding to the first user's voice information include x2, γ k1 12, γ k2Taking λ as an example (i.e., λ = 0.8), let x1, x2, γ k1 and γ k2 Substituting into the above formula (4), the following equation (2) can be generated.
[0184]
[0185] Solving equation (2) above yields μ′. k This means that the voiceprint features corresponding to the voice information of the first user are superimposed on the voiceprint template corresponding to the first user in a proportional manner to obtain a new voiceprint template.
[0186] In this embodiment, when x t When representing voiceprint features, γ kt The value is calculated based on the confidence level corresponding to multiple (e.g., at least two) dimensions of biological information, as well as the influence weight of each dimension of biological information. This can better optimize the proportion of voiceprint template updates, so as to achieve the purpose of updating the voiceprint template based on multi-dimensional fusion information, more objectively reflect the accuracy of each dimension, and further improve the accuracy of voiceprint recognition.
[0187] The above embodiments illustrate the process of updating the voiceprint template of the first user according to the degree of change in the first user's voice in the voiceprint template update method provided in this application. The first user's voice is considered to have changed gradually when the similarity (i.e., the first scoring result) between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user is greater than the voiceprint threshold. The first user's voice is then updated based on the voiceprint features corresponding to the first user's voice information. Conversely, the first user's voice is considered to have changed abruptly when the similarity (i.e., the first scoring result) between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user is less than the voiceprint threshold.
[0188] In other embodiments, the voiceprint threshold may include two, such as a first voiceprint threshold and a second voiceprint threshold, wherein the first voiceprint threshold is greater than the second voiceprint threshold. The process of updating the voiceprint template corresponding to the first user based on the degree of change in the first user's voice may include: obtaining a first score result based on the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user. When the first score result is greater than the first voiceprint threshold, it is considered that the first user's voice has undergone a gradual change, and the voiceprint template corresponding to the first user is updated based on the voiceprint features corresponding to the first user's voice information; when the first score result is less than the second voiceprint threshold, it is considered that the first user's voice has undergone a sudden change, and the voiceprint template corresponding to the first user is re-registered; when the first score result is greater than the second voiceprint threshold but less than the first voiceprint threshold, it is considered that it is impossible to distinguish whether the first user's voice has undergone a sudden change or a gradual change, and no processing is performed, i.e., the voiceprint template corresponding to the first user remains unchanged.
[0189] Similar to the voiceprint threshold described in the previous embodiments, the first voiceprint threshold and the second voiceprint threshold can also be preset values, such as those preset in intelligent robots.
[0190] In some possible implementations, the first and second voiceprint thresholds can also satisfy the following conditions: precision is greater than or equal to the second precision threshold, recall is greater than or equal to the second recall threshold, and false entry rate is less than or equal to the second false entry rate threshold. The second precision threshold can be the same as or different from the first precision threshold; the second recall threshold can be the same as or different from the first recall threshold; and the second false entry rate threshold can be the same as or different from the first false entry rate threshold. For example, the second precision threshold can be 95%, the second recall threshold can be 95%, and the second false entry rate threshold can be 5%.
[0191] Optionally, the determination methods of the second precision threshold, the second recall threshold, and the second false entry threshold can be the same as or similar to the determination methods of the first precision threshold, the first recall threshold, and the first false entry threshold in the foregoing embodiments, and will not be repeated here.
[0192] For example, Figure 4 This is another schematic diagram illustrating the process of updating the voiceprint template of the first user, as provided in an embodiment of this application.
[0193] like Figure 4 As shown, when the speaker is the first user, the steps of updating the voiceprint template corresponding to the first user according to the degree of change in the first user's voice may include: S401-S407.
[0194] S401, Obtain the voice information of the first user.
[0195] S402. Extract the voiceprint features corresponding to the voice information of the first user.
[0196] S401-S402 can be referred to in the above description of S301-S302, and will not be repeated here.
[0197] S403. Determine whether the similarity between the voiceprint feature corresponding to the voice information of the first user and the voiceprint template corresponding to the first user is greater than the first voiceprint threshold.
[0198] If yes, that is, when the similarity between the voiceprint feature corresponding to the first user's voice information and the voiceprint template corresponding to the first user is greater than the first voiceprint threshold, S404 can be executed; if no, that is, when the similarity between the voiceprint feature corresponding to the first user's voice information and the voiceprint template corresponding to the first user is less than or equal to the first voiceprint threshold, S405 can be executed.
[0199] Understandably, before S403, the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user can be calculated first. The process of calculating the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user can refer to the previous embodiment. For example, the first scoring result can be obtained based on the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user, which will not be elaborated here.
[0200] S404. Update the voiceprint template corresponding to the first user based on the voiceprint features corresponding to the first user's voice information.
[0201] In S404, the process of updating the voiceprint template corresponding to the first user based on the voiceprint features corresponding to the first user's voice information can be referred to the previous embodiment and will not be repeated here.
[0202] S405. Determine whether the similarity between the voiceprint feature corresponding to the voice information of the first user and the voiceprint template corresponding to the first user is less than the second voiceprint threshold.
[0203] As mentioned above, the second voiceprint threshold is less than the first voiceprint threshold.
[0204] If yes, that is, when the similarity between the voiceprint feature corresponding to the first user's voice information and the voiceprint template corresponding to the first user is less than the second voiceprint threshold, S406 can be executed; if no, that is, when the similarity between the voiceprint feature corresponding to the first user's voice information and the voiceprint template corresponding to the first user is greater than or equal to the second voiceprint threshold and less than or equal to the first voiceprint threshold, S407 can be executed, or no step can be executed.
[0205] S406. Based on the voiceprint features corresponding to the first user's voice information, re-register the voiceprint template corresponding to the first user.
[0206] The process of re-registering the voiceprint template corresponding to the first user based on the voiceprint features corresponding to the first user's voice information in S406 can be referred to the above-described S305, and will not be repeated here.
[0207] S407. Keep the voiceprint template corresponding to the first user unchanged.
[0208] S407 is equivalent to not performing any steps.
[0209] Optionally, when the first score result equals the second voiceprint threshold, it can be considered that the first user's voice has undergone a sudden change, and the voiceprint template corresponding to the first user is re-registered. Alternatively, it can be considered that the sudden change or gradual change in the first user's voice cannot be distinguished, and the voiceprint template corresponding to the first user remains unchanged. When the first score result equals the first voiceprint threshold, it can be considered that the first user's voice has undergone a gradual change, and the voiceprint template corresponding to the first user is updated according to the voiceprint features corresponding to the first user's voice information. Alternatively, it can be considered that the sudden change or gradual change in the first user's voice cannot be distinguished, and the voiceprint template corresponding to the first user remains unchanged. This application does not limit the cases where the first score result equals the first voiceprint threshold or the second voiceprint threshold.
[0210] In this embodiment, when updating the voiceprint template corresponding to the user based on the degree of change in the user's voice, the distinction between gradual and abrupt changes is made using two thresholds: a first voiceprint threshold and a second voiceprint threshold. The first voiceprint threshold is greater than the second voiceprint threshold. If the similarity between the voiceprint features corresponding to the first user's voice information and the corresponding voiceprint template is between the first and second voiceprint thresholds, it is considered a scenario where abrupt and gradual changes cannot be distinguished, and no further processing is performed. This eliminates some cases of ambiguous recognition, improves the effectiveness of voiceprint template updates, and further enhances the accuracy of voiceprint recognition.
[0211] For example, in one possible implementation, the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user can be between -100 and 100 (related to the scoring result). The first voiceprint threshold can be set to -10, and the second voiceprint threshold can be set to 15. In this embodiment, when the first scoring result (i.e., the similarity between the voiceprint features corresponding to the first user's voice information and the voiceprint template corresponding to the first user) is less than -10, it can be considered a sudden change scenario, and the voiceprint template of the first user is updated; when the first scoring result is greater than 15, it can be considered a gradual change scenario, and the voiceprint template of the first user is re-registered; when the first scoring result is between -10 and 15, it can be considered that the sudden change scenario or the gradual change scenario cannot be distinguished, and no processing is performed.
[0212] Compared to this embodiment, the case described in the previous embodiment where only one voiceprint threshold is used for judgment can also be considered as the case in this embodiment where the first voiceprint threshold and the second voiceprint threshold are the same, that is, Figure 4 In the illustrated embodiment, when the first voiceprint threshold and the second voiceprint threshold are the same, it is equivalent to Figure 3 The situation is illustrated in the embodiments shown.
[0213] The following is combined with Figures 5 to 7 The following is an exemplary description of some implementation logic involved in the embodiments of this application.
[0214] Figure 5 This is a schematic diagram illustrating the implementation logic of voiceprint registration provided in an embodiment of this application. Figure 5 As shown, in some possible examples, the intelligent robot may include: a mel-scale frequency cepstral coefficients (MFCC) feature extraction module 501, a neural network (NN) model 502, and a voiceprint template module 503.
[0215] The MFCC feature extraction module 501 can be used to extract MFCC features from sound information. MFCC features refer to the cepstral parameters extracted in the Mel-scale frequency domain. The Mel scale describes the nonlinear characteristics of human ear frequencies.
[0216] The NN model 502 has the function of outputting the voiceprint features corresponding to the sound information based on the MFCC features. For example, the voiceprint features can be X-vector, I-vector, etc.
[0217] The voiceprint template module 503 can be used to store the voiceprint templates that users register with the intelligent robot.
[0218] Please continue to refer to this. Figure 5 As shown, in this example, the voiceprint registration process may include: inputting cached speech data into the MFCC feature extraction module 501, where the cached speech data may be the speaker's voice information as described in the previous embodiment; the MFCC feature extraction module 501 extracts features from the cached speech data and outputs the MFCC features in the cached speech data to the NN model 502; the NN model 502 processes the received MFCC features and outputs the X vector corresponding to the cached speech data to the voiceprint template module 503; and the voiceprint template module 503 saves the received X vector as the voiceprint template for user registration.
[0219] Optionally, after extracting the MFCC features, the MFCC feature extraction module 501 can preprocess the MFCC features to obtain preprocessed MFCC features, and then output the preprocessed MFCC features to the NN model 502. Preprocessing the MFCC features may include: cepstral mean and variance normalization (CMVN) and voice activity detection (VAD) processing. Voice activity detection can identify and eliminate long silence periods from the audio signal stream of buffered voice data, thus saving voice channel resources without degrading service quality.
[0220] Optionally, the MFCC feature extraction module 501 may include an audio feature extractor, such as a mel filter, through which the MFCC feature extraction module 501 can extract MFCC features.
[0221] Optionally, the NN model 502 can be a model based on a 7-layer time delay neurnos network (TDNN), and the NN model 502 can use the output of the 6th layer of the TDNN as the X vector mentioned above.
[0222] Figure 6 This is a schematic diagram illustrating the implementation logic of voiceprint verification provided in an embodiment of this application. Figure 6 As shown above, in the above Figure 5 Based on the above, the intelligent robot may also include: a backend scoring module 504 and a verification module 505.
[0223] The voiceprint verification process may include: obtaining the X vector corresponding to the speaker's voice information (such as cached speech data) through the MFCC feature extraction module 501 and the NN model 502, the specific principle of which is as described in the above embodiment and will not be repeated here. The NN model 502 outputs the X vector corresponding to the speaker's voice information to the backend scoring module 504. The backend scoring module 504 scores the speaker based on the X vector corresponding to the speaker's voice information and the voiceprint template in the voiceprint template module 503, and outputs the scoring result to the verification module 505. The scoring result can be a specific score. The verification module 505 can determine whether the speaker who emitted the voice information is the user corresponding to the voiceprint template based on the scoring result. If the scoring result is greater than a certain threshold (such as the verification threshold or recognition threshold), it can be determined that the speaker who emitted the voice information is the user corresponding to the voiceprint template; otherwise, it is considered that the speaker who emitted the voice information is not the user corresponding to the voiceprint template.
[0224] Optionally, the backend scoring module 504 may include: a linear dirichlet allocation (LDA) algorithm and a probabilistic linear discriminant analysis (PLDA) algorithm. Among them,
[0225] Linear Discriminant Analysis (LDA) is a commonly used dimensionality reduction method in pattern recognition. LDA uses label information to find the optimal projection direction, minimizing intra-class variance and maximizing inter-class variance in the projected sample set. When applied to speaker identification, the I-Vector of the same speaker represents a class. Minimizing intra-class variance reduces channel-induced variations, while maximizing inter-class variance increases the difference information between speakers.
[0226] Probabilistic Linear Discriminant Analysis (PLDA) is also a channel compensation algorithm, also known as the probabilistic form of LDA. PLDA is typically based on I-Vector features, providing channel compensation for them.
[0227] The backend scoring module 504 can reduce the dimensionality of the X vector and voiceprint template corresponding to the speaker's voice information through a linear discriminant analysis algorithm, and score the dimensionality-reduced X vector through a probabilistic linear discriminant analysis algorithm.
[0228] Figure 7 This is a schematic diagram illustrating the implementation logic of voiceprint template updating provided in an embodiment of this application. Figure 7 As shown above, in the above Figure 6 Based on the above, the intelligent robot may also include: a voiceprint template update module 506.
[0229] The voiceprint template update process may include: obtaining the X vector corresponding to the speaker's voice information (such as cached speech data) through the MFCC feature extraction module 501 and the NN model 502, the specific principle of which is as described in the above embodiment and will not be repeated here. The NN model 502 outputs the X vector corresponding to the speaker's voice information to the backend scoring module 504. The backend scoring module 504 scores the speaker based on the X vector corresponding to the speaker's voice information and the voiceprint template in the voiceprint template module 503, and outputs the scoring result and voiceprint confidence (i.e., the confidence corresponding to the biometric information in the voiceprint dimension) to the voiceprint template update module 506. The voiceprint template update module 506 can determine whether the scoring result is greater than a preset voiceprint threshold. When the scoring result is greater than the preset voiceprint threshold, the voiceprint template update module 506 can calculate a fusion confidence (i.e., λ as described in the above embodiment) based on the voiceprint confidence, visual confidence (i.e., the confidence corresponding to the biological information in the visual dimension), and lip movement confidence (i.e., the confidence corresponding to the biological information in the lip movement dimension), and update the voiceprint template based on the fusion confidence and the X vector output by the NN model 502.
[0230] The above Figures 5 to 7 This is merely one possible implementation logic for intelligent robots and is not intended to limit it.
[0231] It should be understood that the above embodiments can also be applied to other types of first devices. Furthermore, in the embodiments of this application, the user registering the voiceprint template may include one or more users, and each user's voiceprint template may include one or more pieces of that user's voice information; no limitations are imposed here.
[0232] Corresponding to the voiceprint template updating method described in the foregoing embodiments, this application also provides a terminal device that can be used to implement the aforementioned voiceprint template updating method. This terminal device can be a smart robot, mobile phone, tablet computer, smart TV, router, in-vehicle system, watch, desktop computer, laptop computer, handheld computer, notebook computer, super mobile personal computer, netbook, as well as cellular phones, personal digital assistants, augmented reality / virtual reality devices, etc. This application does not limit the specific type of terminal device.
[0233] The functions of this terminal device can be implemented through hardware or by executing corresponding software within the hardware. The hardware or software includes one or more modules or units corresponding to the aforementioned functions.
[0234] For example, Figure 8 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 8 As shown, the terminal device may include a transceiver unit 801 and a processing unit 802.
[0235] The transceiver unit 801 can be used to send and receive information or data, or to communicate with other devices. For example, the transceiver unit 801 may include the camera 220 and the voice acquisition device 230 described in the above embodiments, or it may be connected to the camera 220 and the voice acquisition device 230 described in the above embodiments. The transceiver unit 801 can receive the user's image information and the user's voice data / sound information.
[0236] The processing unit 802 can be used to process data. For example, the processing unit 802 may include the chip 210 described in the above embodiments.
[0237] The transceiver unit 801 and the processing unit 802 can work together to implement the voiceprint template update method described in the embodiments of this application.
[0238] For example, when the terminal device is an intelligent robot (or the first device mentioned above), the transceiver unit 801 can be used to execute S301, S401, etc., as described in the foregoing embodiments. The processing unit 802 can be used to execute S302-S305, S402-S407, etc., as described in the foregoing embodiments.
[0239] The above device (i.e.) should be understood Figure 8 The division of units or modules (hereinafter referred to as units) in the terminal device shown is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, all units in the device can be implemented entirely in software through processing element calls; all units can be implemented entirely in hardware; or some units can be implemented in software through processing element calls, while others can be implemented in hardware.
[0240] For example, each unit can be a separate processing element, or it can be integrated into a chip within the device. Alternatively, it can be stored as a program in memory, invoked and executed by a processing element within the device. Furthermore, these units can be integrated in whole or in part, or implemented independently. The processing element described here can also be called a processor, which can be an integrated circuit with signal processing capabilities. In implementation, each step of the above method or each of the above units can be implemented through integrated logic circuits in the processor element or through software invoked by the processing element.
[0241] In one example, the unit in the above device may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.
[0242] For example, when the units in the device can be implemented through a processing element scheduler, the processing element can be a general-purpose processor, such as a CPU or other processor capable of calling programs. Alternatively, these units can be integrated together to form a system-on-a-chip (SOC).
[0243] In one implementation, the units that implement the corresponding steps in the above method can be implemented in the form of a processing element scheduler. For example, the device may include a processing element and a storage element, wherein the processing element calls a program stored in the storage element to execute the voiceprint template update method described in the above method embodiments. The storage element may be a storage element located on the same chip as the processing element, i.e., an on-chip storage element.
[0244] In another implementation, the program used to perform the above-described voiceprint template update method can be located on a storage element on a different chip than the processing element, i.e., an off-chip storage element. In this case, the processing element calls or loads the program from the off-chip storage element onto the on-chip storage element to call and execute the voiceprint template update method described in the above method embodiments.
[0245] For example, embodiments of this application may also provide a terminal device, which may include: a processor and a memory for storing processor-executable instructions. When the processor is configured to execute the aforementioned instructions, the terminal device implements the voiceprint template update method as described in the foregoing embodiments.
[0246] The terminal device can be a smart robot, mobile phone, tablet computer, smart TV, router, in-vehicle system, watch, desktop computer, laptop computer, handheld computer, notebook computer, super mobile personal computer, netbook, as well as cellular phone, personal digital assistant, augmented reality / virtual reality device, etc. This application does not limit the specific type of terminal device. The memory can be located inside or outside the terminal device, and the processor can include one or more.
[0247] In another implementation, the unit implementing each step of the above voiceprint template update method can be configured as one or more processing elements. These processing elements can be integrated circuits, such as one or more ASICs, one or more DSPs, one or more FPGAs, or combinations of these types of integrated circuits. These integrated circuits can be integrated together to form a chip.
[0248] For example, this application also provides a chip that can be applied to the aforementioned terminal device. The chip includes one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the processor receives and executes computer instructions from the terminal device's memory through the interface circuits to implement the voiceprint template updating method described in the above method embodiments.
[0249] For example, the chip may be chip 210 in the first device described above.
[0250] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0251] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0252] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0253] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0254] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product, such as a program. This software product is stored in a program product, such as a computer-readable storage medium, and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the voiceprint template updating method described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0255] For example, embodiments of this application may also provide a computer-readable storage medium storing computer program instructions. When the computer program instructions are executed by a terminal device, the terminal device implements the voiceprint template updating method as described in the foregoing method embodiments.
[0256] For example, embodiments of this application also provide a computer program product, including: computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a terminal device, the processor in the terminal device implements the voiceprint template update method as described in the above method embodiments.
[0257] Based on the foregoing embodiments Figure 2 In addition to the structure of the intelligent robot shown, this application also provides a terminal device, which may be an intelligent robot, mobile phone, tablet computer, smart TV, router, in-vehicle system, watch, desktop computer, laptop computer, handheld computer, notebook computer, super mobile personal computer, netbook, as well as cellular phone, personal digital assistant, augmented reality / virtual reality device, etc. This application does not limit the specific type of terminal device.
[0258] The terminal device may include a processor, a camera, and a voice acquisition device; the processor is connected to both the camera and the voice acquisition device; the processor is used to invoke the camera and the voice acquisition device to implement the voiceprint template update method as described in the above method embodiments. For example, the processor may... Figure 2 The chip 210 shown is shown.
[0259] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for updating a voiceprint template, characterized in that, The method includes: Obtain the voice information of the first user; the first user is a user who has registered a voiceprint template. Extract the voiceprint features corresponding to the sound information; Based on the voiceprint features and the voiceprint template registered by the first user, a first scoring result is obtained; When the first score is greater than the first voiceprint threshold, the voiceprint template registered by the first user is updated according to the voiceprint features; wherein, the update is implemented based on the voiceprint features and relevant parameters of at least two dimensions of the first user's biometric information; Alternatively, when the first score is less than the second voiceprint threshold, the voiceprint template of the first user is re-registered based on the voiceprint features; the second voiceprint threshold is less than or equal to the first voiceprint threshold.
2. The method according to claim 1, characterized in that, The step of updating the voiceprint template registered by the first user based on the voiceprint features includes: Obtain the product between the voiceprint template registered by the first user and the preset first value to get the first product result; Obtain the product between the voiceprint feature and the second value corresponding to the voiceprint feature to get the second product result; Summing the first product result and the second product result yields the first summation result; Summing the first value and the second value yields a second summation result; Obtain the ratio of the first summation result to the second summation result to get a new voiceprint template; Replace the voiceprint template already registered by the first user with the new voiceprint template.
3. The method according to claim 2, characterized in that, The second value corresponding to the voiceprint feature is related to the confidence level of the first user's biometric information in at least two dimensions, and the influence weight of the biometric information in each of the at least two dimensions.
4. The method according to claim 3, characterized in that, Before obtaining the product between the voiceprint feature and the second value corresponding to the voiceprint feature, the method further includes: The second value corresponding to the voiceprint feature is obtained based on the confidence level of the first user's biometric information in at least two dimensions, and the influence weight of the biometric information in each of the at least two dimensions.
5. The method according to claim 4, characterized in that, The step of obtaining the second value corresponding to the voiceprint feature based on the confidence level corresponding to the first user's biometric information in at least two dimensions, and the influence weight corresponding to the biometric information in each of the at least two dimensions, includes: The influence weights corresponding to the biometric information of the first user in at least two dimensions are summed to obtain a third summation result. Obtain the ratio between the influence weight and the confidence level of the biological information corresponding to each of the at least two dimensions, and obtain the ratio corresponding to the biological information of each dimension; The ratios corresponding to the biological information in each of the at least two dimensions are summed to obtain a fourth summation result; The ratio of the third summation result to the fourth summation result is obtained to obtain the second value corresponding to the voiceprint feature.
6. The method according to any one of claims 1-5, characterized in that, The first voiceprint threshold and the second voiceprint threshold satisfy the following conditions: precision is greater than or equal to the first precision threshold, recall is greater than or equal to the first recall threshold, and false entry rate is less than or equal to the first false entry rate threshold.
7. The method according to any one of claims 1-5, characterized in that, The acquisition of the first user's voice information includes: When someone is detected speaking, image information and the speaker's voice information are acquired; the image information is related to the speaker. Based on the image information and the speaker's voice information, the speaker's image is selected from the image information; Based on the speaker's image and a preset user image, it is determined whether the speaker is the first user who has completed voiceprint registration; When the speaker is the first user who has completed voiceprint registration, the voice information of the first user is obtained.
8. The method according to claim 7, characterized in that, The step of filtering out the speaker's image from the image information based on the image information and the speaker's voice information includes: Based on the image information, lip movement recognition is performed to identify the individuals who made lip movements from the individuals included in the image information; Based on the sound information, the sound source is located to obtain the first location of the speaker; Based on the first position, the speaker is identified from the individuals exhibiting lip movements, and an image of the speaker is acquired.
9. The method according to claim 7, characterized in that, The step of filtering out the speaker's image from the image information based on the image information and the speaker's voice information includes: Lip reading is performed based on the image information to identify individuals who have lip movements from the individuals included in the image information, and the corresponding natural language statements for each individual's lip movements. Speech recognition is performed based on the sound information to obtain the natural language statement corresponding to the sound information; Based on the natural language statement corresponding to the sound information and the natural language statement corresponding to the lip movements of each person who made lip movements, the speaker is identified from the persons who made lip movements, and an image of the speaker is obtained.
10. The method according to claim 7, characterized in that, The step of filtering out the speaker's image from the image information based on the image information and the speaker's voice information includes: Based on the image information, lip movement recognition and lip reading recognition are performed; based on the sound information, sound source localization and speech recognition are performed; and based on the image information, interaction intention recognition is performed. From the people included in the image information, one or more of the following conditions 1, 2, and 3 are identified as the speaker, and the image of the speaker is obtained. Condition 1: Lip movement occurred, and its location was consistent with the sound source localization result; Condition 2: Lip movement occurred, and the natural language statement corresponding to the lip movement is the same as the natural language statement corresponding to the sound information; Condition 3: The willingness to interact meets the preset requirements.
11. A terminal device, characterized in that, include: Processor, camera, and voice capture device; The processor is connected to both the camera and the voice acquisition device. The processor is used to invoke the camera and the voice acquisition device to implement the method as described in any one of claims 1-10.
12. A terminal device, characterized in that, include: A processor, and a memory for storing processor-executable instructions; When the processor is configured to execute the instructions, it causes the terminal device to implement the method as described in any one of claims 1-10.
13. A computer-readable storage medium having computer program instructions stored thereon; characterized in that, When the computer program instructions are executed by the terminal device, the terminal device performs the method as described in any one of claims 1-10.
14. A computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, characterized in that, When the computer-readable code is run in a terminal device, the processor in the terminal device implements the method as described in any one of claims 1-10.
Citation Information
Patent Citations
Man-machine interaction method and device, electronic equipment and storage medium
CN110689889A
Voiceprint recognition method and device
CN112289322A
Automatic registration voiceprint recognition method and device
CN113241080A
Method and apparatus for processing voice data
US20180144742A1