Methods, devices, and terminal devices for identity recognition

Optimizing voiceprint recognition by pre-training feature extraction models and end-to-end loss functions, the problem of high model complexity in the existing technology is solved, and the effect of simplifying training and improving recognition capabilities is achieved.

CN115148213BActive Publication Date: 2025-07-25ALIBABA INNOVATION PRIVATE LIMITED
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110343618.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-30
Publication Date
2025-07-25
Estimated Expiration
2041-03-30

AI Technical Summary

Technical Problem

In the prior art, the voiceprint recognition model is too complex and has not been synchronized, resulting in high system complexity and increased voiceprint comparison complexity.

Method used

The pre-trained feature extraction model is used to process user voice information through the self-attention mechanism and linear layer, and the voiceprint, age and gender characteristics are trained using the end-to-end loss function to achieve parallel optimization.

Benefits of technology

Simplify the model training process, improve the generalization ability of voiceprints, and can simultaneously identify the user's gender and age, reducing system complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115148213B_ABST
    Figure CN115148213B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and terminal device for identity recognition. Among them, the method includes: obtaining user voice information; parsing the user voice through a pre-trained feature extraction model to obtain the voiceprint feature of each sentence in the user voice information; calculating the similarity between the features in the user voice information according to the similarity calculation to obtain the age and gender of the user; generating the identity recognition information of the user based on the voiceprint feature, age and gender. The present invention solves the technical problem that the model used in the related art is too complex and not optimized synchronously.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technologies, and in particular, to a method, apparatus, and terminal device for identity recognition. Background Art

[0002] Voiceprint recognition technology is an important means for smart speakers and smart IoT to perceive user identity and attribute information. In the field of voiceprint recognition, the traditional approach is to use a two-step method to train voiceprint features and a scorer, and these two parts have their own optimization objectives and methods. This approach increases the complexity of the system on the one hand. When deploying the service, two parts of the model need to be deployed, and the scoring is complex during voiceprint comparison.

[0003] On the other hand, since the optimization methods and objectives of the two parts of training voiceprint features and the scorer are inconsistent, it is difficult to obtain the optimal effect.

[0004] In view of the problem that the model used in the related art is too complex and there is no synchronous optimization, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present invention provide a method, apparatus, and terminal device for identity recognition to at least solve the technical problem that the model used in the related art is too complex and there is no synchronous optimization.

[0006] According to one aspect of the embodiments of the present invention, a method for identity recognition is provided, including: obtaining user voice information; parsing the user voice through a pre-trained feature extraction model to obtain the voiceprint features of each sentence in the user voice information; calculating the similarity between the features in the user voice information according to similarity calculation to obtain the age and gender of the user; generating identity recognition information of the user based on the voiceprint features, age, and gender.

[0007] Optionally, the pre-trained feature extraction model includes: setting a preset number of encoding layers for each convolutional layer; synthesizing all frame features of the input voice passing through the convolutional network through the added self-attention mechanism on the encoding layer to obtain the synthesized features; normalizing the synthesized features after passing through the linear layer to obtain a first loss function and a second loss function; training according to the first loss function and the second loss function to obtain the trained feature extraction model, where the first loss function is used to extract the voiceprint features in the user voice information, and the second loss function is used to extract the age and gender in the user voice information.

[0008] Optionally, before setting a preset number of encoding layers for each convolutional layer, the method further includes: configuring the convolutional neural network, where configuring the convolutional neural network includes: setting the step size of each convolutional layer to a preset value.

[0009] Optionally, training is performed based on the first loss function and the second loss function, and the obtained trained feature extraction model includes: performing weighted summation on the first loss function and the second loss function to obtain a loss function to be trained; performing sample training based on the loss function to be trained, and correcting the weights of the first loss function and the second loss function to obtain the trained feature extraction model.

[0010] Optionally, obtaining user voice information includes: collecting the user's voice information through a client; obtaining the voice signal spectrum based on the voice information through spectrum transformation; obtaining acoustic features based on the voice signal spectrum through acoustic processing; and determining the acoustic features as the user voice information.

[0011] Optionally, calculating the similarity between each feature in the user voice information according to similarity calculation to obtain the age and gender of the user includes: calculating each feature in the user voice information through cosine similarity calculation to obtain a similarity score; and matching the age and gender of the user based on the similarity score.

[0012] According to another aspect of the embodiments of the present invention, there is provided an identity recognition device, including: an acquisition module for acquiring user voice information; an analysis module for analyzing the user voice through a pre-trained feature extraction model to obtain the voiceprint feature of each sentence in the user voice information; a calculation module for calculating the similarity between each feature in the user voice information according to similarity calculation to obtain the age and gender of the user; and an identification module for generating the identity recognition information of the user based on the voiceprint feature, age, and gender.

[0013] According to another aspect of the embodiments of the present invention, there is provided a terminal device, including: an intelligent terminal to which the above method is applied.

[0014] Optionally, the intelligent terminal includes: a smart speaker, or a terminal carrying a client with a voiceprint recognition function.

[0015] According to another aspect of the embodiments of the present invention, there is provided a non-volatile storage medium, where the non-volatile storage medium includes a stored program, and when the program runs, it controls the device where the non-volatile storage medium is located to execute the above method.

[0016] According to another aspect of the embodiments of the present invention, there is provided a processor, where the processor is used to run a program, and when the program runs, it executes the above method.

[0017] In an embodiment of the present invention, user voice information is obtained; the user voice is parsed by a feature extraction model obtained through pre-training to obtain the voiceprint feature of each sentence in the user voice information; the similarity between the features in the user voice information is calculated according to similarity calculation to obtain the age and gender of the user; and user identity recognition information is generated based on the voiceprint feature, age, and gender, achieving the purpose of implementing parallel training using a single data model, thereby realizing the enhancement of the generalization ability of the voiceprint and simultaneously providing the technical effect of a service for gender and age recognition using voice signals, and further solving the technical problem that the models used in the related art are too complex and not optimized synchronously. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0019] Figure 1 is a hardware structure block diagram of a computer terminal for a method of identity recognition according to an embodiment of the present invention;

[0020] Figure 2 is a flowchart of a method of identity recognition according to Embodiment 1 of the present invention;

[0021] Figure 3 is a schematic flowchart of obtaining a loss function in a method of identity recognition according to Embodiment 1 of the present invention;

[0022] Figure 4 is a schematic diagram of a device for identity recognition according to Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0024] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] According to an embodiment of the present invention, there is also provided a method embodiment of an identity recognition method. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0026] The method embodiment provided in the first embodiment of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking the execution on a computer terminal as an example, Figure 1 is a hardware structure block diagram of a computer terminal for an identity recognition method according to an embodiment of the present invention. As Figure 1 shown, the computer terminal 10 may include one or more (only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown.

[0027] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / modules corresponding to the identity recognition method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizes the identity recognition method of the above-mentioned application program. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0028] The transmission module 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the computer terminal 10. In one instance, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission module 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0029] Under the above operating environment, the present application provides an identity recognition method as Figure 2 shown. Figure 2 It is a flowchart of the identity recognition method according to Embodiment 1 of the present invention. The identity recognition method provided by the embodiments of the present application includes the following steps:

[0030] Step S202, obtain user voice information;

[0031] In the above step S202 of the present application, the identity recognition method provided by the embodiments of the present application can be applied to a natural language processing scenario (Natural Language Processing, abbreviated as NLP). In the process of obtaining user voice information, user voice signals can be collected through a client or a terminal device, and the acoustic features and semantics of the user can be obtained by analyzing the user voice signals, and then the user voice information can be obtained. The terminal device may be a device installed with software APPs having functions of voice collection, analysis, interaction, and communication, or a device having functions of voice collection, analysis, interaction, and communication; the client may be a software APP having functions of voice collection, analysis, interaction, and communication.

[0032] Among them, the terminal device may include: smart phones, tablet computers, laptop computers, desktop computers, supercomputers, smart wearable devices (such as smart watches, smart bracelets, smart glasses), and smart speakers.

[0033] In a preferred solution, obtaining the user voice information in step S202 includes: collecting the user's voice information through the client; obtaining the voice signal spectrum through spectrum transformation based on the voice information; obtaining the acoustic features through acoustic processing based on the voice signal spectrum; and determining the acoustic features as the user voice information.

[0034] Specifically, collecting the user's voice information through the client to obtain the voice signal spectrum through short-time Fourier transform, and then using methods such as Mel scale windowing, energy extraction, logarithm taking, and DCT (discrete cosine transform) to obtain the acoustic features, and further obtaining the user voice information.

[0035] Step S204, parsing the user voice through the pre-trained feature extraction model to obtain the voiceprint features of each sentence in the user voice information;

[0036] The pre-trained feature extraction model in step S204 of the present application is trained based on the CNN and Transformer models. Among them, the encoder-decoder architecture is adopted in the Transformer, which is widely used in the NLP field, such as machine translation, question answering systems, text summarization, and speech recognition.

[0037] Among them, the pre-trained feature extraction model includes: setting a preset number of encoding layers for each convolutional layer; synthesizing all the frame features of the input voice that have passed through the convolutional network through the added self-attention mechanism on the encoding layer to obtain the synthesized features; normalizing the synthesized features after passing through the linear layer to obtain the first loss function and the second loss function; training according to the first loss function and the second loss function to obtain the trained feature extraction model, where the first loss function is used to extract the voiceprint features in the user voice information, and the second loss function is used to extract the age and gender in the user voice information.

[0038] Optionally, before setting a preset number of encoding layers for each convolutional layer, the identity recognition method provided by the embodiments of the present application further includes: configuring the convolutional neural network, where configuring the convolutional neural network includes: setting the stride of each convolutional layer to a preset value.

[0039] Optionally, training is performed based on the first loss function and the second loss function, and the trained feature extraction model obtained includes: performing a weighted sum of the first loss function and the second loss function to obtain a loss function to be trained; performing sample training based on the loss function to be trained, and correcting the weights of the first loss function and the second loss function to obtain the trained feature extraction model.

[0040] Specifically, in voiceprint feature modeling (i.e., the feature extraction model pre-trained in the embodiments of the present application), first, three convolutional layers (Convolutional Neural Network, CNN) are used, and the stride of each convolutional layer is 2, so the final number of frames is 1 / 8 of the original; 3 encoder layers in 3 transformers (i.e., the encoding layer in the embodiments of the present application) are used on the convolutional layer, and the output dimension is 128; a pooling method of self-attention is added on the encoder layer to synthesize all the frame features passing through the network model in a sentence (i.e., at least two frame features passing through the network model in the embodiments of the present application) into a final feature, and three linear layers are passed through; after normalization, there are two loss functions. One is the end-to-end loss function loss1 of the speaker (i.e., the first loss function in the embodiments of the present application), and the other is the cross-entropy loss function loss2 after softmax for gender and age recognition of the person (i.e., the second loss function in the embodiments of the present application).

[0041] Based on the obtained first loss function and second loss function, a weighted sum of the first loss function and the second loss function is performed: loss = a1 * loss1 + a2 * loss2, where a1 and a2 are weight values, and the obtained loss is the input quantity in the embodiments of the present application and is debugged during training; after training is completed, there is no need to train a scorer. The voiceprint feature (voiceprint embedding) of each sentence is obtained.

[0042] In summary, Figure 3 is a schematic flow diagram of obtaining a loss function in the identity recognition method according to Embodiment 1 of the present invention. As Figure 3 shown, the identity recognition method provided by the embodiments of the present application is specifically as follows:

[0043] Step1, collect user voice information through a client or a terminal device, and process the user voice information through Short-time Fourier transform (STFT) to obtain sound features;

[0044] Step 2, through multi-layer convolution calculation, identify the voiceprint features (Voiceprint Embedding);

[0045] Step 3, based on the output of the multi-layer convolution calculation in Step 2, obtain the first loss function (marked as speaker e2eloss) and the second loss function (marked as gender and age softmax loss).

[0046] Step S206, calculate the similarity between each feature in the user voice information according to the similarity calculation, and obtain the age and gender of the user;

[0047] In a preferred example, calculating the similarity between each feature in the user voice information according to the similarity calculation in Step S206 to obtain the age and gender of the user includes: calculating each feature in the user voice information through cosine similarity calculation to obtain a similarity score; matching the age and gender of the user according to the similarity score.

[0048] Specifically, after obtaining the voiceprint features (voiceprint embedding) of each sentence based on Step S204, the cosine distance similarity (CDS) can be used to obtain the similarity score between two features, and the recognition results of gender and age can be output simultaneously.

[0049] Step S208, generate the identity recognition information of the user based on the voiceprint features, age, and gender.

[0050] The identity recognition method provided by the embodiments of the present application uses an end-to-end loss function to train the voiceprint feature extraction model and the scorer simultaneously under one optimization objective, thereby solving the problem that the models used in the related technologies are too complex and not optimized synchronously. In addition, a multi-task learning mechanism for gender and age is added to increase the generalization ability of the voiceprint and provide services for gender and age recognition using voice signals at the same time.

[0051] In the embodiments of the present invention, by obtaining the user voice information; parsing the user voice through a pre-trained feature extraction model to obtain the voiceprint features of each sentence in the user voice information; calculating the similarity between each feature in the user voice information according to the similarity calculation to obtain the age and gender of the user; generating the identity recognition information of the user based on the voiceprint features, age, and gender, the purpose of parallel training using one data model is achieved, thereby achieving the technical effect of increasing the generalization ability of the voiceprint and providing services for gender and age recognition using voice signals at the same time, and further solving the technical problem that the models used in the related technologies are too complex and not optimized synchronously.

[0052] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0053] Through the description of the above embodiments, those skilled in the art can clearly understand that the method for identity recognition according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0054] According to an embodiment of the present invention, there is also provided a device for implementing the above method for identity recognition. Figure 4 It is a schematic diagram of the device for identity recognition according to Embodiment 2 of the present invention. As Figure 4 shown, the device for identity recognition provided by the embodiment of the present application includes: an acquisition module 42 for acquiring user voice information; an analysis module 44 for analyzing the user voice through a pre-trained feature extraction model to obtain the voiceprint feature of each sentence in the user voice information; a calculation module 46 for calculating the similarity between the features in the user voice information according to similarity calculation to obtain the age and gender of the user; and an identification module 48 for generating the identity recognition information of the user based on the voiceprint feature, age and gender.

[0055] Optionally, the pre-trained feature extraction model includes: setting a preset number of encoding layers for each convolutional layer; synthesizing all the frame features passing through the convolutional network in the input voice through the added self-attention mechanism on the encoding layer to obtain the synthesized features; normalizing the synthesized features after passing through a linear layer to obtain a first loss function and a second loss function; and training according to the first loss function and the second loss function to obtain the trained feature extraction model, where the first loss function is used to extract the voiceprint feature in the user voice information, and the second loss function is used to extract the age and gender in the user voice information.

[0056] Optionally, before setting a preset number of encoding layers for each convolutional layer, configure the convolutional neural network, where configuring the convolutional neural network includes: setting the stride of each convolutional layer to a preset value.

[0057] Optionally, training according to the first loss function and the second loss function to obtain a trained feature extraction model includes: performing weighted summation on the first loss function and the second loss function to obtain a loss function to be trained; performing sample training according to the loss function to be trained, and correcting the weights of the first loss function and the second loss function to obtain a trained feature extraction model.

[0058] Optionally, the obtaining module 42 includes: collecting the voice information of the user through the client; obtaining the voice signal spectrum through spectrum transformation according to the voice information; obtaining the acoustic features through acoustic processing according to the voice signal spectrum; and determining the acoustic features as the user voice information.

[0059] Optionally, the calculation module 46 includes: calculating each feature in the user voice information through cosine similarity calculation to obtain a similarity score; and matching the age and gender of the user according to the similarity score.

[0060] According to another aspect of the embodiments of the present invention, a terminal device is provided, including: an intelligent terminal to which the above identity recognition method is applied.

[0061] Optionally, the intelligent terminal includes: an intelligent speaker, or a terminal carrying a client with a voiceprint recognition function.

[0062] According to another aspect of the embodiments of the present invention, a non-volatile storage medium is further provided, where the non-volatile storage medium includes a stored program, and when the program runs, it controls the device where the non-volatile storage medium is located to execute the above identity recognition method.

[0063] According to another aspect of the embodiments of the present invention, a processor is further provided, where the processor is used to run a program, and when the program runs, it executes the above identity recognition method.

[0064] An embodiment of the present invention further provides a storage medium. Optionally, in this embodiment, the above storage medium can be used to save the program code executed by the identity recognition method provided in the first embodiment above.

[0065] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group.

[0066] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining user voice information; parsing the user voice through a pre-trained feature extraction model to obtain the voiceprint feature of each sentence in the user voice information; calculating the similarity between the features in the user voice information according to similarity calculation to obtain the age and gender of the user; generating the identity recognition information of the user based on the voiceprint feature, age and gender.

[0067] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: The pre-trained feature extraction model includes: setting a preset number of encoding layers for each convolutional layer; synthesizing all frame features passing through the convolutional network in the input voice through the added self-attention mechanism on the encoding layer to obtain the synthesized feature; normalizing the synthesized feature after passing through the linear layer to obtain the first loss function and the second loss function; training according to the first loss function and the second loss function to obtain the trained feature extraction model, where the first loss function is used to extract the voiceprint feature in the user voice information, and the second loss function is used to extract the age and gender in the user voice information.

[0068] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: Before setting a preset number of encoding layers for each convolutional layer, configure the convolutional neural network, where configuring the convolutional neural network includes: setting the stride of each convolutional layer to a preset value.

[0069] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: Training according to the first loss function and the second loss function to obtain the trained feature extraction model includes: performing weighted summation on the first loss function and the second loss function to obtain the loss function to be trained; performing sample training according to the loss function to be trained and correcting the weights of the first loss function and the second loss function to obtain the trained feature extraction model.

[0070] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: Obtaining user voice information includes: collecting the voice information of the user through the client; obtaining the voice signal spectrum according to the voice information through spectrum transformation; obtaining the acoustic feature according to the voice signal spectrum through acoustic processing; determining the acoustic feature as the user voice information.

[0071] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: calculating the similarity between features in the user voice information according to similarity calculation to obtain the age and gender of the user, including: calculating the features in the user voice information through cosine similarity calculation to obtain a similarity score; and matching the age and gender of the user based on the similarity score.

[0072] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0073] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0074] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in an electrical or other form.

[0075] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0076] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0077] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.

[0078] The foregoing are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for identity recognition, comprising: Obtaining user voice information; Parsing the user voice through a pre-trained feature extraction model to obtain the voiceprint feature of each sentence in the user voice information. The feature extraction model is trained based on a first loss function and a second loss function. The first loss function is used to extract the voiceprint feature in the user voice information, and the second loss function is used to extract the age and gender in the user voice information; Calculating the similarity between the features in the user voice information according to the similarity calculation to obtain the age and gender of the user; Generating the identity recognition information of the user based on the voiceprint feature, age and gender.

2. The method according to claim 1, wherein The pre-trained feature extraction model includes: Setting a preset number of encoding layers for each convolutional layer; Synthesizing all the frame features passing through the convolutional network in the input voice through the added self-attention mechanism on the encoding layer to obtain the synthesized feature; Normalizing the synthesized feature after passing through the linear layer to obtain the first loss function and the second loss function.

3. The method according to claim 2, wherein, Before setting a preset number of encoding layers for each convolutional layer, the method further includes: Configuring the convolutional neural network, where configuring the convolutional neural network includes: setting the stride of each convolutional layer to a preset value.

4. The method according to claim 2, wherein, Training based on the first loss function and the second loss function to obtain the trained feature extraction model includes: Performing weighted summation on the first loss function and the second loss function to obtain the loss function to be trained; Performing sample training based on the loss function to be trained and correcting the weights of the first loss function and the second loss function to obtain the trained feature extraction model.

5. The method according to claim 1, wherein, The obtaining user voice information includes: Collecting the voice information of the user through the client; Obtaining the voice signal spectrum through spectrum transformation based on the voice information; Obtaining the acoustic feature through acoustic processing based on the voice signal spectrum; Determining the acoustic feature as the user voice information.

6. The method according to claim 1, wherein The calculating the similarity between the features in the user voice information according to the similarity calculation to obtain the age and gender of the user includes: Calculating the features in the user voice information through cosine similarity calculation to obtain the similarity score; Matching the age and gender of the user based on the similarity score.

7. An identity recognition device, comprising: An obtaining module, configured to obtain user voice information; An analysis module, configured to parse the user voice through a pre-trained feature extraction model to obtain the voiceprint feature of each sentence in the user voice information. The feature extraction model is trained based on a first loss function and a second loss function. The first loss function is used to extract the voiceprint feature in the user voice information, and the second loss function is used to extract the age and gender in the user voice information; A calculation module, configured to calculate the similarity between the features in the user voice information according to the similarity calculation to obtain the age and gender of the user; An identification module for generating identity identification information of the user based on the voiceprint feature, age and gender.

8. A terminal device, comprising: An intelligent terminal applied to the method according to any one of claims 1 to 6.

9. The terminal device according to claim 8, wherein, The intelligent terminal includes: a smart speaker, or a terminal carrying a voiceprint recognition function client.

10. A non-volatile storage medium, wherein, The non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute the method according to claim 1.

11. A processor, wherein, The processor is used to run a program, wherein when the program runs, it executes the method according to claim 1.

Citation Information

Patent Citations

  • Automatic registration method and device as well as intelligent equipment

    CN110689894A