Speech recognition method, device and equipment and computer readable storage medium

By combining voiceprint recognition and age prediction models, using voiceprint vectors and speech features for prediction processing, and matching them with voiceprint library, the problem of low accuracy of speech age recognition in the prior art is solved, and higher recognition accuracy and model generalization ability are achieved.

CN120220656APending Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311809227.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing speech recognition technologies have problems in the quality and scale of data sets directly determine the recognition performance in terms of speech age recognition, resulting in limited generalization ability of the model and low prediction accuracy.

Method used

By obtaining the voiceprint recognition model to be recognized, the trained voiceprint recognition model and the age prediction model, the voiceprint vector and voice feature prediction processing are carried out, and auxiliary judgment is performed in combination with the voiceprint library to improve the accuracy of voice recognition.

Benefits of technology

It effectively improves the recognition accuracy of the corresponding age of speech, enhances the performance and generalization ability of the model, and improves the accuracy of speech recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220656A_ABST
    Figure CN120220656A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method, device and equipment and a computer readable storage medium. The method comprises the following steps: acquiring to-be-recognized voice, a trained voiceprint recognition model and a trained age prediction model; performing feature extraction on the to-be-recognized voice to obtain to-be-recognized voice features; performing prediction processing on the to-be-recognized voice features by using the trained voiceprint recognition model to obtain a first voiceprint vector corresponding to the to-be-recognized voice; performing prediction processing on the first voiceprint vector and the to-be-recognized voice feature by using a trained age prediction model to obtain an initial prediction result; and determining a target prediction result of the to-be-recognized voice based on the initial prediction result, and outputting the target prediction result. The voice recognition method and device can be applied to an anti-addiction system, and the accuracy of voice recognition can be improved through the voice recognition method and device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular, to a speech recognition method, apparatus, device, and computer-readable storage medium. Background Art

[0002] Age recognition based on speech aims to recognize the age of a target speech through deep learning methods, which has always been one of the research hotspots in the speech field. The existing technical solutions are mainly end-to-end recognition. First, in the constructed dataset, each speech sample corresponds to a corresponding age label, and then deep learning methods are used for modeling, and classification or regression training is carried out in a supervised manner. As a pure data-driven method, this method has a fatal problem. The quality and scale of the dataset will directly determine the recognition performance, and the cost of obtaining a large amount of labeled speech data is very expensive, which leads to limited generalization ability of the model, resulting in low prediction accuracy of the model. Summary of the Invention

[0003] Embodiments of this application provide a speech recognition method, apparatus, and computer-readable storage medium, which can effectively improve the recognition accuracy of the age corresponding to the speech.

[0004] The technical solution of the embodiments of this application is implemented as follows:

[0005] Embodiments of this application provide a speech recognition method, the method comprising:

[0006] Obtain the speech to be recognized, a trained voiceprint recognition model, and a trained age prediction model, where the trained age prediction model is trained using training voiceprint vectors and training speech features;

[0007] Extract features from the speech to be recognized to obtain speech features to be recognized;

[0008] Use the trained voiceprint recognition model to perform prediction processing on the speech features to be recognized to obtain a first voiceprint vector corresponding to the speech to be recognized;

[0009] Use the trained age prediction model to perform prediction processing on the first voiceprint vector and the speech features to be recognized to obtain an initial prediction result;

[0010] Determine a target prediction result of the speech to be recognized based on the initial prediction result, and output the target prediction result.

[0011] Embodiments of this application provide a speech recognition apparatus, comprising:

[0012] A first acquisition module, configured to acquire a voice to be recognized, a trained voiceprint recognition model, and a trained age prediction model, where the trained age prediction model is trained using training voiceprint vectors and training voice features;

[0013] A first extraction module, configured to extract features from the voice to be recognized to obtain voice features to be recognized;

[0014] A first prediction module, configured to perform prediction processing on the voice features to be recognized using the trained voiceprint recognition model to obtain a first voiceprint vector corresponding to the voice to be recognized;

[0015] A second prediction module, configured to perform prediction processing on the first voiceprint vector and the voice features to be recognized using the trained age prediction model to obtain an initial prediction result;

[0016] An output module, configured to determine a target prediction result of the voice to be recognized based on the initial prediction result and output the target prediction result.

[0017] An embodiment of the present application provides an electronic device, where the electronic device includes:

[0018] A memory, configured to store computer-executable instructions;

[0019] A processor, configured to implement the voice recognition method provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0020] An embodiment of the present application provides a computer-readable storage medium, storing a computer program or computer-executable instructions, which are configured to implement the voice recognition method provided by the embodiment of the present application when being executed by a processor.

[0021] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, where the computer program or computer-executable instructions implement the voice recognition method provided by the embodiment of the present application when being executed by a processor.

[0022] The embodiment of the present application has the following beneficial effects:

[0023] When age recognition of the speech to be recognized is required, first obtain a trained voiceprint recognition model and a trained age prediction model. The trained voiceprint recognition model is used to predict a first voiceprint vector based on the speech features of the speech to be recognized. Since in addition to training speech features when training the age prediction model, the training voiceprint vector of the deep representation is also added, so that the limited training data can be better utilized, the performance and generalization ability of the model can be improved, and when using the trained age prediction model for age prediction, it is to predict the first voiceprint vector of the speech to be recognized and the speech features to be recognized. Therefore, the accuracy of the speech recognition result can be improved. Brief Description of the Drawings

[0024] Figure 1 is a schematic structural diagram of the application data processing system 100 provided by an embodiment of the present application;

[0025] Figure 2 is a schematic structural diagram of the server 400 provided by an embodiment of the present application;

[0026] Figure 3A is a schematic implementation flowchart of an application data processing method provided by an embodiment of the present application;

[0027] Figure 3B is a schematic implementation flowchart of determining the target prediction result of the speech to be recognized provided by an embodiment of the present application;

[0028] Figure 3C is a schematic implementation flowchart of determining the target clustering cluster to which the first voiceprint vector belongs provided by an embodiment of the present application;

[0029] Figure 3D is a schematic implementation flowchart of updating the reference age information provided by an embodiment of the present application;

[0030] Figure 3E is a schematic implementation flowchart of determining the target prediction result of the speech to be recognized provided by an embodiment of the present application;

[0031] Figure 3F is a schematic implementation flowchart of training the voiceprint recognition model provided by an embodiment of the present application;

[0032] Figure 3G is a schematic implementation flowchart of training the age prediction model provided by an embodiment of the present application;

[0033] Figure 3H is a schematic implementation flowchart of the specific training of the age prediction model provided by an embodiment of the present application;

[0034] Figure 4 is another schematic implementation flowchart of the speech recognition method provided by an embodiment of the present application;

[0035] Figure 5 Schematic flow diagram of the ECAPA-TDNN network provided by the embodiments of the present application;

[0036] Figure 6 Schematic flow diagram of the SE-Res2Block part provided by the embodiments of the present application;

[0037] Figure 7 Schematic implementation flow diagram of age prediction using the age prediction model provided by the embodiments of the present application;

[0038] Figure 8 Schematic structural diagram of feature extraction provided by the embodiments of the present application;

[0039] Figure 9 Schematic specific structural diagram of the conformer layer provided by the embodiments of the present application;

[0040] Figure 10 Schematic structural diagram of the forward propagation sub-module provided by the embodiments of the present application;

[0041] Figure 11 Schematic structural diagram of the multi-head attention module provided by the embodiments of the present application. Detailed implementation manners

[0042] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0043] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0044] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0045] In the embodiments of the present application, data related to user information, etc. is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0046] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by those of ordinary skill in the art to which this application pertains. The terms used in the embodiments of this application are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0047] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are explained, and the nouns and terms involved in the embodiments of this application are applicable to the following explanations.

[0048] 1) Mel Frequency Cepstral Coefficients (MFCC): A commonly used feature extraction method in speech signal processing. It mimics the auditory characteristics of the human ear and is widely used in fields such as speech recognition and speaker recognition. It performs a short-time Fourier transform on the speech signal and then passes it through a series of Mel filters, making the extracted features more in line with the auditory characteristics of the human ear. This feature extraction method is widely used in the speech field, can better capture the important features of the speech signal, and has a certain robustness to noise and other interferences. Therefore, Mel Frequency Cepstral Coefficients play an important role in tasks such as speech recognition and speaker recognition and have become an important technical means in the field of speech signal processing.

[0049] 2) Self-Attention (SA): A mechanism widely used in the fields of natural language processing and computer vision. It is a technique for modeling long-range dependencies in sequential data. Compared with traditional recurrent neural networks or convolutional neural networks, the self-attention mechanism can consider all positions in the input sequence simultaneously and explicitly encode the relationships between different positions as weights. The core of the self-attention mechanism is to determine the weights by calculating the similarities between queries, keys, and values. The query is the position for finding relevant information, and the key and value are the positions for providing information. By calculating the similarity between the query and the key and normalizing it to weights, the association degree of each query with all other positions can be obtained. Next, multiplying the weights by the values and performing weighted summation can obtain a new representation with global context information.

[0050] 3) Multi-Headed self-attention module (MSA): A technique commonly used in the field of deep learning, which is a further improvement based on the self-attention mechanism. Compared with the self-attention mechanism with a single attention head, MSA calculates the attention on the input in parallel by introducing multiple attention heads, enabling the model to simultaneously learn different attention representations. Each attention head can focus on different parts of the input sequence and learn different semantic information. Finally, by combining the results of multiple heads, the model's understanding of the input is enriched. This parallel computing method can enhance the model's ability to model long-range dependencies and strengthen the model's ability to capture and represent global information. Therefore, it has been widely used in natural language processing and computer vision tasks and has become an important technical means to improve model performance.

[0051] 4) Feed Forward Module (FFM): The feed forward module is a commonly used module in deep learning for non-linear transformation and feature extraction of the input. This module usually consists of two linear transformation layers (or fully connected layers) and a non-linear activation function. Through the feed forward module, the model can gradually extract higher-level and more abstract representations by learning to transform the input features layer by layer. The core idea of the feed forward module is to use linear transformation and non-linear activation functions to map the input. First, the input data passes through the first linear transformation, mapping it from the original feature space to an intermediate representation space. Then, by applying the non-linear activation function, the result of the linear transformation is non-linearly transformed. Finally, after passing through the second linear transformation, the intermediate representation is mapped back to the original feature space to generate the final output.

[0052] 5) Convolutional Neural Network (CNN): A deep learning model specifically designed to process data with a similar grid structure. The core idea of CNN is to automatically extract features from the input data through convolutional layers and pooling layers, and perform classification or regression prediction through fully connected layers and softmax layers. The convolutional layer extracts features in the local receptive field using convolutional operations, reducing the number of parameters through weight sharing and local connections, thereby effectively capturing the spatial structure information of the input data; while the pooling layer reduces the size of the feature map through downsampling operations, reducing the computational complexity while retaining key information. Since CNN performs well in fields such as images, videos, and speech, it has been widely used in fields such as image recognition, object detection, and face recognition, and has achieved remarkable results in many tasks.

[0053] 6) Frame Division: An important technique in speech signal processing. By dividing the original speech signal into small segments of fixed length (usually called frames), independent feature extraction and analysis can be performed on each small segment. Each frame usually has a length of 10 milliseconds to 30 milliseconds. This can improve the local stability of the signal while retaining the short-term characteristics of the speech signal, and provides a basis for subsequent signal processing and feature extraction. The frame division technique has a wide range of applications in the fields of speech recognition, speech synthesis, and speech signal analysis, and is an indispensable part of speech processing.

[0054] 7) Windowing: A process performed after signal frame division to eliminate the spectral leakage effect. It is achieved by multiplying each frame of the speech signal by a window function. This method can reduce the spectral leakage caused by signal truncation, and at the same time maintain the smooth transition and continuity between frames, which is beneficial to extracting accurate spectral features. In speech processing, each frame is brought into the window function to form the windowed speech signal sw(n) = s(n) * w(n). Commonly used window functions include the rectangular window and the Hamming window, which can effectively reduce the discontinuity of the signal at the frame boundary and improve the accuracy and stability of signal processing.

[0055] In the prior art, when performing speech recognition, it is mainly end-to-end recognition. First, data collection is carried out, and each speech sample corresponds to a corresponding age label. Subsequently, deep learning methods are used for modeling, and classification or regression training is carried out in a supervised manner. In the field of speech recognition, end-to-end deep learning modeling is the most common method. Currently, there is a key problem: due to the boundary effect of speech (the speech similarity between speakers aged 12 to 18 and adults is extremely high, especially for females; or, the speech similarity between the elderly over 60 years old and middle-aged people aged 35 to 59 is also extremely high), that is, it is difficult for end-to-end deep learning methods to show good recognition ability under a limited data set. How to further improve the generalization and recognition ability of the model has become the most important problem at present.

[0056] To solve the problems existing in the prior art, the embodiments of this application propose a speech recognition method based on speaker deep representation. Aiming at the problem of low generalization and accuracy of the current end-to-end deep learning method, the embodiments of this application introduce speaker deep representation into speech recognition, enabling the neural network to additionally obtain the voiceprint characteristics of the speech during training; at the same time, the embodiments of this application build a voiceprint library for each speaker and perform auxiliary decision through voiceprint comparison. The embodiments of this application significantly improve the performance of speech recognition.

[0057] The embodiments of the present application provide a voice recognition method, device, equipment, computer-readable storage medium and computer program product, which can improve the accuracy of real-time voice corresponding age recognition in the voice field by introducing voiceprint vectors.

[0058] The following describes the exemplary applications of the electronic device provided by the embodiments of the present application. The device provided by the embodiments of the present application can be implemented as various types of user terminals such as laptops, tablets, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), smartphones, smart speakers, smart watches, smart TVs, in-vehicle terminals, etc., or can be implemented as a server. Below, the exemplary applications when the device is implemented as a server will be described.

[0059] See Figure 1 , Figure 1 is a schematic diagram of the architecture of the application data processing system 100 provided by the embodiments of the present application. To implement a support for an exemplary application, the terminal 200 is connected to the server 400 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two. The data processing system 100 may further include a database 500 for voice data of the voice to be recognized, training voice data, label data, trained voiceprint recognition model data, trained age prediction model data, etc. The database 500 can be independent of the server 400 or located inside the server 400. In Figure 1 the example where the database 500 is independent of the server 400 is described.

[0060] The server 400 receives the voice to be recognized sent by the terminal 200, and the server 400 extracts features from the voice to be recognized to obtain the voice feature to be recognized. After determining the trained voiceprint recognition model, the server 400 uses the trained voiceprint recognition model to perform prediction processing on the voice feature to be recognized to obtain the first voiceprint vector corresponding to the voice to be recognized. Then, after determining the trained age prediction model, the server 400 uses the trained age prediction model to perform prediction processing on the first voiceprint vector and the voice feature to be recognized to obtain an initial prediction result. Finally, the server 400 determines the target prediction result of the voice to be recognized based on the initial prediction result and outputs the target prediction result to the terminal 200. So that the terminal 200 can determine whether the object corresponding to the voice to be recognized meets the age condition according to the target prediction result.

[0061] In some embodiments, the server 400 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal 200 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and there is no limitation in the embodiments of the present application.

[0062] See Figure 2 , Figure 2 is a schematic structural diagram of the server 400 provided by the embodiments of the present application. Figure 2 The server 400 shown in Figure 2 includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 440.

[0063] The processor 410 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a Digital Signal Processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.

[0064] The user interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, and other input buttons and controls.

[0065] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically located away from the processor 410.

[0066] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0067] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.

[0068] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0069] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, wireless fidelity (WiFi), and universal serial bus (USB), etc.;

[0070] The presentation module 453 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.);

[0071] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one of one or more input devices 432.

[0072] In some embodiments, the device provided in the embodiments of the present application may be implemented in software. Figure 2 Shown is a voice recognition device 455 stored in the memory 450, which may be software in the form of programs and plugins, etc., including the following software modules: a first acquisition module 4551, a first extraction module 4552, a first prediction module 4553, a second prediction module 4554, and an output module 4555. These modules are logical, and thus can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0073] In some other embodiments, the device provided by the embodiments of the present application may be implemented in a hardware manner. As an example, the device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the voice recognition method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0074] Next, the application data processing method provided by the embodiments of the present application will be described. As mentioned above, the electronic device for implementing the application data processing method of the embodiments of the present application may be a server. Therefore, the execution subject of each step will not be repeated hereinafter. Refer to Figure 3A , Figure 3A is a schematic flowchart of an implementation process of the application data processing method provided by the embodiments of the present application, which will be described in conjunction with Figure 3A the steps shown. Figure 3A The subject of the step is the server.

[0075] In step 101, the voice to be recognized, the trained voiceprint recognition model, and the trained age prediction model are obtained.

[0076] In some embodiments, the voice to be recognized refers to the voice data for which voiceprint recognition or age prediction needs to be performed. The voice to be recognized may be sent by the terminal to the server. The voice to be recognized may be collected in real time by the terminal using its own voice collection device, or downloaded by the terminal from the network, or sent by other terminals. For example, a microphone or a voice interface is used to collect the voice to be recognized. By establishing communication with these devices, real-time sound input can be received, and then these sound signals can be recorded and converted into digital data for subsequent prediction.

[0077] In some embodiments, the voiceprint recognition model and the age prediction model can be neural network models. The neural network models can be convolutional neural network models, deep learning neural network models, etc. For example, the voiceprint recognition model can be an ECAPA-TDNN network, and the age prediction model can be a neural network model composed of a conformer module and a fully connected layer module. The voiceprint recognition model and the age prediction model can be models pre-trained by the server using training data, or models that have been trained and obtained by the server from other external servers. They can also be downloaded by connecting to other systems or databases, and these models may be stored in the form of files on a remote server and can be downloaded through a network connection. After obtaining these models, the server loads them into memory for use in subsequent voiceprint recognition and age prediction tasks.

[0078] In step 102, feature extraction is performed on the speech to be recognized to obtain the speech feature to be recognized.

[0079] In some embodiments, before the prediction process, a series of numerical values representing sound features need to be extracted from the speech to be recognized, and a series of algorithms and technologies are used to perform feature extraction on the speech to be recognized. First, the speech to be recognized is converted into digital data, which is usually achieved through processes such as sampling and quantization. Then, signal processing and audio processing technologies are applied to these digital speech data to extract information related to sound features. In some embodiments, feature extraction can be achieved through technologies such as Short-time Fourier Transform (STFT), Cepstral Analysis, and Linear Predictive Coding (LPC). These technologies can decompose the speech signal into different frequency, phase, and amplitude components, thereby capturing the time-domain and frequency-domain features of the sound. The speech feature to be recognized generated during the feature extraction process is usually a vector of a series of numerical values. These vectors can contain features such as the spectrum information, formant information, speech rate, and intonation of the sound. Each vector corresponds to the feature description of the speech signal at a certain time period.

[0080] In step 103, the trained voiceprint recognition model is used to perform prediction processing on the speech feature to be recognized to obtain the first voiceprint vector corresponding to the speech to be recognized.

[0081] In some embodiments, a trained voiceprint recognition model is used to predict the features of the speech to be recognized, thereby generating a first vector representing the voiceprint features.

[0082] In this process, the speech features to be recognized are input into the trained voiceprint recognition model. The voiceprint recognition model is learned through the training process and has a certain ability to understand the relationships between different speech features. For prediction processing, the speech features to be recognized are input into the voiceprint recognition model, and the speech feature vectors are encoded and decoded to obtain the corresponding first voiceprint vector. The first voiceprint vector can contain the position and feature information of the speech to be recognized in the voiceprint space. This first voiceprint vector can be used in subsequent voiceprint comparison and age recognition tasks.

[0083] In step 104, the trained age prediction model is used to perform prediction processing on the first voiceprint vector and the speech features to be recognized, obtaining an initial prediction result.

[0084] In some embodiments, the trained age prediction model is trained using training voiceprint vectors and training speech features. That is to say, when training the age prediction model, in addition to the training speech features, the training voiceprint vectors with deep representations are added, so that the limited training data can be better utilized, and the performance and generalization ability of the model can be improved.

[0085] In some embodiments, the age prediction model may include a conformer module, a fully connected layer, and an output layer. When using the trained age prediction model to perform prediction processing on the first voiceprint vector and the speech features to be recognized, the conformer module can first be used to perform encoding and decoding processing on the speech features to be recognized, and then the output vector of the conformer module is concatenated with the first voiceprint vector to obtain a concatenated vector. Then, the fully connected layer is used to perform a fully connected process on the concatenated vector, and finally, the output layer is used to process the vector obtained after the fully connected process to obtain the initial prediction result of the speech to be recognized. It should be noted that the age prediction result is inferred based on the trained model and there may be a certain error.

[0086] In step 105, based on the initial prediction result, the target prediction result of the speech to be recognized is determined, and the target prediction result is output.

[0087] In some embodiments, the initial prediction result is used to determine the target prediction result of the speech to be recognized and output it. In this process, first, the initial prediction result is obtained by predicting the first voiceprint vector and the speech features to be recognized, and is represented as a specific age value. Based on this preliminary age prediction result, the age information in the voiceprint library is combined to determine the target prediction result of the speech to be recognized. In the process of determining the target prediction result, the age or other relevant information in the voiceprint library is used to further verify and adjust the initial prediction result. Once the target prediction result of the speech to be recognized is determined, the server outputs the target prediction result to the terminal or related devices, and the terminal can present the target prediction result in the form of numbers, graphics or other forms. This output result can be used to identify the age range of the speaker in the real-time audio, such as determining whether the age of the speaker meets the conditions for using relevant software or games, so as to give feedback such as allowing continued use or interrupting use to the object.

[0088] In some embodiments, refer to Figure 3B , Figure 3B which is a schematic diagram of the implementation process for determining the target prediction result of the speech to be recognized provided by the embodiments of the present application. Figure 3A The step 105 shown can be implemented by steps 1051 to 1055 as shown in Figure 3B . The following is a specific description with reference to Figure 3B .

[0089] Step 1051: Obtain the object identifier corresponding to the speech to be recognized.

[0090] In some embodiments, the object identifier can be a unique number, symbol or other form of identifier. Exemplarily, the object identifier can be the account information of the object.

[0091] Step 1052: Obtain multiple reference voiceprint vectors corresponding to the object identifier from the pre-constructed voiceprint library.

[0092] Among them, multiple reference voiceprint vectors form at least one reference clustering cluster. In some embodiments, it is necessary to obtain multiple voiceprint vectors associated with the object identifier from a pre-prepared voiceprint library. A voiceprint vector is a mathematical description used to represent the voice characteristics of a speaker. First, the voiceprint library is a pre-constructed database that contains the voiceprint information of multiple objects. This voiceprint information may be obtained through specific signal processing and feature extraction algorithms on the voice samples of the objects. Each object has a unique object identifier corresponding to it in the voiceprint library. Secondly, according to the object identifier corresponding to the voice to be recognized, the corresponding object voiceprint information will be retrieved in the voiceprint library. Specifically, the set of voiceprint vectors stored for the object identifier in the voiceprint library will be found according to the object identifier. An object identifier can correspond to multiple reference voiceprint vectors. There may be multiple and different speakers under this object identifier, or the characteristics of the voice may vary due to factors such as environment and emotion. For example, different objects use the same account to play games or browse software. These voiceprint vectors can be regarded as different voice expressions of the object information in various situations.

[0093] Furthermore, multiple reference voiceprint vectors can be organized into at least one reference clustering cluster. The clustering cluster is obtained by clustering multiple reference voiceprint vectors according to the similarity between the reference voiceprint vectors in the voiceprint library. By clustering the reference voiceprint vectors, the voice change range of the object can be represented, and an accurate reference can be provided for determining the subsequent target prediction result.

[0094] Step 1053: Determine the similarity between the first voiceprint vector and each reference voiceprint vector.

[0095] In some embodiments, the similarity comparison between the first voiceprint vector and each reference voiceprint vector includes steps such as similarity calculation and similarity evaluation. The similarity can be determined by calculating the Euclidean distance or cosine distance between the first voiceprint vector and the reference voiceprint vector. The Euclidean distance and cosine distance can measure the degree of difference between two vectors. Then, for each reference voiceprint vector, a similarity value can be obtained, indicating the similarity degree between the first voiceprint vector and the reference voiceprint vector. These similarity values can be normalized according to certain rules. For example, the similarity values can be converted into probability values between 0 and 1. Finally, after obtaining the similarity values between each reference voiceprint vector and the first voiceprint vector, the most similar reference voiceprint vector can be determined by comparing these values, or a comprehensive similarity score can be calculated. According to the specific application scenario, a threshold can be set for judgment, and the voiceprint vectors exceeding the threshold are considered to be successfully matched.

[0096] Step 1054: Based on each similarity, determine the target clustering cluster to which the first voiceprint vector belongs, and obtain the reference age information corresponding to the target clustering cluster.

[0097] Among them, the reference age information of the target clustering cluster includes at least two age intervals and the proportion of the number corresponding to each age interval. Exemplarily, the reference age information of the target clustering cluster may include two age intervals: [0, 18), [18, 150], that is, one of these two age intervals represents minors and the other represents adults. The proportion of the number corresponding to each age region is obtained by dividing the number of reference voiceprint vectors included in the age interval by the total number of reference voiceprint vectors, and the proportion of the number corresponding to each age interval is a real number between 0 and 1. Exemplarily, assuming that the total number of reference voiceprint vectors corresponding to the object identifier is 10, and the number of reference voiceprint vectors corresponding to the age interval [0, 18) is 2, then the proportion of the number corresponding to the age interval [0, 18) is 0.2. Similarly, the proportion of the number corresponding to the age interval [18, 150] is 0.8. Exemplarily, the reference age information of the target clustering cluster may also include four age intervals: [0, 18), [18, 35], [35, 59], [60, 120], and these four age intervals represent minors, young people, middle-aged people, and the elderly respectively.

[0098] In some embodiments, referring to Figure 3C , Figure 3C is a schematic diagram of the implementation process for determining the target clustering cluster to which the first voiceprint vector belongs provided by the embodiments of the present application. Figure 3B The step 1054 shown can be implemented through steps 541 to 546 as shown in Figure 3B , and the following will be specifically described in combination with Figure 3C .

[0099] Step 541: Determine the maximum similarity from each similarity.

[0100] In some embodiments, the maximum similarity can be determined from each similarity through a preset algorithm. For example, the preset algorithm can be a sorting algorithm or a maximum value determination algorithm.

[0101] Step 542: Judge the magnitude relationship between the maximum similarity and the similarity threshold.

[0102] In some embodiments, when the maximum similarity is greater than the preset similarity threshold, go to step 543; when the maximum similarity is less than or equal to the similarity threshold, go to step 545.

[0103] Step 543: Obtain the reference voiceprint vector corresponding to the maximum similarity.

[0104] In some embodiments, when the maximum similarity is greater than a preset similarity threshold, it indicates that the similarity between the reference voiceprint vector corresponding to the highest similarity and the first voiceprint vector corresponding to the voice to be recognized is very high. At this time, the reference voiceprint vector corresponding to the highest similarity is obtained. In implementation, data access technologies such as indexes or pointers can be used to quickly obtain the reference voiceprint vector corresponding to the maximum similarity.

[0105] Step 544: Determine the clustering cluster where the reference voiceprint vector corresponding to the maximum similarity is located as the target clustering cluster to which the first voiceprint vector belongs.

[0106] In some embodiments, when determining the target clustering cluster to which the first voiceprint vector belongs, a voiceprint comparison algorithm is used to compare the first voiceprint vector with the reference voiceprint vector and calculate the similarity score between them. During this period, a similarity threshold is set to determine whether the similarity between the first voiceprint vector and the reference voiceprint vector corresponding to the highest similarity reaches the standard for being recognized as the same clustering cluster. If the highest similarity is greater than the similarity threshold, it indicates that the first voiceprint vector can be added to the clustering cluster where the reference voiceprint vector corresponding to the highest similarity is located. Therefore, the clustering cluster where the reference voiceprint vector corresponding to the highest similarity is located is determined as the target clustering cluster of the first voiceprint vector.

[0107] In some embodiments, after step 544 determines the target clustering cluster to which the first voiceprint vector belongs, it is also necessary to Figure 3D update the reference age information corresponding to the object identifier through the steps 61 to 64 shown in Figure 3D which is a schematic flowchart of the implementation process for updating the reference age information provided by the embodiments of the present application. The following will be specifically described in conjunction with Figure 3D this.

[0108] Step 61: Add the first voiceprint vector to the target clustering cluster.

[0109] Step 62: Determine the first age interval to which the predicted age included in the initial prediction result belongs, and update the number of first voiceprint vectors corresponding to the first age interval.

[0110] In some embodiments, the predicted age included in the initial prediction result is a specific age value, such as 15 years old, 28 years old, 66 years old, etc. According to the predicted age information, the first voiceprint vector is assigned to the corresponding age interval, and the number of first voiceprint vectors contained in this age interval is updated. To implement this process, the predicted age information is compared with the preset age intervals. This comparison process may involve judging the predicted age and the boundaries of each age interval to determine the specific age interval to which the predicted age belongs. Once the first age interval to which the predicted age belongs is determined, the corresponding data record in the voiceprint database is searched, which may include the existing number information of the first voiceprint vectors. Then, the number of the first voiceprint vectors corresponding to this age interval is updated, possibly by increasing the existing quantity, or creating a new data record and initializing the value.

[0111] Step 63: Obtain the number of second voiceprint vectors corresponding to the second age interval in the reference age information.

[0112] In some embodiments, it is necessary to extract data of a specific age interval from the reference age information and obtain the number of second voiceprint vectors corresponding to this age interval. To implement this process, first, access the database or file storing the reference age information. In this data, each age interval may be associated with a specific number of voiceprint vectors. Next, according to the preset rules or indexes, locate the data record corresponding to the second age interval in the reference age information. This process involves comparing and matching the boundaries of the age intervals. Once the server successfully locates the data record corresponding to the second age interval, the number information of the second voiceprint vectors stored in this data record is extracted.

[0113] Step 64: Update the proportion of the first number in the first age interval and the proportion of the second number in the second age region based on the updated number of the first voiceprint vectors and the number of the second voiceprint vectors.

[0114] In some embodiments, the proportion of the number of voiceprint vectors in a specific age range in the voiceprint database is calculated and adjusted. To achieve this process, first, the updated number of first voiceprint vectors and the number of second voiceprint vectors are obtained. These data may be obtained from the reference age information through the methods mentioned above, or statistically derived from new voiceprint data. Next, the corresponding data record of the first age range in the voiceprint database is located, and the proportion information of the first number stored in this record is extracted. Similarly, the corresponding data record of the second age range in the voiceprint database is located, and the proportion information of the second number stored in this record is extracted. Then, based on the updated number of first voiceprint vectors and the number of second voiceprint vectors, the proportion of the first number in the first age range and the proportion of the second number in the second age range are recalculated. This may involve simple mathematical operations such as ratio calculation or percentage adjustment. Finally, the updated proportion of the first number in the first age range and the proportion of the second number in the second age range are stored back into the corresponding data records in the voiceprint database to ensure that the data in the voiceprint database can accurately reflect the latest number and proportion information of voiceprint vectors.

[0115] Step 545: Create a new clustering cluster.

[0116] Step 546: Determine the newly created clustering cluster as the target clustering cluster to which the first voiceprint vector belongs.

[0117] In some embodiments, when the maximum similarity is less than or equal to the similarity threshold, it indicates that the similarity between the first voiceprint vector and the existing reference voiceprint vectors is not sufficient to classify it into the existing clustering clusters. Therefore, a new clustering cluster needs to be created to accommodate this first voiceprint vector. During the process of creating a new clustering cluster, a unique clustering cluster identifier can be generated to distinguish different clustering clusters.

[0118] In some embodiments, after step 546, it is also necessary to add the first voiceprint vector that does not exist in the voiceprint library to the newly created target clustering cluster, and determine the reference age information of the target clustering cluster based on the initial prediction result corresponding to the speech to be recognized. The initial prediction result corresponding to the speech to be recognized is an age value. First, determine the age range where the age value is located, and then update the proportion of the number of the age range where the age value is located. Since there is only one first voiceprint vector in the newly created target clustering cluster at this time, the proportion of the number of the age range where the age value is located is 100%, and the proportion of the number of other age ranges is 0. Exemplarily, the initial prediction result corresponding to the speech to be recognized is 10 years old, and the age range where it is located is [0, 18). At this time, the number of reference voiceprint vectors corresponding to the age range [0, 18) is 1, and the total number of reference voiceprint vectors in the target clustering cluster is also 1. Then the proportion of the number of the age range [0, 18) is 100%, and the number of reference voiceprint vectors corresponding to the age range [18, 150] is 0. Then the proportion of the number of the age range [18, 150] is 0. That is to say, the reference age information of the target clustering cluster at this time includes: the number of reference voiceprint vectors included in [0, 18) is 1, and the corresponding proportion of the number is 100%, and the number of reference voiceprint vectors included in [18, 150] is 0, and the corresponding proportion of the number is 0.

[0119] In the above steps 541 to 546, first, by determining the target clustering cluster to which the first voiceprint vector belongs based on each similarity, the clustering classification of the voiceprint vectors can be realized. This enables the voices with similar voiceprint characteristics to be effectively classified into the same clustering cluster, thereby realizing the effective organization and management of the voiceprint data. Secondly, when the maximum similarity is greater than the preset similarity threshold, the reference voiceprint vector corresponding to the maximum similarity can be obtained, and the clustering cluster where it is located is determined as the target clustering cluster to which the first voiceprint vector belongs. In addition, when the maximum similarity is less than or equal to the similarity threshold, a new clustering cluster is created, and the newly created clustering cluster is determined as the target clustering cluster to which the first voiceprint vector belongs, so as to ensure the accuracy of the determined target clustering cluster to which the first voiceprint vector belongs, and provide an accurate data basis for determining the target prediction result of the speech to be recognized based on the reference age information of the target clustering cluster.

[0120] Step 1055: Determine the target prediction result of the speech to be recognized based on the reference age information and the initial prediction result.

[0121] In some embodiments, referring to Figure 3E , Figure 3E is a schematic flow chart of realizing the determination of the target prediction result of the speech to be recognized provided by the embodiment of the present application, Figure 3B The step 1055 shown can be as Figure 3EThe steps 551 to 555 shown are implemented. The following will specifically describe in conjunction with Figure 3E specific details.

[0122] Step 551: Determine whether the proportion of the number corresponding to each age interval is the same.

[0123] In some embodiments, the target clustering cluster may include multiple reference voiceprint vectors. The predicted ages corresponding to different reference voiceprint vectors may be the same or different. The reference age information of the target clustering cluster includes at least two age intervals and the proportion of the number corresponding to each of the age intervals. In some embodiments, the reference age information corresponding to the target clustering cluster may further include the number of reference voiceprint vectors corresponding to each age interval.

[0124] Exemplarily, four age intervals can be preset, which are: [0, 18), [18, 35], [35, 59], [60, 120]. The target clustering cluster includes 10 reference voiceprint vectors, which are reference voiceprint vector 1 to reference voiceprint vector 10; the predicted age corresponding to reference voiceprint vector 1 is 58 years old, the predicted age corresponding to reference voiceprint vector 2 is 49 years old, the predicted age corresponding to reference voiceprint vector 3 is 62 years old, the predicted age corresponding to reference voiceprint vector 4 is 55 years old, the predicted age corresponding to reference voiceprint vector 5 is 61 years old, the predicted age corresponding to reference voiceprint vector 6 is 64 years old, the predicted age corresponding to reference voiceprint vector 7 is 68 years old, the predicted age corresponding to reference voiceprint vector 8 is 59 years old, the predicted age corresponding to reference voiceprint vector 9 is 71 years old, and the predicted age corresponding to reference voiceprint vector 10 is 34 years old. Then the reference age information of this target clustering cluster is: the number of reference voiceprint vectors in the interval [0, 18] is 0, and the proportion is 0%, the number of reference voiceprint vectors in the interval [18, 35] is 1, and the proportion is 10%, the number of reference voiceprint vectors in the interval [35, 59] is 4, and the proportion is 40%, and the number of reference voiceprint vectors in the interval [60, 120] is 5, and the proportion is 50%.

[0125] Exemplarily, two age intervals can also be preset, which are [0, 18) and [18, 150] respectively. There are 10 reference voiceprint vectors in the target clustering cluster, which are reference voiceprint vector 1 to reference voiceprint vector 10; the predicted age corresponding to reference voiceprint vector 1 is 20 years old, the predicted age corresponding to reference voiceprint vector 2 is 22 years old, the predicted age corresponding to reference voiceprint vector 3 is 21 years old, the predicted age corresponding to reference voiceprint vector 4 is 20 years old, the predicted age corresponding to reference voiceprint vector 5 is 12 years old, the predicted age corresponding to reference voiceprint vector 6 is 20 years old, the predicted age corresponding to reference voiceprint vector 7 is 15 years old, the predicted age corresponding to reference voiceprint vector 8 is 20 years old, the predicted age corresponding to reference voiceprint vector 9 is 25 years old, and the predicted age corresponding to reference voiceprint vector 10 is 14 years old. Then the reference age information of this target clustering cluster is: the number of reference voiceprint vectors in the [0, 18) interval is 3, and the proportion is 30%, and the number of reference voiceprint vectors in the [18, 150] interval is 7, and the proportion is 70%.

[0126] In this step, it is judged whether the proportion corresponding to each age interval is the same. If the reference age information includes two age intervals, that is, it is judged whether the proportion corresponding to these two age intervals is the same. If the reference age information includes three or more age intervals, then it is judged whether the proportion of all age intervals is the same. Suppose the reference age information includes three age intervals: [0, 12), [13, 18), [18, 150], the proportion corresponding to [0, 12) is 0.1, the proportion corresponding to [13, 18) is 0.45, and the proportion corresponding to [18, 150] is 0.45. Then at this time, it is determined that the proportion corresponding to the three age intervals is not the same. When the proportion corresponding to each age interval is not the same, that is, when the proportion corresponding to at least two age intervals is different, go to step 552. When the proportion corresponding to each age interval is the same, go to step 554.

[0127] Step 552: Determine the maximum proportion.

[0128] In some embodiments, the proportion corresponding to each age interval will be compared to determine the maximum proportion.

[0129] Step 553: Based on the age interval corresponding to the maximum proportion and the initial prediction result, determine the target prediction result of the voice to be recognized.

[0130] In some embodiments, when there is only one age range corresponding to the largest percentage, the age label of this age range can be obtained. Here, the age label corresponding to the age range is pre-set. The age label can be minor, adult, or it can also be child, juvenile, young, middle-aged, elderly, etc. Suppose there are two age ranges: [0, 18) and [18, 150]. Then the age label corresponding to [0, 18) can be minor, and the age label corresponding to [18, 150] can be adult. Then, the age label of this age range is determined as the target prediction result of the speech to be recognized. In some embodiments, the age range corresponding to the largest percentage can also be directly determined as the target prediction result of the speech to be recognized.

[0131] When there are two or more age ranges corresponding to the largest percentage, the predicted age in the initial prediction result of the speech to be recognized can be obtained, and then the target prediction result of the speech to be recognized is determined based on the age range where the predicted age is located.

[0132] Exemplarily, the reference age information of the target clustering cluster where the first voiceprint vector is located includes three age ranges: [0, 12), [13, 18), [18, 150]. Among them, the percentage corresponding to [0, 12) is 0.1, the percentage corresponding to [13, 18) is 0.45, the percentage corresponding to [18, 150] is 0.45, and the largest percentage is 0.45. However, there are two age ranges corresponding to this largest percentage, which are [13, 18) and [18, 150] respectively. At this time, the predicted age in the initial prediction result of the speech to be recognized needs to be obtained. Suppose it is 21 years old. Then the age range where the predicted age is located is [18, 150]. Then, the age label of the age range [18, 150] is determined as the target prediction result of the speech to be recognized.

[0133] Step 554: Determine the average age corresponding to the target clustering cluster based on the reference ages corresponding to the respective reference voiceprint vectors in the target clustering cluster.

[0134] In some embodiments, when the percentages corresponding to each age range are the same, first, obtain the reference ages of all the reference voiceprint vectors from the target clustering cluster. Then, the average age of the target clustering cluster is determined by calculating the average value of the reference ages corresponding to all the reference voiceprint vectors in the target clustering cluster.

[0135] Step 555: Determine the target prediction result of the speech to be recognized based on the average age.

[0136] In some embodiments, when the proportion of the number corresponding to two age intervals is the same, the average age is calculated based on the reference ages corresponding to the respective reference voiceprint vectors in the target clustering cluster, and then this average age is used as the age prediction result of the voice to be recognized for final confirmation. Determining the target prediction result of the voice to be recognized based on the average age may be to determine the age interval in which the average age is located, and then use the age interval in which the average age is located as the target prediction result, or use the age label of the age interval in which the average age is located as the target prediction result. Exemplarily, the predicted age of the voice to be recognized is determined by analyzing the reference age information of the target clustering cluster where the first voiceprint vector is located. Suppose the reference age information includes two age intervals: [0, 18) and [18, 150]. There are two data, 16 years old and 17 years old, in the age interval [0, 18), and one data, 22 years old, in the age interval [18, 150]. If the initial predicted age of the voice to be recognized is 21 years old, then the data in the two age intervals [0, 18) and [18, 150] need to be compared. First, calculate the proportion of the data in the age interval [0, 18) and the proportion of the data in the age interval [18, 150], and both are obtained as 0.5. Next, add up the age data of 16 years old, 17 years old, 21 years old, and 22 years old and take the average, and the result obtained is 19 years old. Since 19 years old belongs to the age interval [18, 150], according to the above logic, the age interval [18, 150] is used as the target prediction result.

[0137] In some embodiments, in addition to the method of steps 1051 to 1055 above, it is also possible to avoid the complexity of comparing one by one with all the vectors in the clustering cluster by comparing the similarity between the first voiceprint vector and the voiceprint vectors of the cluster centers of each clustering cluster. For example, first, obtain the voiceprint vectors of the cluster centers of the multiple clustering clusters corresponding to the object identifier, and these cluster center vectors represent the average voiceprint features of each clustering cluster. Then, compare the similarity between the first voiceprint vector and each cluster center voiceprint vector to find the highest similarity. In the process of calculating the similarity, distance measurement methods such as Euclidean distance or cosine similarity are usually used to measure the similarity between two voiceprint vectors. By calculating the similarity between the first voiceprint vector and each cluster center voiceprint vector, a set of highest similarity values can be obtained. Finally, compare the highest similarity with a predefined threshold. If the highest similarity is greater than the set threshold, it means that the first voiceprint vector has a high similarity with the voiceprint vector of the cluster center of this clustering cluster, that is, this voiceprint vector is considered to belong to this clustering cluster. This voiceprint clustering method based on similarity comparison can effectively reduce the computational complexity and achieve fast and accurate voiceprint clustering. By comparing the similarity between the voiceprint vector and the cluster center voiceprint vector, the voiceprint vector most similar to a specific clustering cluster can be identified, and its belonging can be judged according to the set threshold, which can improve the efficiency and accuracy of the system.

[0138] In the above steps 1051 to 1055, it helps to improve the accuracy and reliability of the voiceprint recognition model. First, by obtaining the object identifier corresponding to the voice to be recognized, personalized voiceprint recognition can be performed for a specific object, avoiding the situation of confusion or misrecognition among multiple people and improving the recognition accuracy. Second, obtaining multiple reference voiceprint vectors corresponding to the object identifier in the pre-constructed voiceprint library and forming at least one reference clustering cluster can more comprehensively reflect the voice characteristics of a specific speaker and improve the system's recognition ability for individual differences. Then, by determining the similarity between the first voiceprint vector and each reference voiceprint vector and determining the target clustering cluster to which the first voiceprint vector belongs based on the similarity, the voice to be recognized can be more accurately classified into the corresponding voiceprint feature group, enhancing the recognition reliability. Finally, by combining the reference age information corresponding to the target clustering cluster and the initial prediction result, the target prediction result of the voice to be recognized is determined. This comprehensive utilization of age information further improves the accuracy and comprehensiveness of voiceprint recognition.

[0139] In some embodiments, before step 101, it is also necessary to obtain a trained voiceprint recognition model through Figure 3F the steps 201 to 204 shown, Figure 3F which is a schematic diagram of the implementation process for training the voiceprint recognition model provided by the embodiments of this application. The following will be specifically described in conjunction with Figure 3F this.

[0140] Step 201: Obtain a first training data set and a preset voiceprint recognition model.

[0141] Among them, the first training data set includes multiple first training voice data and the first label data corresponding to each first training voice data. In some embodiments, the first training data set can be collected from various sources, and these data sets include multiple first training voice data and their corresponding first label data. The first training voice data can come from different voice collection devices or existing voice libraries, and the first label data can be used to indicate the identity information or characteristics corresponding to each voice data. For example, the data set corresponding to the voiceprint recognition model can be aishell2, voxceleb, etc. Among them, the training voice data comes from the data set corresponding to the voiceprint recognition model. In the preset voiceprint recognition model of the embodiments of this application, the label data only contains one annotation information, that is, the identity identifier of the speaker, and does not contain other information such as age and gender.

[0142] Step 202: Use the voiceprint recognition model to perform prediction processing on each first training voice data to obtain the first prediction result corresponding to each first training voice data.

[0143] In some embodiments, to train a voiceprint recognition model, the voiceprint recognition model is used to perform prediction processing on each first training voice data to obtain a first prediction result corresponding to each voice data. First, feature extraction is performed on each first training voice data to convert the voice signal into a numerical feature representation, and these features may include acoustic features such as the spectrum information and formants of the voice. Then, the trained voiceprint recognition model is used to process the extracted features. By learning a large number of labeled voice data, this model can identify and extract unique features in the voiceprint for distinguishing speakers of different age groups. By performing prediction processing on each first training voice data, a first prediction result corresponding to each voice data is obtained, and these results can be a series of numerical values or vectors, representing the recognition prediction made by the voiceprint model for each voice data.

[0144] Step 203: Based on the first loss function corresponding to the voiceprint recognition model, the first prediction results corresponding to each first training voice data, and the first label data corresponding to each first training voice data, determine the first loss value.

[0145] In some embodiments, according to the design of the voiceprint recognition model, a loss function is defined to measure the difference between the prediction result and the actual label data. This loss function can be cross-entropy loss, mean squared error, etc. Through the aforementioned voiceprint recognition model, prediction processing is performed on each first training voice data to obtain the corresponding first prediction results, and these prediction results can be a series of numerical values or vectors, representing the recognition prediction of the model for the voice data. At the same time, there is also the first label data corresponding to each first training voice data, that is, the true identity label of the voice data, and these label data are obtained through manual annotation or other means. During the training process, the difference between the first prediction result and the first label data will be calculated, and the first loss value is obtained through the first loss function. This loss value reflects the error size of the model on the current training samples and can be used as a basis for evaluating the model performance and optimization.

[0146] Step 204: Backpropagate the first loss value to the voiceprint recognition model to adjust the parameters of the voiceprint recognition model until the training end condition is reached, and obtain the trained voiceprint recognition model.

[0147] In some embodiments, through the backpropagation algorithm, the model parameters are adjusted according to the loss value of each training sample, enabling the model to gradually improve its voiceprint recognition ability. During this process, gradient information is utilized, and through continuous iteration and parameter update, the trained voiceprint recognition model is finally obtained.

[0148] At the beginning of training, the parameters of the voiceprint recognition model are initialized. These parameters represent the connection weights between the various layers in the model. The first training voice data is input into the voiceprint recognition model, and the first prediction result is obtained through forward propagation. Then, the first prediction result is compared with the first label data to calculate the first loss value. Using the backpropagation algorithm, starting from the output layer, the first loss value is propagated forward layer by layer to calculate the gradient of each parameter with respect to the loss value. In this process, the server calculates the gradient of each layer according to the chain rule to determine the degree of influence of each parameter on the loss value. Next, according to the calculated parameter gradients, using an optimization algorithm (such as gradient descent), the parameters of the model are adjusted. By reducing the step size in the gradient direction, the parameters of the model are gradually adjusted in the direction that can reduce the loss value. Then, the above process is repeated. The subsequent training voice data is input into the model, and the loss, gradient, and parameter updates are calculated to continuously adjust the model parameters to gradually reduce the loss value on the overall training set. During the training process, an end condition is set, such as reaching a certain number of training epochs or a convergence threshold of the loss value. When the end condition is met, the training process stops, and a trained voiceprint recognition model is obtained.

[0149] In some embodiments, before step 101, it is also necessary to obtain a trained age prediction model through Figure 3G the steps 301 to 305 shown below, Figure 3G which is a schematic diagram of the implementation process for training the age prediction model provided by the embodiments of this application. The following will be specifically described in conjunction with Figure 3G this.

[0150] Step 301: Obtain a second training data set, a trained voiceprint recognition model, and a preset age prediction model.

[0151] Among them, the second training data set includes multiple second training voice data and the corresponding second label data for each second training voice data. The second label data contains the actual age information corresponding to each voice sample. In some embodiments, in order to train the age prediction model, first, a large number of second training voice data are collected, and the corresponding age label data are prepared for each voice data. These voice data and label data will form the second training data set for training the age prediction model. At the same time, a trained voiceprint recognition model is obtained, which can be obtained through the previous voiceprint recognition model training process or obtained as a ready-made voiceprint recognition model through other channels. In addition, it is also necessary to obtain a preset age prediction model, which may involve obtaining an already established age prediction model from a professional research institution or an open-source project, or designing and configuring an age prediction model according to specific requirements. After obtaining these data and models, the training of the age prediction model is carried out.

[0152] Step 302: Obtain each second training speech data for feature extraction to obtain the training speech features corresponding to each second training speech data.

[0153] In some embodiments, in the process of obtaining each second training speech data for feature extraction, based on signal processing and digital signal processing technologies, by performing mathematical operations and transformations on the speech signal, the speech data is converted into a set of numerical features. Among them, the training speech features include the acoustic feature information of each speech sample, such as spectrum, Mel-frequency cepstral coefficients, etc. Specifically, in the process of acoustic feature extraction, a series of signal processing algorithms will be used, such as Fourier transform, cepstrum analysis, linear prediction analysis, etc. These algorithms can analyze features such as the spectrum, pitch, and formants of the speech signal and convert them into a mathematical representation. For example, each second training speech data is segmented into short-time speech frames, and then the Fourier transform is applied to each speech frame to convert the speech signal from the time domain to the frequency domain representation. Then, features such as the spectral envelope and pitch contour of each speech frame are calculated, and the voice features related to age are extracted. In addition to frequency domain analysis, techniques such as linear prediction analysis can also be used to perform more in-depth feature extraction on the speech data. Linear prediction analysis can estimate information such as the formants and vocal tract parameters of the speech signal, thereby capturing the speech features related to the speaker's age.

[0154] Step 303: Use the trained voiceprint recognition model to perform prediction processing on each training speech feature to obtain the training voiceprint vectors corresponding to each second training speech data.

[0155] In some embodiments, in the process of training the age prediction model, the trained voiceprint recognition model is used to perform prediction processing on each training speech feature to obtain the training voiceprint vectors corresponding to each second training speech data. First, the obtained training speech features are used as input data. Second, the already trained voiceprint recognition model is used. This model has learned and trained with a large amount of speech data, can extract unique voiceprint information from the speech features, and at the same time has learned the mapping relationship between the speech features and the speaker's identity. This model may adopt deep learning methods, such as convolutional neural networks or recurrent neural networks, to learn the complex relationship between speech features and voiceprints. By inputting the training speech features into the voiceprint recognition model, prediction processing is performed on each speech feature to obtain the corresponding training voiceprint vectors. These training voiceprint vectors integrate the individual features of the voice and the speaker's identity information. Finally, through the prediction processing of the voiceprint recognition model, the corresponding training voiceprint vectors are generated for each second training speech data.

[0156] Step 304: Use the training speech features, training voiceprint vectors corresponding to each second training speech data, and the second label data to train the age prediction model to obtain the trained age prediction model.

[0157] In some embodiments, it is necessary to combine voice features, voiceprint vectors, and corresponding age labels for an age prediction model to predict the age information of a speaker. During the training process, the model uses training voice features and training voiceprint vectors as input features to learn age-related information contained in the voice data. At the same time, the model also uses the second label data as the training target, and adjusts the model parameters by comparing with the age predicted by the model, so that the model can more accurately predict the age of the speaker corresponding to the voice. This training process usually adopts the supervised learning method in machine learning and may use model structures such as deep neural networks. The model continuously adjusts the parameters through the backpropagation algorithm to minimize the error between the predicted age and the actual age as much as possible. Through multiple iterative trainings, the model gradually improves its learning ability for age information in voice data, and finally obtains a trained age prediction model.

[0158] In some embodiments, referring to Figure 3H , Figure 3H is a schematic diagram of the specific training implementation process of the age prediction model provided by the embodiments of the present application. Figure 3G The step 304 shown can be implemented by steps 3041 to 3044 as shown in Figure 3H . The following will be specifically described in conjunction with Figure 3H .

[0159] Step 3041: Encode the training voice features corresponding to each second training voice data to obtain respective encoding results, and splice each of the encoding results and the training voiceprint vector to obtain respective spliced training vectors.

[0160] In some embodiments, the training voice features and the training voiceprint vector are spliced element by element to form a longer vector. The spliced vector will contain a comprehensive representation of voice features and voiceprint information. Specifically, the corresponding elements of the training voice features and the training voiceprint vector can be spliced one by one in a certain order. For example, for each voice sample, the first elements of its voice features and voiceprint vector are spliced together to form the first element of the spliced vector; then the second elements are spliced, and so on, until all corresponding elements are spliced. In this way, a spliced training vector containing voice features and voiceprint information is obtained. By performing the above processing on each second training voice data, a series of spliced training vectors are obtained, and each vector represents the features of a voice sample.

[0161] Step 3042: Use the age prediction model to perform prediction processing on each spliced training vector to obtain second prediction results for each second training voice data.

[0162] In some embodiments, each spliced training vector is input into a pre-trained deep neural network for forward propagation to perform prediction processing. This deep neural network is usually a model with multiple hidden layers. By learning a large amount of labeled training data, the network can automatically learn the complex non-linear mapping relationship between speech features and age. During the forward propagation process, each spliced training vector will successively undergo non-linear transformations of each hidden layer, calculate and transmit information layer by layer, and finally obtain an output representing the prediction result. This output is usually a vector with multiple dimensions, where each dimension corresponds to a possible age range. The age range corresponding to the dimension with the largest value in the output vector is selected as the second prediction result of the spliced training vector, that is, the age of the speaker inferred by the model based on the input vector.

[0163] Step 3043: Determine the second loss value based on the second loss function corresponding to the age prediction model, the second prediction results corresponding to each second training speech data, and the second label data corresponding to each second training speech data.

[0164] In some embodiments, a second loss function is used to measure the difference between the second prediction result corresponding to each second training speech data and the second label data. This second loss function is usually a function such as the cross-entropy loss function or the mean squared error loss function, which is used to measure the distance between the prediction result and the true label. For each second training speech data, it is input into the trained age prediction model for forward propagation to obtain the corresponding second prediction result. Then, based on the second label data and the second prediction result, the second loss function is used to calculate the second loss value. This process is repeated until the second loss values corresponding to all second training speech data are obtained. During the determination of the second loss value, a deep neural network model is mainly used to perform forward propagation to calculate the second prediction result, and the selected second loss function is combined to measure the error between the prediction result and the true label, thereby obtaining the second loss value. The entire process is completely automated without any human intervention. For example, in the process of determining the second loss value, the method of mean squared error is adopted. This method calculates the difference between the prediction result and the actual label data and squares the difference to obtain an index measuring the prediction accuracy. First, the second prediction result of each second training speech data is compared with the corresponding second label data, which can be achieved by taking the prediction result and the label data as inputs, passing through the calculation process of the model, and obtaining the difference between the model output and the actual value. Then, the loss function is used to measure the difference between the prediction result and the label data. In the age prediction model, the loss function used is the mean squared error, which calculates the square of the difference between the prediction result and the label data and averages the difference values of all samples.

[0165] Step 3044: Backpropagate the second loss value to the age prediction model to adjust the parameters of the age prediction model until the training end condition is reached, obtaining a trained age prediction model.

[0166] In some embodiments, backpropagating the second loss value to the age prediction model to adjust the parameters of the model. Specifically, the backpropagation algorithm is used to calculate the gradient of each parameter with respect to the second loss value. The gradient represents the sensitivity of the parameter to the change in the loss function. By adjusting the parameter in the opposite direction of the gradient, the loss value can be gradually reduced and the prediction accuracy of the model can be improved. To implement backpropagation, first calculate the gradient of the second loss value with respect to the model output, and then calculate the gradients of the parameters of each layer layer by layer according to the chain rule. The automatic differentiation technique is used in this process and no manual intervention is required. Once the gradient of each parameter is obtained, an optimization algorithm (such as stochastic gradient descent) is used to update the value of the parameter, adjusting it in the direction of reducing the loss value. In this way, through multiple iterations, continuously repeating the process of backpropagation and parameter update until the training end condition is reached, that is, the performance of the model cannot be further improved or the predetermined number of training epochs is reached. By adjusting the model parameters through multiple iterations, the age prediction model is gradually optimized and a trained model is obtained.

[0167] In the above steps 3041 to 3043, first, by concatenating the features of the second training speech data and the training voiceprint vector, a concatenated training vector is obtained, realizing the effective integration of the voiceprint information and the speech features, thereby improving the model's comprehensive processing ability of the voiceprint features and speech information. Secondly, using the age prediction model to perform a prediction process on the concatenated training vector, obtaining the second prediction result of the second training speech data, realizing the prediction of the age information in the speech data, and providing a basis for subsequent parameter adjustment. Then, based on the second loss function, the second prediction result and the second label data corresponding to the age prediction model, the second loss value is determined, thereby evaluating the difference between the prediction result and the true label, and providing guidance for the optimization of the model. Finally, backpropagate the second loss value to the age prediction model to adjust the parameters of the age prediction model. Through multiple iterations of training, a trained age prediction model is finally obtained, enabling it to more accurately predict the age information corresponding to the speech data. In summary, the method of steps 3041 to 3043 can effectively improve the accuracy and robustness of the age prediction model by integrating voiceprint and speech features, performing age prediction and loss function evaluation, and backpropagating parameter adjustment.

[0168] In the above manner, embodiments of the present application can extract voiceprint features and perform prediction processing on the voice to be recognized by using a trained voiceprint recognition model. During this process, the voiceprint recognition model can convert the voice to be recognized into a first voiceprint vector, which can effectively distinguish the differences between different speakers and improve the accuracy of speaker identification. Secondly, by using the trained age prediction model, prediction processing can be performed on the first voiceprint vector and the voice features to be recognized to obtain an initial prediction result. The age prediction model can predict the age range of the speaker based on the voice features and speech content. By combining voiceprint information and age prediction, the identity and age characteristics of the speaker of the voice to be recognized can be more comprehensively understood. During this process, by establishing a voiceprint database, associating the voiceprint features of each object with its corresponding identity information, and using the voiceprint database for age judgment can improve the efficiency of age prediction. The establishment of the voiceprint database can also enable the model to better understand each voiceprint feature and the corresponding age information. Finally, based on the initial prediction result, the target prediction result of the voice to be recognized can be further determined. By combining multiple information sources, the accuracy and reliability of the prediction result can be improved. Therefore, through the present application, the accuracy of voice recognition in the anti-addiction system can be improved.

[0169] Next, an exemplary application of embodiments of the present application in a practical application scenario will be described.

[0170] The voice recognition method provided by embodiments of the present application can be applied to scenarios that require age recognition detection. For example, it can be applied to the scenario of minor identification. In response to the country's protection policy for minors and to prevent the harm caused by games to minors, voices can be collected during the game session of players, and the age of the players can be predicted by using the voice recognition method provided by embodiments of the present application. If it is determined that the player is a minor, their game time can be restricted, thereby providing help for the protection of minors. If it is determined that the player is an elderly person, their game time can also be restricted to prevent the elderly from being addicted to games. The voice recognition method provided by embodiments of the present application can also be applied to the scenario of malicious refund. For example, after a player recharges in a game application and requests a refund on the grounds that the recharge process was operated by a minor, the voices of the player before and after recharge can be obtained, and age detection can be performed based on the player's voice to determine the player's age and voiceprint information to verify the authenticity of the refund reason and prevent players from maliciously refunding.

[0171] See Figure 4 , Figure 4 which is another schematic diagram of the implementation process of the voice recognition method provided by embodiments of the present application. Next, specific descriptions will be made in combination with the steps shown in Figure 4 shown below.

[0172] Step 401: Data preprocessing.

[0173] Before training the voiceprint model and age prediction model, it is necessary to preprocess the two datasets of the voiceprint model and age prediction model.

[0174] First, for the training datasets of the voiceprint recognition model such as aishell2, voxceleb2, cn-celeb, etc., the following preprocessing steps are required: First, sample and extract features from the original speech data to convert the speech signal into a numerical feature representation, such as Mel Frequency Cepstral Coefficients or Mel spectrogram, etc. Then, perform speaker identification processing on the training dataset of the voiceprint recognition model to ensure that each speech sample is associated with the corresponding speaker ID. Then, expand the training dataset through data augmentation techniques to increase the diversity and richness of the data, including methods such as variable speed, adding noise, and speech rate expansion. Finally, perform data normalization and standardization processing to ensure that the data input into the voiceprint recognition model has a consistent feature distribution and range.

[0175] Second, for the training datasets of the age prediction model such as common voice, datatang-200h Chinese dataset, etc., a similar preprocessing process is also required: Similarly, it is necessary to sample and extract features from the original speech data to convert the speech signal into a numerical feature representation. Then, perform age annotation processing on the speech samples in the training dataset to ensure that each speech sample is annotated with the corresponding age information. Then, data augmentation techniques can also be used to expand the training dataset to increase the diversity and richness of the data. Finally, perform data normalization and standardization processing to ensure that the data input into the age prediction model has a consistent feature distribution and range.

[0176] Step 402: Use the voiceprint recognition model to extract the voiceprint vector of the preprocessed speech data.

[0177] The network structure of the voiceprint recognition model mainly consists of the ECAPA-TDNN network, see Figure 5 , Figure 5 which is the flow diagram of the ECAPA-TDNN network provided by the embodiment of the present application. The ECAPA-TDNN network is an optimized TDNN (Time Delay Neural Network) structure, with the characteristics of a time delay neural network, and integrates the self-attention mechanism and the Squeeze-Excitation (SE) module to learn the key temporal information in the input features.

[0178] In addition to the three-layer SE-Res2Block part, the ECAPA-TDNN network structure also includes an input layer, a frame-level feature extraction layer, a TDNN layer, a mean and variance normalization layer, a dimensionality reduction layer, an activation function layer, a pooling layer, a final classification layer, an Attentive-stat pool part, and a loss function part.

[0179] Among them, the input layer receives acoustic features as the input of the model, and the frame-level feature extraction layer preprocesses and extracts features from the input speech signal. The TDNN layer is the core layer of the ECAPA-TDNN network. It captures the temporal information in the input features through filters with different time delays and realizes the temporal modeling of features through convolutional operations. The mean and variance normalization layer normalizes the mean and variance of the features output by the TDNN layer to enhance the robustness and distinguishability of the features. The dimensionality reduction layer uses linear transformation or projection operations to map high-dimensional features to a low-dimensional space to reduce the computational complexity and improve the generalization ability of the model. The activation function layer performs a non-linear transformation on the dimensionality-reduced features. Common activation functions include ReLU, etc. The SE-Res2Block part is composed of multiple SE modules, introducing a channel attention mechanism and a temporal self-attention mechanism to enhance the feature representation ability and robustness. The pooling layer is used to perform pooling operations on the features, such as global average pooling or statistical pooling, to extract global statistical information and reduce the feature dimension. The final classification layer usually uses a fully connected layer or a softmax layer to map the extracted features to the final classification result and output the probability distribution corresponding to different speaker IDs.

[0180] The Attentive-stat pooling part introduces a self-attention mechanism, enabling the neural network to focus in the time dimension and accumulate the information of different channels on the time axis. Specifically, this part performs weighted accumulation on the features of different channels through weighted average and weighted variance, thereby enhancing the robustness and distinguishability of the model for voiceprint features. This adaptive weighted accumulation can better capture the dynamic changes and speech expression details of voiceprint features and improve the performance of voiceprint recognition.

[0181] In the loss function part, the voiceprint recognition model adopts the AAM-softmax loss function. This loss function is a loss function designed for classification tasks. It adjusts the size of the angle by setting constants s and m, as shown in formula 1:

[0182]

[0183] The AAM-softmax loss function can reduce the angle between samples of the same class while increasing the angle θ between samples of different classes, thereby enhancing the distinguishability between voiceprint features. Here, N represents the total number of classes, s is a constant used to control the angular margin between classes, m is a constant used to control the angular margin within a class, i ranges from 1 to N, representing the i-th sample. j also ranges from 1 to N, but excludes the true class to which the current sample belongs. By optimizing this loss function, the voiceprint recognition model can better learn the voiceprint features of the speaker and achieve a more accurate voiceprint recognition task.

[0184] These layers are interconnected through forward propagation and backward propagation. Forward propagation passes the input features from one layer to the next and performs calculations and transformations through the parameters of each layer. Backward propagation calculates the gradients based on the loss function and updates the parameters using the gradients, enabling the network to gradually optimize and learn effective voiceprint recognition features. Through this collaborative effect between layers, the ECAPA-TDNN network can fully exploit the temporal information and channel correlation in acoustic features to accurately model and recognize the voiceprint features of the speaker.

[0185] See Figure 6 , Figure 6This is a schematic flowchart of the SE-Res2Block part provided by the embodiments of this application. The SE-Res2Block consists of a channel attention module (Squeeze Excitation, SE), and the SE module includes: a channel attention unit, a temporal self-attention unit, and a residual connection unit. First, in the channel attention unit, the importance weights of each channel are learned. It uses global average pooling to convert the feature map into global statistical information, then performs nonlinear transformation and adjustment through two fully connected layers, and finally uses the sigmoid function to limit the obtained weights between 0 and 1. In this way, the weights of each channel can be dynamically adjusted according to their important contribution degrees to the feature representation. Next, in the temporal self-attention unit, the self-attention mechanism is used to capture the correlation of features in the time dimension. It maps the temporal features to an attention weight vector through global average pooling and fully connected layers, and this vector represents the importance of each time step. Then, through nonlinear transformation and adjustment of the attention weight vector, and multiplying it with the original temporal features, the weighted temporal features are obtained. Finally, in the residual connection unit, the weighted temporal features are element-wise added to the original input features to retain the information of the original features. This residual connection can effectively alleviate the problem of gradient disappearance and promote information flow, which helps to extract more discriminative voiceprint features. By introducing the SE-Res2Block part, the ECAPA-TDNN network can adaptively adjust the channel weights and temporal weights to capture important voiceprint features and suppress noise and redundant information. This can improve the model's ability to model the speaker's voiceprint and enhance the accuracy and robustness of voiceprint recognition.

[0186] Step 403: Perform age prediction using the age prediction model.

[0187] The age prediction model can be a deep learning network model. The structure of the age prediction model includes a feature extraction module, a conformer module, and a fully connected layer module. When training the age prediction model, first use the feature extraction module to extract the Mel features of the training speech, and obtain the 192-dimensional voiceprint vector of the training speech in the dataset corresponding to the age model as the input. These features will be input into the conformer module for model training. In the previous layer of the network, the voiceprint vector will be merged with the nodes of the current network to make full use of the voiceprint information and combine other features. The entire network is trained using a regression model, where the label is the specific age. During the training process, the age prediction model will gradually adjust the network parameters through an optimization algorithm, enabling the model to accurately predict the age information corresponding to the input speech. After obtaining the trained age prediction model, it can be applied to game applications. For example, in a game, obtain the voice data of the player, and then use the trained model to identify and predict the age of the player, so as to realize functions such as personalizing the game experience according to age characteristics or performing age-related data statistical analysis. See Figure 7 , Figure 7 FIG. Figure 7 is a schematic diagram of the implementation process of age prediction using the age prediction model provided by the embodiment of the present application, Figure 4 The step 403 shown can be implemented by steps 4031 to 4034 as shown in Figure 7 FIG. Figure 7 . The following will be specifically described in conjunction with Figure 7 FIG. Figure 7 .

[0188] Step 4031: Feature extraction.

[0189] In some embodiments, the speech data is subjected to feature extraction to obtain speech features. During the feature extraction process, the processing of the speech features in the game speech involves six main parts. The feature extraction process is as shown in Figure 8 FIG. Figure 8 , Figure 8 which is a schematic diagram of the structure of feature extraction provided by the embodiment of the present application. The features used in feature extraction can be Mel spectrum features. For the extraction process of speech features, it includes six parts: pre-emphasis of the signal, framing, windowing, Fourier transform, Mel filtering, and taking the logarithm.

[0190] First is the pre-emphasis of the signal. This step aims to emphasize the high-frequency part, reduce the high-frequency attenuation phenomenon in the speech signal, and improve the signal-to-noise ratio and speech quality. Before audio signal processing, a high-pass filter such as Equation 2 is usually used to pre-emphasize the signal. The purpose of pre-emphasis is to balance the spectrum and highlight the high frequency. The time-domain expression is as shown in Equation 3, and α is generally taken as 0.97:

[0191] H(z) = 1 - μz -1 (2)

[0192] y(n) = x(n) - αx(n - 1) (3)

[0193] Then comes windowing. By multiplying the speech signal of each frame by a window function (such as the Hamming window), the boundaries are smoothed to avoid spectral leakage and reduce spectral fluctuations. After frame segmentation, to achieve a smooth transition between adjacent frames, that is, to eliminate the signal discontinuity (i.e., spectral leakage) that may occur at both ends of each frame, the window function can reduce the impact of truncation. We apply each frame to the window function, and the windowed speech signal sw(n) = s(n) * w(n). The commonly used window in speech processing is the Hamming window, and the formula for the Hamming window is shown in Equation 4, where N is the window length, 0 ≤ n ≤ N - 1, and a is a constant:

[0194]

[0195] Next is the Fourier transform, which converts the windowed speech signal of each frame to the frequency domain and obtains its spectral information. The fast Fourier transform converts the time-domain signal to the frequency domain, and the formula is shown in Equation 5, where N is the number of FFT points:

[0196]

[0197] Immediately following is Mel filtering. A set of triangular filters is used to filter the signal after Fourier transform to simulate the sensitivity of the human ear to different frequencies and extract features related to human hearing. The spectral signal is passed through the triangular filters on the Mel scale, and the frequency response of the triangular filter is shown in Equation 6, where f(m) is the center frequency of the filter, and M usually takes values between 22 and 26:

[0198]

[0199] Where:

[0200]

[0201] Finally, taking the logarithm of the filtered signal to narrow the dynamic range, enhance the signal-to-noise ratio, and obtain the final Mel spectral features as the input features for training the age prediction model. These feature extraction steps can effectively extract the key information in the speech.

[0202] Step 4032: Use the conformer module to encode the speech features.

[0203] Figure 9 This is the specific structural schematic diagram of the conformer module provided by the embodiments of the present application. In Figure 9 In, Figure 9 The left figure in shows the entire structure of the conformer module, Figure 9The right figure in [description] is a schematic structural diagram of the conformer layer. Figure 9 The left figure in [description] respectively includes a convolutional layer 901, a fully connected layer 902, a dropout layer 903, and N conformer layers 904. The convolutional layer 901 is used for downsampling, the fully connected layer 902 is used for dimensional transformation and providing non-linearity, and the dropout layer 903 is used to prevent overfitting. Figure 9 The right figure in [description] is the internal structure of the conformer layer 904, which respectively includes a forward propagation sub-module 9041, a multi-head attention module 9042, a convolutional layer 9043, and a normalization layer 9044. The main modules are the forward propagation sub-module (feed forward module, FFM) 9041 and the multi-head self-attention (Multi-head Self-Attention, MSA) module 9042, as shown in Figure 10 and Figure 11 shown. The role of the forward propagation sub-module is to provide non-linear changes and enhance the robustness of the model. The forward propagation sub-module is as shown in Figure 10 shown. Figure 10 This is the schematic structural diagram of the forward propagation sub-module provided by the embodiment of the present application. This module is mainly composed of two convolutional layers and a Swish activation function, and is used for non-linear transformation and mapping of the input features. This can help the network learn richer and more complex feature representations to improve the performance of classification or regression tasks. The forward propagation sub-module can introduce non-linear transformation and sequence information during the feature extraction process, thereby enhancing the expression ability and context modeling ability of the model; the multi-head attention module is as shown in Figure 11 shown. Figure 11 This is the schematic structural diagram of the multi-head attention module provided by the embodiment of the present application. Through the self-attention mechanism provided by the multi-head attention module, the neural network can fully learn the forward and backward correlations in the audio signal, enabling the model to focus on the features that are more beneficial to the result and less on useless features. The role of the multi-head attention mechanism here is to allocate different spaces to potential different features, so that different features can be captured during training and the expression ability of the features can be improved.

[0204] Step 4033: Perform a fully connected process on the voiceprint vector and the output vector of the conformer module using the fully connected layer.

[0205] The fully connected layer is a common neural network layer in the age prediction model. It consists of multiple neurons, and each neuron is connected to all the inputs of the previous layer. The goal of the fully connected layer is to perform linear transformation and non-linear mapping on the features of the previous layer to generate a higher-level abstract representation. In the age prediction model, the fully connected layer is used to integrate and compress the features extracted by each previous module to capture more representative feature representations. By learning appropriate weight parameters, the fully connected layer can map these features to a lower-dimensional feature space with better discriminative ability.

[0206] Step 4034: Output the age recognition result.

[0207] The output module is the last part of the age prediction model. It receives the feature representation from the fully connected layer and maps it to the final age recognition result. Usually, the output module uses one or more fully connected layers to perform classification or regression tasks. For classification problems, this module usually contains a fully connected layer with a softmax activation function to calculate the probability distribution of different age categories. The model will select the most likely age category as the final recognition result based on the predicted probability. For regression problems, the output module may contain a fully connected layer with a linear activation function to directly predict a continuous value of age. By training and optimizing the parameters of these fully connected layers, the output module can convert the feature representation into an age recognition result, thus completing the entire age recognition task.

[0208] The embodiments of this application combine voiceprint technology and age recognition well, making age recognition more accurate. In addition, voiceprint technology can also assist in the identification of minors, making minor identification more accurate.

[0209] In addition, after training the voiceprint recognition model and the age prediction model, a voiceprint library needs to be established. The voiceprint library is a database that stores the voiceprint features of objects and is used for voiceprint recognition tasks. It contains the voiceprint vectors of each object and the corresponding predicted ages.

[0210] First, a batch of voice data of a group of objects needs to be collected. These voice data should contain different speech contents of each object. For each object, at least several seconds of voice samples are required. Then, through the voiceprint model, the collected voice data is converted into voiceprint vectors. A voiceprint vector is a high-dimensional vector that represents the voice characteristics of the object. Associating the voiceprint vector of each object with its corresponding object identifier allows the voiceprint vectors in the voiceprint library to be indexed by the object identifier. When there is a new voice sample for voiceprint recognition, it is converted into a voiceprint vector. Then, using cosine similarity or other similarity measurement methods, the similarity between the new voiceprint vector and all the voiceprint vectors in the voiceprint library is calculated. The result range of cosine similarity calculation is between [-1, 1]. The closer the value is to 1, the more similar the two voiceprint vectors are. Set a similarity threshold, which is determined according to actual needs and application scenarios. When the similarity between the new voiceprint vector and a certain voiceprint vector in the voiceprint library exceeds the threshold, it can be considered that they belong to the same object. When the similarity between the new voiceprint vector and the voiceprint vectors of this object in the voiceprint library does not exceed the threshold, a new cluster needs to be established with the new voiceprint vector, and the predicted age is added through the age prediction model. For the object identifier corresponding to the voiceprint vectors similar to the new voiceprint vector, the corresponding predicted age can be obtained. Based on these predicted ages, the age distribution of this cluster can be observed, and it can be judged whether this cluster is of minors. If the proportion of minors in the age distribution is relatively high, it can be determined that the object corresponding to the new voiceprint vector is a minor.

[0211] It should be noted that when establishing the voiceprint library, the situation where multiple people may use the same object identifier needs to be considered. Therefore, when establishing the voiceprint library, there may be multiple different voiceprint vectors under the same object identifier, and each similar voiceprint vector will generate a cluster. After real-time voice enters, by calculating the similarity with the similar voiceprint vectors in the voiceprint library, the age corresponding to the new voiceprint vector can be determined. First, observe the proportion of ages. If the age proportions are the same, the age of the new voiceprint vector is calculated according to the average age. In short, the establishment and use of the voiceprint library are to associate the object identifier with the voiceprint vector, and then predict the age of the speaker of the input voice through similarity calculation. This method can effectively perform voiceprint recognition and age judgment tasks.

[0212] In summary, by training a voiceprint recognition model, generating voiceprint vectors, and training an age prediction model in combination with voice data, accurate recognition of the speaker's identity and age can be achieved. When constructing a voiceprint library, there are multiple similar voiceprint vectors under each object identifier, forming a cluster, which can better handle voice changes in different environments and improve the robustness and accuracy of voiceprint recognition. During application, when real-time voice enters the system, the voiceprint recognition model and the age prediction model will calculate the real-time voiceprint vector and the specific age respectively. Then, by calculating the similarity with the similar voiceprint vectors in the voiceprint library, the age corresponding to the new voiceprint vector can be determined. In addition, the proportion of ages is observed preferentially. If the age proportions are the same, the age of the new voiceprint vector is calculated according to the average age.

[0213] It can be understood that in the embodiments of the present application, data related to user information, user voice, etc. are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0214] Next, the implementation of the voice recognition device 455 provided by the embodiments of the present application as an exemplary structure of software modules will be continued. In some embodiments, as Figure 2 shown, the software modules stored in the voice recognition device 455 in the memory 450 may include:

[0215] A first acquisition module 4551, configured to acquire the voice to be recognized, a trained voiceprint recognition model, and a trained age prediction model, where the trained age prediction model is obtained by training using training voiceprint vectors and training voice features;

[0216] A first extraction module 4552, configured to extract features from the voice to be recognized to obtain voice features to be recognized;

[0217] A first prediction module 4553, configured to perform prediction processing on the voice features to be recognized by using the trained voiceprint recognition model to obtain a first voiceprint vector corresponding to the voice to be recognized;

[0218] A second prediction module 4554, configured to perform prediction processing on the first voiceprint vector and the voice features to be recognized by using the trained age prediction model to obtain an initial prediction result;

[0219] An output module 4555, configured to determine a target prediction result of the voice to be recognized based on the initial prediction result and output the target prediction result.

[0220] In some embodiments, the speech recognition device further includes: a second acquisition module, configured to acquire a first training data set and a preset voiceprint recognition model, where the first training data set includes a plurality of first training voice data and first label data corresponding to each of the first training voice data; a first prediction module, configured to perform prediction processing on each of the first training voice data by using the voiceprint recognition model to obtain a first prediction result corresponding to each of the first training voice data; a first determination module, configured to determine a first loss value based on a first loss function corresponding to the voiceprint recognition model, the first prediction result corresponding to each of the first training voice data, and the first label data corresponding to each of the first training voice data; a backpropagation module, configured to backpropagate the first loss value to the voiceprint recognition model to adjust parameters of the voiceprint recognition model until a training end condition is reached, so as to obtain a trained voiceprint recognition model.

[0221] In some embodiments, the speech recognition device further includes: a third acquisition module, configured to acquire a second training data set, a trained voiceprint recognition model, and a preset age prediction model, where the second training data set includes a plurality of second training voice data and second label data corresponding to each of the second training voice data; a second extraction module, configured to perform feature extraction on each of the second training voice data to obtain training voice features corresponding to each of the second training voice data; a second prediction module, configured to perform prediction processing on each of the training voice features by using the trained voiceprint recognition model to obtain a trained voiceprint vector corresponding to each of the second training voice data; a training module, configured to train the age prediction model by using the training voice features, the trained voiceprint vector, and the second label data corresponding to each of the second training voice data, so as to obtain a trained age prediction model.

[0222] In some embodiments, the training module is further configured to perform encoding processing on the training voice features corresponding to each of the second training voice data to obtain respective encoding results, and perform splicing processing on each of the encoding results and the trained voiceprint vector to obtain respective spliced training vectors; perform prediction processing on each of the spliced training vectors by using the age prediction model to obtain a second prediction result corresponding to each of the second training voice data; determine a second loss value based on a second loss function corresponding to the age prediction model, the second prediction result corresponding to each of the second training voice data, and the second label data corresponding to each of the second training voice data; backpropagate the second loss value to the age prediction model to adjust parameters of the age prediction model until a training end condition is reached, so as to obtain a trained age prediction model.

[0223] In some embodiments, the output module 4555 is further configured to obtain the object identifier corresponding to the voice to be recognized; obtain a plurality of reference voiceprint vectors corresponding to the object identifier from a pre-constructed voiceprint library, and the plurality of reference voiceprint vectors form at least one reference clustering cluster; determine the similarity between the first voiceprint vector and each of the reference voiceprint vectors; based on each of the similarities, determine the target clustering cluster to which the first voiceprint vector belongs, and obtain the reference age information corresponding to the target clustering cluster; and determine the target prediction result of the voice to be recognized based on the reference age information and the initial prediction result.

[0224] In some embodiments, the output module 4555 is further configured to determine the maximum similarity from each of the similarities; when the maximum similarity is greater than a preset similarity threshold, obtain the reference voiceprint vector corresponding to the maximum similarity; and determine the clustering cluster where the reference voiceprint vector corresponding to the maximum similarity is located as the target clustering cluster to which the first voiceprint vector belongs.

[0225] In some embodiments, the output module 4555 is further configured to determine the maximum number ratio when the number ratios corresponding to each of the age intervals are different; and determine the target prediction result of the voice to be recognized based on the age interval corresponding to the maximum number ratio and the initial prediction result.

[0226] In some embodiments, the output module 4555 is further configured to, when the number ratios corresponding to each of the age intervals are the same, determine the average age corresponding to the target clustering cluster based on the reference ages corresponding to each of the reference voiceprint vectors in the target clustering cluster; and determine the target prediction result of the voice to be recognized based on the average age.

[0227] In some embodiments, the voice recognition device further includes: a first adding module, configured to add the first voiceprint vector to the target clustering cluster; a first updating module, configured to determine the first age interval to which the predicted age belongs, and update the number of first voiceprint vectors corresponding to the first age interval; a fourth obtaining module, configured to obtain the number of second voiceprint vectors corresponding to a second age interval in the reference age information; and a second updating module, configured to update the first number ratio of the first age interval and the second number ratio of the second age region based on the updated number of first voiceprint vectors and the number of second voiceprint vectors.

[0228] In some embodiments, the output module 4555 is further configured to, when the maximum similarity is less than or equal to the similarity threshold, create a new clustering cluster; and determine the newly created clustering cluster as the target clustering cluster to which the first voiceprint vector belongs.

[0229] In some embodiments, the voice recognition device further includes: a second increasing module, configured to add the first voiceprint vector to the target clustering cluster; and a second determining module, configured to determine reference age information of the target clustering cluster based on an initial prediction result corresponding to the voice to be recognized.

[0230] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the voice recognition method described above in the embodiments of the present application.

[0231] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, where computer-executable instructions or a computer program are stored, and when the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the voice recognition method provided in the embodiments of the present application. For example, as Figures 3A to 3H shown in the voice recognition method.

[0232] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0233] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted language, or declarative or procedural language), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0234] As an example, the computer-executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that stores other programs or data, for example, stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).

[0235] As an example, the computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0236] In summary, the voiceprint recognition model trained through the embodiments of the present application extracts and predicts the voiceprint features of the voice to be recognized, obtaining the first voiceprint vector to distinguish different speakers and improve the identity accuracy. In combination with the age prediction model, the voiceprint vector and voice features are predicted to obtain the initial prediction result, so as to comprehensively understand the identity and age features of the speaker. An object voiceprint library is established to associate the voiceprint features and identity information, improving the age prediction efficiency and enabling the model to better understand the voiceprint features and age information. Finally, by combining multiple information sources, the target prediction result is determined to improve the prediction accuracy and reliability.

[0237] The above description is only for the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A speech recognition method, characterized in that, The method includes: Obtaining the speech to be recognized, a trained voiceprint recognition model, and a trained age prediction model, where the trained age prediction model is obtained by training using training voiceprint vectors and training speech features; Performing feature extraction on the speech to be recognized to obtain speech features to be recognized; Performing prediction processing on the speech features to be recognized by using the trained voiceprint recognition model to obtain a first voiceprint vector corresponding to the speech to be recognized; Performing prediction processing on the first voiceprint vector and the speech features to be recognized by using the trained age prediction model to obtain an initial prediction result; Determining a target prediction result of the speech to be recognized based on the initial prediction result, and outputting the target prediction result.

2. The method according to claim 1, characterized in that The method further includes: Obtaining a first training data set and a preset voiceprint recognition model, where the first training data set includes multiple first training speech data and first label data corresponding to each of the first training speech data; Performing prediction processing on each of the first training speech data by using the voiceprint recognition model to obtain a first prediction result corresponding to each of the first training speech data; Determining a first loss value based on a first loss function corresponding to the voiceprint recognition model, the first prediction results corresponding to each of the first training speech data, and the first label data corresponding to each of the first training speech data; Backpropagating the first loss value to the voiceprint recognition model to adjust parameters of the voiceprint recognition model until a training end condition is reached, thereby obtaining a trained voiceprint recognition model.

3. The method according to claim 2, characterized in that, The method further includes: Obtaining a second training data set, a trained voiceprint recognition model, and a preset age prediction model, where the second training data set includes multiple second training speech data and second label data corresponding to each of the second training speech data; Performing feature extraction on each of the second training speech data to obtain training speech features corresponding to each of the second training speech data; Performing prediction processing on each of the training speech features by using the trained voiceprint recognition model to obtain a training voiceprint vector corresponding to each of the second training speech data; Training the age prediction model by using the training speech features, the training voiceprint vectors, and the second label data corresponding to each of the second training speech data, thereby obtaining a trained age prediction model.

4. The method according to claim 3, characterized in that, The training the age prediction model by using the training speech features, the training voiceprint vectors, and the second label data corresponding to each of the second training speech data to obtain a trained age prediction model includes: Performing encoding processing on the training speech features corresponding to each of the second training speech data to obtain respective encoding results, and performing splicing processing on each of the encoding results and the training voiceprint vectors to obtain respective spliced training vectors; Performing prediction processing on each of the spliced training vectors by using the age prediction model to obtain a second prediction result corresponding to each of the second training speech data; Determine a second loss value based on the second loss function corresponding to the age prediction model, the second prediction results corresponding to each of the second training voice data, and the second label data corresponding to each of the second training voice data; Backpropagate the second loss value to the age prediction model to adjust the parameters of the age prediction model until a training end condition is reached, and obtain a trained age prediction model.

5. The method according to any one of claims 1 to 4, characterized in that, The determining the target prediction result of the voice to be recognized based on the initial prediction result includes: Obtain the object identifier corresponding to the voice to be recognized; Obtain a plurality of reference voiceprint vectors corresponding to the object identifier from a pre-constructed voiceprint library, and the plurality of reference voiceprint vectors form at least one reference clustering cluster; Determine the similarity between the first voiceprint vector and each of the reference voiceprint vectors; Based on each of the similarities, determine the target clustering cluster to which the first voiceprint vector belongs, and obtain the reference age information corresponding to the target clustering cluster; Based on the reference age information and the initial prediction result, determine the target prediction result of the voice to be recognized.

6. The method according to claim 5, wherein The determining the target clustering cluster to which the first voiceprint vector belongs based on each of the similarities includes: Determine the maximum similarity from each of the similarities; When the maximum similarity is greater than a preset similarity threshold, obtain the reference voiceprint vector corresponding to the maximum similarity; Determine the clustering cluster where the reference voiceprint vector corresponding to the maximum similarity is located as the target clustering cluster to which the first voiceprint vector belongs.

7. The method according to claim 6, characterized in that, The reference age information of the target clustering cluster includes at least two age intervals and the proportion of the number corresponding to each age interval, and the determining the target prediction result of the voice to be recognized based on the reference age information and the initial prediction result includes: When the proportion of the number corresponding to each age interval is different, determine the maximum proportion of the number; Based on the age interval corresponding to the maximum proportion of the number and the initial prediction result, determine the target prediction result of the voice to be recognized.

8. The method according to claim 7, wherein The determining the target prediction result of the voice to be recognized based on the reference age information and the initial prediction result includes: When the proportion of the number corresponding to each age interval is the same, determine the average age corresponding to the target clustering cluster based on the reference ages corresponding to each reference voiceprint vector in the target clustering cluster; Based on the average age, determine the target prediction result of the voice to be recognized.

9. The method according to claim 6, wherein The initial prediction result includes a predicted age, and the method further includes: Add the first voiceprint vector to the target clustering cluster; Determine the first age interval to which the predicted age belongs, and update the number of first voiceprint vectors corresponding to the first age interval; Obtain the number of second voiceprint vectors corresponding to a second age interval in the reference age information, where the second age interval is other age intervals in the reference age information except the first age interval; Based on the updated number of first voiceprint vectors and the number of second voiceprint vectors, update the first proportion of the number in the first age interval and the second proportion of the number in the second age region.

10. The method according to claim 6, characterized in that, Determining the target clustering cluster to which the first voiceprint vector belongs based on each of the similarities includes: When the maximum similarity is less than or equal to the similarity threshold, a new clustering cluster is created; The newly created clustering cluster is determined as the target clustering cluster to which the first voiceprint vector belongs.

11. The method according to claim 10, characterized in that, The method further includes: Adding the first voiceprint vector to the target clustering cluster; Based on the initial prediction result corresponding to the voice to be recognized, determining the reference age information of the target clustering cluster.

12. A voice recognition device, characterized in that, The device includes: A first acquisition module, configured to acquire a voice to be recognized, a trained voiceprint recognition model, and a trained age prediction model, where the trained age prediction model is trained using training voiceprint vectors and training voice features; A first extraction module, configured to perform feature extraction on the voice to be recognized to obtain voice features to be recognized; A first prediction module, configured to perform prediction processing on the voice features to be recognized using the trained voiceprint recognition model to obtain a first voiceprint vector corresponding to the voice to be recognized; A second prediction module, configured to perform prediction processing on the first voiceprint vector and the voice features to be recognized using the trained age prediction model to obtain an initial prediction result; An output module, configured to determine a target prediction result of the voice to be recognized based on the initial prediction result, and output the target prediction result.

13. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions; A processor, configured to implement the method according to any one of claims 1 to 11 when executing the computer-executable instructions stored in the memory.

14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or computer program, when executed by a processor, implement the method according to any one of claims 1 to 11.

15. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or computer program, when executed by a processor, implement the method according to any one of claims 1 to 11.

Citation Information

Cited By

  • Speech recognition method, terminal equipment and storage medium

    CN119107955A

  • Speech recognition method, terminal device and storage medium

    CN119107955B