Voice object recognition method and device, computer equipment and storage medium
By using the trained speech object recognition model in the speech object verification technology to perform feature extraction and difference processing on the recognized speech, and combining position coding and multi-layer perceptron calculation verification scores, the accuracy problem of the prior art in multi-person speech mixing or noise interference scenarios is solved, achieving higher robustness and accuracy.
Patent Information
- Application Number
- CN202510251872.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
AI Technical Summary
The existing voice object verification technology has shown a significant decline in scenarios where multiple people's voice mixing or severe noise interference, making it difficult to accurately distinguish and identify the target voice object, resulting in a significant reduction in verification accuracy.
The trained speech object recognition model is used to extract features for recognized speech. By subtracting the extracted speech feature vector from the registered feature vector, the target difference vector is obtained, and position encoding is added to it, input it into the encoder for processing, and the target difference vector encoding is generated. Then, the verification score is calculated using a multi-layer perceptron and the preset activation function to determine the probability that the target speech object appears in the speech to be recognized.
It significantly improves the robustness and accuracy of speech object verification, can more accurately quantify the possibility of target speech objects appearing in test speech, and effectively handle speech signals in complex environments.
Smart Images

Figure CN120089129A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and is applicable to the medical and financial fields. In particular, it relates to a method, device, computer device and storage medium for voice object recognition. Background Art
[0002] In the field of voice object verification, current mainstream technologies usually convert variable-length voice signals into fixed-length voice object vectors, and then complete the verification process by calculating the cosine similarity between the embedding vectors corresponding to two voice segments. However, this technology shows obvious limitations when facing complex environmental challenges. Especially in scenarios of multi-person voice mixing or severe noise interference, its performance will decline significantly. Although current voice object verification technologies, such as the x-vector model, can show good verification effects under clear voice conditions, they are difficult to accurately distinguish and identify target voice objects when encountering multi-person voice mixing or a noisy background, resulting in a significant reduction in verification accuracy.
[0003] For example, in the scenario of medical voice recognition and verification, especially in a noisy hospital environment, such as an emergency room or an intensive care unit, there are multiple overlapping voices or severe background noise generated by medical equipment. Verifying the patient's identity or the content of medical orders by calculating the cosine similarity between the feature vectors of two voices is difficult to accurately distinguish and identify the target voice due to the overlapping multi-person voices or high-noise background in the hospital environment. In the scenario of remote financial services, customers may be in a noisy public place or conduct transactions through a telephone conference. At this time, the voice signal may be mixed with the voices of others or severely interfered by telephone line noise. If the cosine similarity between voice feature vectors is calculated to verify the customer's identity or authorization instructions, in the face of multi-person voice mixing or a high-noise background, its ability to accurately distinguish and identify the target voice is severely affected, resulting in a significant drop in verification accuracy.
[0004] Therefore, it is urgent to develop new technologies to solve the above problems and improve the accuracy and robustness of voice object verification. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to propose a method, device, computer device and storage medium for voice object recognition, which can improve the accuracy and robustness of voice object verification.
[0006] To solve the above technical problems, the embodiments of the present application provide a method for voice object recognition, which adopts the following technical solutions:
[0007] Obtain the voice to be recognized, and use the trained voice object recognition model to extract features from the voice to be recognized to obtain a voice feature vector;
[0008] Subtract the speech feature vector from the registered feature vector in the trained speech object recognition model to obtain a target difference vector, where the registered feature vector is a feature vector generated by registering and embedding the trained speech object recognition model with the separate speech of the target speech object in advance;
[0009] Add a position encoding to the target difference vector, and input the position-encoded target difference vector into the encoder of the trained speech object recognition model for processing to obtain a target difference vector encoding;
[0010] Use a multi-layer perceptron and a preset activation function to calculate a verification score based on the target difference vector encoding, and determine the probability that the target speech object appears in the speech to be recognized according to the verification score.
[0011] To solve the above technical problems, an embodiment of the present application also provides a speech object recognition device, which adopts the following technical solutions:
[0012] An acquisition module, configured to acquire the speech to be recognized, and use the trained speech object recognition model to extract features from the speech to be recognized to obtain a speech feature vector;
[0013] A subtraction module, configured to subtract the speech feature vector from the registered feature vector in the trained speech object recognition model to obtain a target difference vector, where the registered feature vector is a feature vector generated by registering and embedding the trained speech object recognition model with the separate speech of the target speech object in advance;
[0014] An encoding module, configured to add a position encoding to the target difference vector, and input the position-encoded target difference vector into the encoder of the trained speech object recognition model for processing to obtain a target difference vector encoding;
[0015] A calculation module, configured to use a multi-layer perceptron and a preset activation function to calculate a verification score based on the target difference vector encoding, and determine the probability that the target speech object appears in the speech to be recognized according to the verification score.
[0016] To solve the above technical problems, an embodiment of the present application also provides a computer device, which adopts the following technical solutions: The computer device includes a memory and a processor, and computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, the steps of the speech object recognition method described in any one of the above are implemented.
[0017] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solutions: Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the voice object recognition method described in any one of the above are implemented.
[0018] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:
[0019] The solution of the embodiment of the present application extracts features from the voice to be recognized by using the trained voice object recognition model, can accurately capture the key information in the voice signal, and convert it into a voice feature vector with high discrimination; by subtracting the extracted voice feature vector from the registered feature vector in the voice object recognition model, a target difference vector is obtained, which can highlight the difference between the voice to be recognized and the registered voice, and helps the subsequent recognition of the target voice object; by adding position encoding to the target difference vector, the timing information in the voice signal can be introduced, which is convenient for the subsequent use of the encoder to process continuous voice signals and capture the relationship between voice features. By inputting the position-encoded target difference vector into the encoder for processing to further capture the timing features and internal structure of the voice signal, a target difference vector encoding is obtained, which can generate a feature representation with stronger expression ability and robustness; by using a multi-layer perceptron and a preset activation function, a verification score is calculated according to the target difference vector encoding, and the probability that the target voice object appears in the voice to be recognized is determined according to the verification score. The multi-layer perceptron can be used to learn complex non-linear relationships and map the target difference vector encoding to the verification score, so as to effectively judge whether the target voice object appears in the voice to be recognized. Based on the solution of the present application, the possibility of the target voice object appearing in the test voice can be more accurately quantified, thus significantly improving the robustness and accuracy of voice object verification. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;
[0022] Figure 2 is a flowchart of an embodiment of the voice object recognition method of the present application;
[0023] Figure 3It is a schematic structural diagram of an embodiment of the voice object recognition device of the present application;
[0024] Figure 4 It is a schematic structural diagram of an embodiment of the computer device of the present application. Detailed implementation manners
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0026] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase may not necessarily refer to the same embodiment when it appears in various places in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0027] To enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0028] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0029] Users can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.
[0030] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a desktop computer, and the like.
[0031] The server 103 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal device 101.
[0032] It should be noted that the voice object recognition method provided by the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the voice object recognition device is generally set in the server / terminal device.
[0033] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in
[0034] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 2 Continuing to refer to
[0035] Figure [FIGURE NUMBER] (not shown), a flowchart of an embodiment of the question-and-answer method according to the present application is shown. The voice object recognition method includes the following steps:
[0036] In this embodiment, the electronic device (such as Figure 1 the server / terminal device shown in [[FIGURE REFERENCE]]) on which the voice object recognition method runs can obtain a target question containing context information through a wired connection or a wireless connection. It should be noted that the above wireless connection methods can include, but are not limited to, 2G / 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.
[0037] (Note: In the above translation, [FIGURE NUMBER] and [FIGURE REFERENCE] need to be replaced with the actual figure number and figure reference in the original text.)Specifically, the speech to be recognized is obtained from a recording device or a pre-collected audio file library. The speech to be recognized refers to the audio data that needs to be processed for speech object recognition. In this embodiment, a trained speech object recognition model is obtained. The speech object recognition model is a deep learning model capable of extracting feature vectors from speech signals and recognizing speech objects based on the feature vectors. For the obtained speech to be recognized, the speech to be recognized is input into the trained speech object recognition model, and the trained speech object recognition model is used to extract features from the speech to be recognized, obtaining a speech feature vector. Among them, the speech feature vector is a numerical vector used to represent the key information contained in the speech to be recognized, and may include, but is not limited to, the time domain, frequency domain or other features of the speech signal.
[0038] In some alternative implementation manners, before the above-mentioned obtaining of the speech to be recognized, the following steps may further be included:
[0039] Obtain pre-collected training samples, where the training samples include a number of association pairs of registered speech and test speech, and the registered speech is composed of the individual speech of the training speech object;
[0040] Specifically, obtain pre-collected training samples. The training samples include a number of association pairs of registered speech and test speech. Among them, the registered speech is audio data composed of the individual speech of the training speech object, and the individual speech refers to the speech that only contains the speech signal of the training speech object; the test speech refers to the audio data used for training speech object recognition processing.
[0041] Input the training samples into a pre-trained speech object recognition model. The pre-trained speech object recognition model includes a registration embedding network and a feature extraction network;
[0042] Specifically, input the training samples (including a number of association pairs of registered speech and test speech) into the pre-trained speech object recognition model in sequence. Among them, the pre-trained speech object recognition model includes a registration embedding network and a feature extraction network. The registration embedding network is used to convert the input registered speech into a feature vector of a fixed length (i.e., the registration feature vector); the feature extraction network is used to convert the input test speech into a feature vector (i.e., the test feature vector).
[0043] Perform registration embedding processing on the registered speech by using the registration embedding network to obtain the registration feature vector of the training speech object;
[0044] Specifically, the registration embedding network in the voice object recognition model is used to extract features from the registered voice, and the extracted voice feature vector is mapped into the registration space to achieve the registration embedding of the voice feature vector of the training voice object. Among them, the voice feature vector mapped into the registration space is the registration feature vector of the training voice object. This registration feature vector can be used to uniquely identify the training voice object. Among them, the registration embedding network can be constructed using the architecture of the time delay neural network TDNN.
[0045] In some optional implementation manners, the above-mentioned use of the registration embedding network to perform registration embedding processing according to the registered voice to obtain the registration feature vector of the training voice object may include the following steps:
[0046] Format the registered voice according to the input format of the registration embedding network to obtain formatted voice data;
[0047] Input the formatted voice data into the registration embedding network, and use the convolutional layer of the registration embedding network to extract features from the formatted voice data to obtain local registration features;
[0048] Use the fully connected layer of the registration embedding network to map the local registration features into a preset registration feature space to obtain global registration features, which are used as the registration feature vectors of the training voice objects.
[0049] Specifically, obtain the input format of the registration embedding network, format the registered voice according to the input format of the registration embedding network, so as to convert the registered voice into a format suitable for input into the registration embedding network to obtain formatted voice data. Input the formatted voice data into the registration embedding network, and use multiple convolutional layers in the registration embedding network to perform feature extraction on the formatted voice data using dot product operations to obtain local registration features. Among them, different convolutional layers focus on extracting different types of local registration features. Transmit the extracted local registration features to the fully connected layer of the registration embedding network, and use the fully connected layer to map the local registration features into a preset registration feature space to combine and obtain global registration features, and use this global registration feature as the registration feature vector of the training voice object to uniquely identify the training voice object.
[0050] In the embodiments of the present application, the registered speech is formatted according to the input format of the registration embedding network to ensure that the input data meets the processing requirements of the network. The convolutional layer of the registration embedding network uses mechanisms such as dot product operations to extract features from the formatted speech data, obtaining local registration features, which can capture local details and patterns in the speech signal and help the model more accurately distinguish different speech objects. The fully connected layer maps the local registration features into a preset registration feature space to form a global feature vector that can represent the entire training speech object. Based on the embodiments of the present application, by extracting local features through the convolutional layer and performing global mapping through the fully connected layer, the feature information of the speech object can be captured and expressed more comprehensively.
[0051] Use the feature extraction network to extract features from the test speech to obtain the test feature vector of the test speech;
[0052] Specifically, use the feature extraction network in the speech object recognition model to extract frame-by-frame features from the test speech to obtain the test feature vector of the test speech, where the feature extraction network can be constructed using the architecture of the time-delay neural network TDNN.
[0053] Subtract the test feature vector from the registration feature vector of the training speech object to obtain a training difference vector;
[0054] Specifically, align the test feature vector extracted based on the test speech with the registration feature vector of the training speech object, and subtract the aligned test feature vector and the registration feature vector of the training speech object to obtain a training difference vector.
[0055] Optionally, when the feature extraction network used to extract the test feature vector is different from the registration embedding network used to extract the registration feature vector, that is, the test feature vector extracted by the model and the registration feature vector do not have the same feature dimension, then first perform dimensionality reduction or dimensionality increase processing on the test feature vector according to the feature dimension of the registration feature vector, so that the test feature vector and the registration feature vector have the same feature dimension, and then align the test feature vector and the registration feature vector, and subtract based on the aligned test feature vector and registration feature vector to obtain a training difference vector.
[0056] In some alternative implementation manners, to ensure the compatibility between the registration feature vector and the test feature vector, the registration embedding network and the feature extraction network can adopt the same network dimension structure, so that the test feature vector output by the feature extraction network is consistent with the registration feature vector in dimension. The above-mentioned subtracting the test feature vector from the registration feature vector of the training speech object to obtain a training difference vector may include the following steps:
[0057] Align each first element in the test feature vector with each second element in the registered feature vector of the training voice object to obtain an association pair of the first element and the second element;
[0058] Subtract the second element from the first element in the association pair to obtain a training difference vector.
[0059] Specifically, when the registration embedding network and the feature extraction network adopt the same network dimension structure, align and associate each first element in the test feature vector with each second element in the registered feature vector of the training voice object one by one to generate several association pairs of the first element and the second element. Subtract the second element from the first element in each association pair in turn to obtain a training difference vector.
[0060] Add a position encoding to the training difference vector, and input the position-encoded training difference vector into the encoder of the pre-trained voice object recognition model for processing to obtain a training difference vector encoding;
[0061] Specifically, in order to introduce the timing information of the voice signal, add a position encoding to the training difference vector to obtain a position-encoded training difference vector. Among them, the way to add a position encoding to the training difference vector can be to generate a corresponding position encoding according to the timing position information of each element in the training difference vector, or to use a deep learning model to dynamically generate a better position representation (i.e., learnable position encoding) according to the timing position information of the training difference vector. Input the position-encoded training difference vector into the encoder of the pre-trained voice object recognition model. Through the encoder using the self-attention mechanism, further process the position-encoded training difference vector to extract a more expressive feature representation, that is, the training difference vector encoding.
[0062] In some optional implementation manners, adding a position encoding to the training difference vector and inputting the position-encoded training difference vector into the encoder of the pre-trained voice object recognition model for processing to obtain a training difference vector encoding may include the following steps:
[0063] Generate a corresponding position encoding according to each third element in the training difference vector, and splice the position encoding correspondingly to each third element to obtain a position-encoded training difference vector;
[0064] Input the position-encoded training difference vector into the encoder of the pre-trained voice object recognition model, and use the self-attention mechanism component of the encoder to capture the correlation between each third element in the position-encoded training difference vector to obtain the correlation value between the third elements;
[0065] Using the feedforward neural network of the encoder, encode and transform the correlation values between the third elements to obtain the training difference vector encoding.
[0066] Specifically, a combination of sine and cosine functions is used to generate corresponding position encodings for each third element in the training difference vector, that is, sine-cosine position encodings. The position encodings are correspondingly spliced onto each third element of the training difference vector to obtain the training difference vector after position encoding. This training difference vector not only contains the training difference relationship between speech features but also contains the position relationship between each element in the difference vector. The training difference vector after position encoding is input into the encoder of the pre-trained speech object recognition model. Using the self-attention mechanism component in the encoder, calculate the correlation values between each third element in the training difference vector after position encoding to capture the correlation between each third element, and obtain the correlation values between the third elements. These correlation values are represented in the form of a matrix or a vector. The correlation values between the third elements are fed into the feedforward neural network (Feedforward Neural Network, FFN) of the encoder. Using the fully connected layer and non-linear activation function in the feedforward neural network, encode and transform the correlation values between the third elements to obtain the training difference vector encoding.
[0067] In the embodiment of the present application, by generating corresponding position encodings for each third element in the training difference vector and correspondingly splicing the position encodings onto each third element, the training difference vector after position encoding is obtained. The training difference vector after position encoding not only reflects the training difference relationship between speech features but also reflects the position relationship between each element in the difference vector, enabling the model to better capture the temporal features and internal structure of the speech signal, thereby improving the accuracy of speech object recognition. Through the self-attention mechanism component of the encoder, the correlation values between each third element in the training difference vector after position encoding can be calculated to capture the correlation between elements, enabling the model to more flexibly process the dependency relationships in the input data, especially those spanning long distances; through the feedforward neural network (FFN) in the encoder, the correlation values between the third elements can be encoded and transformed to obtain the training difference vector encoding, enabling the model to more efficiently process the input data and generate useful representations. By adding position encodings to the training difference vector and using the encoder of the pre-trained speech object recognition model for processing, more accurate understanding and representation of the speech signal are achieved, and the performance of the model is improved.
[0068] Calculate the loss function value by using a preset loss function and the training difference vector encoding, calculate the gradient according to the loss function value, and update the parameters of the pre-trained speech object recognition model according to the gradient;
[0069] Iteratively train the pre-trained speech object recognition model until the convergence condition is met, stop the model training, and obtain the trained speech object recognition model.
[0070] Specifically, obtain a preset loss function, such as the binary cross-entropy (BCE) loss function; use the loss function to calculate the loss function value according to the training difference vector encoding and the actual difference vector encoding. Calculate the gradient according to the calculated loss function value. The gradient is a vector that points in the direction where the loss function value decreases fastest. According to the calculated gradient, use the gradient descent method to update the parameters of the pre-trained speech object recognition model to make the prediction result of the model closer to the actual result. Repeat the above process to iteratively train the pre-trained speech object recognition model. When the model meets the convergence condition (that is, the loss function value drops to a certain level or no longer decreases significantly, or the number of iterations reaches the preset upper limit), stop the model training and obtain the trained speech object recognition model.
[0071] Step S202: Subtract the speech feature vector from the registered feature vector in the trained speech object recognition model to obtain a target difference vector, where the registered feature vector is a feature vector generated by registering and embedding the trained speech object recognition model with the individual speech of the target speech object in advance;
[0072] In this embodiment, the trained speech object recognition model is registered and embedded with the individual speech of the target speech object in advance, that is, the trained speech object recognition model is used to extract the speech feature vector of the individual speech and map the extracted speech feature vector into the registration space to realize the registration and embedding of the speech feature vector of the target speech object. The speech feature vector mapped into the registration space is the registered feature vector. This registered feature vector can be used to uniquely identify the target speech object.
[0073] Specifically, align the speech feature vector extracted from the speech to be recognized with the registered feature vector in the trained speech object recognition model, and subtract based on the aligned speech feature vector and the registered feature vector to obtain a target difference vector.
[0074] Optionally, when the network dimension structure for extracting the speech feature vector in the speech object recognition model is the same as the network dimension structure for extracting the registration feature vector, that is, when the speech feature vector extracted by the speech object recognition model and the registration feature vector have the same feature dimension, the speech feature vector and the registration feature vector are aligned, and based on the aligned speech feature vector and registration feature vector, subtraction is performed to obtain the target difference vector.
[0075] Optionally, when the network dimension structure for extracting the speech feature vector in the speech object recognition model is different from the network dimension structure for extracting the registration feature vector, that is, when the speech feature vector extracted by the speech object recognition model and the registration feature vector do not have the same feature dimension, the speech feature vector is first subjected to dimensionality reduction or dimensionality increase processing according to the feature dimension of the registration feature vector, so that the speech feature vector and the registration feature vector have the same feature dimension, and then the speech feature vector and the registration feature vector are aligned, and based on the aligned speech feature vector and registration feature vector, subtraction is performed to obtain the target difference vector.
[0076] Step S203: Add position encoding to the target difference vector, and input the target difference vector after position encoding into the encoder of the trained speech object recognition model for processing to obtain the target difference vector encoding;
[0077] In this embodiment, in order to introduce the timing information of the speech signal, position encoding is added to the target difference vector to obtain the target difference vector after position encoding. Among them, the position encoding is used to add position information to the input sequence (that is, the target difference vector) to make up for the lack of sense of order in the self-attention mechanism in the encoder. The method of adding position encoding to the target difference vector can be to generate corresponding position encoding according to the timing position information of each element in the target difference vector, or to use a deep learning model to dynamically generate a better position representation (that is, learnable position encoding) according to the timing position information of the target difference vector.
[0078] It should be noted that since the speech signal is a kind of timing information and there is a time sequence order between the elements it contains (such as speech frames), it is necessary to introduce timing position information when processing the speech signal. By adding position encoding to the target difference vector, the model can better understand the timing relationship in the speech signal, so as to capture the timing features and internal structure of the speech signal, thereby improving the accuracy of speech object recognition.
[0079] Specifically, the position-encoded target difference vector is input into the encoder of the trained speech object recognition model, and the encoder uses a self-attention mechanism to further process the position-encoded target difference vector to extract a more expressive feature representation, namely, the target difference vector encoding.
[0080] Step S204, using a multilayer perceptron and a preset activation function, calculating a verification score according to the target difference vector encoding, and determining the probability that the target speech object appears in the speech to be recognized according to the verification score.
[0081] In this embodiment, the target difference vector code is input into a pre-trained multi-layer perceptron (MLP), and the target difference vector code is nonlinearly transformed through each perceptual layer and a preset activation function in the multi-layer perceptron to obtain a number of verification scores, which reflect the possibility of the target voice object appearing in the speech to be recognized, and the value range is between 0 and 1. The various verification scores are summarized, and the probability of the target voice object appearing in the speech to be recognized is determined based on the summary result, for example, the summary result is compared with a preset threshold to determine whether the target voice object appears in the speech to be recognized.
[0082] In some optional implementations, the above-mentioned use of a multilayer perceptron and a preset activation function to calculate a verification score according to the target difference vector encoding may include the following steps:
[0083] Inputting the target difference vector code into a multi-layer perceptron, and processing the target difference vector code using a weight matrix set in each layer of the multi-layer perceptron to obtain an output value;
[0084] The output value is calculated using a preset activation function to obtain a verification score.
[0085] Specifically, the target difference vector code is input into a pre-trained multi-layer perceptron (MLP), and the target difference vector code is weighted by the weight matrix set in each perceptual layer of the multi-layer perceptron to obtain an output value. A preset activation function, such as a Sigmoid function, is obtained, and the output value is nonlinearly transformed using the activation function to obtain a number of verification scores, which reflect the possibility of the target speech object appearing in the speech to be recognized, and its value range is between 0-1.
[0086] In some optional implementations, the determining, according to the verification score, the probability that the target voice object appears in the voice to be recognized may include the following steps:
[0087] Obtaining a preset score threshold, and comparing the verification score with the score threshold;
[0088] If the verification score is greater than or equal to the score threshold, it is determined that the probability of the target voice object appearing in the voice to be recognized is relatively high;
[0089] If the verification score is less than the score threshold, it is determined that the probability of the target voice object appearing in the voice to be recognized is relatively low.
[0090] Specifically, a preset score threshold is obtained, which is used to measure whether the target voice object appears in the voice to be recognized. The verification score is compared with the score threshold. When the verification score is greater than or equal to the score threshold, it is determined that the probability of the target voice object appearing in the voice to be recognized is relatively high; when the verification score is less than the score threshold, it is determined that the probability of the target voice object appearing in the voice to be recognized is relatively low.
[0091] Through the above solution in the embodiments of the present application, specifically, by using the trained voice object recognition model to extract features from the voice to be recognized, the key information in the voice signal can be accurately captured and converted into a voice feature vector with high discrimination; by subtracting the extracted voice feature vector from the registered feature vector in the voice object recognition model, a target difference vector is obtained, which can highlight the difference between the voice to be recognized and the registered voice, and is helpful for the subsequent recognition of the target voice object; by adding a position encoding to the target difference vector, the timing information in the voice signal can be introduced, which is convenient for the subsequent use of the encoder to process continuous voice signals and capture the relationship between voice features. By inputting the position-encoded target difference vector into the encoder for processing to further capture the timing features and internal structure of the voice signal, a target difference vector encoding is obtained, which can generate a feature representation with stronger expression ability and robustness; by using a multi-layer perceptron and a preset activation function, the verification score is calculated according to the target difference vector encoding, and the probability of the target voice object appearing in the voice to be recognized is determined according to the verification score. The multi-layer perceptron can be used to learn complex non-linear relationships and map the target difference vector encoding to the verification score, so as to effectively judge whether the target voice object appears in the voice to be recognized. Based on the solution of the present application, the possibility of the target voice object appearing in the test voice can be more accurately quantified, thus significantly improving the robustness and accuracy of voice object verification.
[0092] Exemplarily, in a telemedicine consultation and intelligent medical record system in the medical field, accurate and robust voice object recognition technology is crucial for protecting patient privacy and ensuring the correct execution of medical instructions. First, collect a large number of clear voice samples of target medical personnel (such as doctors, nurses), as well as various noise backgrounds that may appear in a simulated medical environment (such as equipment operation sounds, conversations of other people, etc.). Use this data to train a voice object recognition model so that it can extract discriminative voice feature vectors. For each medical personnel, use the individual voice sample of the medical personnel to register the trained model to generate corresponding registered feature vectors. These vectors will be used as benchmarks in the subsequent verification process. During a telemedicine consultation, when a voice request from a patient is received, use the trained model to extract the feature vector of this voice. Subtract the feature vector of the voice to be recognized from the registered feature vector of the corresponding medical personnel to obtain a target difference vector. Subsequently, add a position encoding to the target difference vector, and input the position-encoded target difference vector into the encoder of the model for processing to obtain a target difference vector encoding. Use a multi-layer perceptron and a preset activation function (such as ReLU or Sigmoid) to calculate a verification score based on the target difference vector encoding. The higher the verification score, the greater the probability that the target medical personnel appears in the voice to be recognized. Finally, set a threshold based on the verification score. When the verification score exceeds this threshold, it is considered that the verification passes and the medical consultation is allowed to continue; otherwise, it is prompted that the verification fails and additional identity verification measures may need to be taken.
[0093] Through the above method, even in a noisy medical environment, the voice of the target medical personnel can be effectively recognized, improving the accuracy and robustness of voice object verification, thereby ensuring the security and efficiency of telemedicine consultation.
[0094] Exemplarily, in the financial field, especially in telephone banking and intelligent customer service systems, accurately verifying the customer's identity is the key to ensuring transaction security. First, collect clear voice samples of a large number of customers, as well as various noise backgrounds that may appear in a simulated financial transaction environment (such as telephone line noise, background music, etc.). Use this data to train a voice object recognition model so that it can extract discriminative customer voice feature vectors. For each customer, use the customer's individual voice sample to register the trained model to generate corresponding registered feature vectors. These vectors will serve as a benchmark in the subsequent verification process. When a customer conducts a transaction through telephone banking or an intelligent customer service system, use the trained model to extract the feature vectors of the voice. Subtract the feature vectors of the voice to be verified from the customer's registered feature vectors to obtain a target difference vector. Subsequently, add position encoding to the target difference vector, and input the position-encoded target difference vector into the encoder of the model for processing to obtain a target difference vector encoding. Use a multi-layer perceptron and a preset activation function (such as ReLU or Sigmoid) to calculate a verification score based on the target difference vector encoding. The higher the verification score, the greater the probability that the customer's identity appears in the voice to be verified. Finally, set a threshold based on the verification score. When the verification score exceeds this threshold, it is considered that the customer identity verification is passed and the transaction is allowed to continue; otherwise, the transaction is rejected and the customer is prompted to perform additional identity verification. At the same time, record all verification attempts and results for subsequent auditing and security analysis.
[0095] Through the above method, even in a noisy financial transaction environment, the customer's voice can be effectively identified, improving the accuracy and robustness of voice object verification, thereby ensuring the security and reliability of financial transactions.
[0096] In addition, in the field of property insurance, by introducing the Neural Scoring method based on the TDNN and Transformer architectures, the accuracy of the home security monitoring system can be improved. By accurately identifying the voices of family members, false alarms can be reduced and the response ability of the home anti-theft system can be enhanced. In the field of financial services, especially in mobile payment and e-banking services, this technology, as an additional security verification means, can accurately identify the customer's identity during the transaction process, reduce the fraud risk, and ensure transaction security. In addition, the advantage of this technology in dealing with mixed voice and noise interference enables it to maintain high accuracy in complex environments, which plays an important role for enterprises to improve customer service quality, ensure transaction security, and enhance user experience. Through this innovative speaker verification method, enterprises can provide more secure and reliable financial services for customers, and at the same time open up a new path for the technological innovation and business development of enterprises.
[0097] It should be emphasized that, to further ensure the privacy and security of the above-mentioned voice information to be recognized and other relevant data, the above-mentioned voice information to be recognized and other relevant data can also be stored in a node of a blockchain.
[0098] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0099] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned method embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a Read-Only Memory (ROM), etc., or a Random Access Memory (RAM), etc.
[0100] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. Their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0101] Further referring to Figure 3 As an implementation of the above-mentioned Figure 2 shown method, an embodiment of a voice object recognition device is provided in this application. This device embodiment corresponds to the Figure 2 shown method embodiment, and this device can be specifically applied to various electronic devices.
[0102] As Figure 3 shown, the voice object recognition device 400 described in this embodiment includes: an acquisition module 401, a subtraction module 402, an encoding module 403, and a calculation module 404. Among them:
[0103] An acquisition module 401, configured to acquire a speech to be recognized, extract features from the speech to be recognized by using a trained speech object recognition model, and obtain a speech feature vector;
[0104] A subtraction module 402, configured to subtract the speech feature vector from a registered feature vector in the trained speech object recognition model to obtain a target difference vector, where the registered feature vector is a feature vector generated by registering and embedding a separate speech of a target speech object into the trained speech object recognition model in advance;
[0105] An encoding module 403, configured to add position encoding to the target difference vector, input the target difference vector after position encoding into an encoder of the trained speech object recognition model for processing, and obtain a target difference vector encoding;
[0106] A calculation module 404, configured to calculate a verification score according to the target difference vector encoding by using a multi-layer perceptron and a preset activation function, and determine a probability that the target speech object appears in the speech to be recognized according to the verification score.
[0107] In this embodiment, the acquisition module 401 is configured to acquire a speech to be recognized, where the speech to be recognized refers to audio data that needs to be processed for speech object recognition. In this embodiment, a trained speech object recognition model is acquired, and the speech object recognition model is a deep learning model capable of extracting feature vectors from speech signals and recognizing speech objects according to the feature vectors. For the acquired speech to be recognized, the speech to be recognized is input into the trained speech object recognition model, and the trained speech object recognition model is used to extract features from the speech to be recognized to obtain a speech feature vector. Among them, the speech feature vector is a numerical vector used to represent key information included in the speech to be recognized, and may include, but is not limited to, time domain, frequency domain or other features of the speech signal.
[0108] Specifically, the subtraction module 402 is configured to register and embed a separate speech of a target speech object into the trained speech object recognition model in advance, that is, use the trained speech object recognition model to extract features from the separate speech, and map the extracted speech feature vector into a registration space to implement registration and embedding of the speech feature vector of the target speech object, where the speech feature vector mapped into the registration space is the registered feature vector. The registered feature vector can be used to uniquely identify the target speech object.
[0109] Specifically, the speech feature vector extracted from the speech to be recognized is aligned with the registered feature vector in the trained speech object recognition model, and the aligned speech feature vector and the registered feature vector are subtracted to obtain a target difference vector.
[0110] Optionally, when the network dimension structure for extracting the speech feature vector in the speech object recognition model is the same as the network dimension structure for extracting the registered feature vector, that is, when the speech feature vector extracted by the speech object recognition model and the registered feature vector have the same feature dimension, the speech feature vector and the registered feature vector are aligned, and the aligned speech feature vector and the registered feature vector are subtracted to obtain a target difference vector.
[0111] Optionally, when the network dimension structure for extracting the speech feature vector in the speech object recognition model is different from the network dimension structure for extracting the registered feature vector, that is, when the speech feature vector extracted by the speech object recognition model and the registered feature vector do not have the same feature dimension, the speech feature vector is first dimension-reduced or dimension-increased according to the feature dimension of the registered feature vector so that the speech feature vector and the registered feature vector have the same feature dimension, and then the speech feature vector and the registered feature vector are aligned, and the aligned speech feature vector and the registered feature vector are subtracted to obtain a target difference vector.
[0112] Specifically, in order to introduce the timing information of the speech signal, the encoding module 403 is used to add position encoding to the target difference vector to obtain the target difference vector after position encoding, where the position encoding is used to add position information to the input sequence (i.e., the target difference vector) to make up for the lack of sense of order in the self-attention mechanism in the encoder. The way to add position encoding to the target difference vector can be to generate corresponding position encoding according to the timing position information of each element in the target difference vector, or to use a deep learning model to dynamically generate a better position representation (i.e., learnable position encoding) according to the timing position information of the target difference vector.
[0113] It should be noted that since the speech signal is a kind of timing information and there is a temporal order between the elements (such as speech frames) it contains, it is necessary to introduce timing position information when processing the speech signal. By adding position encoding to the target difference vector, the model can better understand the timing relationship in the speech signal, so as to capture the timing features and internal structure of the speech signal, thereby improving the accuracy of speech object recognition.
[0114] Specifically, the target difference vector after position encoding is input into the encoder of the trained voice object recognition model. Through the encoder using the self-attention mechanism, the target difference vector after position encoding is further processed to extract a more expressive feature representation, that is, the target difference vector encoding.
[0115] Specifically, the calculation module 404 is configured to input the target difference vector encoding into a pre-trained multi-layer perceptron (MLP). Through each perception layer in the multi-layer perceptron and a preset activation function, a non-linear transformation is performed on the target difference vector encoding to obtain a number of verification scores. The verification score reflects the possibility of the target voice object appearing in the speech to be recognized, and its value range is between 0 and 1. Summarize each verification score, and determine the probability of the target voice object appearing in the speech to be recognized according to the summary result. For example, compare the summary result with a preset threshold to determine whether the target voice object appears in the speech to be recognized.
[0116] The voice object recognition device 400 according to the embodiment of the present application can accurately capture key information in the speech signal and convert it into a highly discriminative speech feature vector by using the trained voice object recognition model to extract features from the speech to be recognized; by subtracting the extracted speech feature vector from the registered feature vector in the voice object recognition model, a target difference vector is obtained, which can highlight the difference between the speech to be recognized and the registered speech, and is helpful for subsequent recognition of the target voice object; by adding position encoding to the target difference vector, the timing information in the speech signal can be introduced, which is convenient for subsequent processing of continuous speech signals by the encoder and capturing the relationship between speech features. By inputting the target difference vector after position encoding into the encoder for processing to further capture the timing features and internal structure of the speech signal, and obtaining the target difference vector encoding, a feature representation with stronger expressiveness and robustness can be generated; by using a multi-layer perceptron and a preset activation function, verification scores are calculated according to the target difference vector encoding, and the probability of the target voice object appearing in the speech to be recognized is determined according to the verification scores. The complex non-linear relationship can be learned by using the multi-layer perceptron to map the target difference vector encoding to the verification scores, so as to effectively judge whether the target voice object appears in the speech to be recognized. Based on the solution of the present application, the possibility of the target voice object appearing in the test speech can be more accurately quantified, thereby significantly improving the robustness and accuracy of voice object verification.
[0117] In some optional implementation manners of this embodiment, the voice object recognition device 400 may further include a training module. Among them:
[0118] A training module for obtaining pre-collected training samples, where the training samples include a number of associated pairs of registered voices and test voices, and the registered voices are composed of the individual voices of the training voice objects;
[0119] Input the training samples into a pre-trained voice object recognition model, where the pre-trained voice object recognition model includes a registration embedding network and a feature extraction network;
[0120] Use the registration embedding network to perform registration embedding processing based on the registered voices to obtain the registration feature vectors of the training voice objects;
[0121] Use the feature extraction network to perform feature extraction based on the test voices to obtain the test feature vectors of the test voices;
[0122] Subtract the test feature vectors from the registration feature vectors of the training voice objects to obtain training difference vectors;
[0123] Add positional encoding to the training difference vectors, and input the training difference vectors after positional encoding into the encoder of the pre-trained voice object recognition model for processing to obtain training difference vector encodings;
[0124] Use a preset loss function to calculate a loss function value with the training difference vector encodings, calculate gradients based on the loss function value, and update the parameters of the pre-trained voice object recognition model according to the gradients;
[0125] Iteratively train the pre-trained voice object recognition model until a convergence condition is met, stop the model training, and obtain a trained voice object recognition model.
[0126] Specifically, the training module is used to obtain pre-collected training samples, and the training samples contain a number of associated pairs of registered voices and test voices. Among them, the registered voices are audio data composed of the individual voices of the training voice objects, and the individual voices refer to voices that only contain the voice signals of the training voice objects; the test voices refer to audio data used for training voice object recognition processing.
[0127] Specifically, the training module is used to sequentially input the training samples (including a number of associated pairs of registered voices and test voices) into a pre-trained voice object recognition model. Among them, the pre-trained voice object recognition model includes a registration embedding network and a feature extraction network. The registration embedding network is used to convert the input registered voices into fixed-length feature vectors (i.e., registration feature vectors); the feature extraction network is used to convert the input test voices into feature vectors (i.e., test feature vectors).
[0128] Specifically, the training module is used to extract features from the registered speech using the registration embedding network in the speech object recognition model, and map the extracted speech feature vectors into the registration space to achieve the registration embedding of the speech feature vectors of the training speech objects. Among them, the speech feature vectors mapped into the registration space are the registration feature vectors of the training speech objects. The registration feature vectors can be used to uniquely identify the training speech objects. Among them, the registration embedding network can be constructed using the architecture of the time delay neural network TDNN.
[0129] Specifically, the training module is used to extract frame-by-frame features from the test speech using the feature extraction network in the speech object recognition model, and obtain the test feature vectors of the test speech. Among them, the feature extraction network can be constructed using the architecture of the time delay neural network TDNN.
[0130] Specifically, the training module is used to align the test feature vectors extracted based on the test speech with the registration feature vectors of the training speech objects, and subtract the aligned test feature vectors from the registration feature vectors of the training speech objects to obtain the training difference vectors.
[0131] Optionally, when the feature extraction network used to extract the test feature vectors is different from the registration embedding network used to extract the registration feature vectors, that is, when the test feature vectors extracted by the model and the registration feature vectors do not have the same feature dimensions, the test feature vectors are first processed by dimensionality reduction or dimensionality increase according to the feature dimensions of the registration feature vectors, so that the test feature vectors and the registration feature vectors have the same feature dimensions, and then the test feature vectors are aligned with the registration feature vectors, and the aligned test feature vectors and registration feature vectors are subtracted to obtain the training difference vectors.
[0132] Specifically, in order to introduce the timing information of the speech signal, the training module is used to add position encoding to the training difference vectors to obtain the training difference vectors after position encoding. Among them, the method of adding position encoding to the training difference vectors can be to generate corresponding position encoding according to the timing position information of each element in the training difference vectors, or to use a deep learning model to dynamically generate better position representations (i.e., learnable position encoding) according to the timing position information of the training difference vectors. The training difference vectors after position encoding are input into the encoder of the pre-trained speech object recognition model, and through the encoder using the self-attention mechanism, the training difference vectors after position encoding are further processed to extract more expressive feature representations, that is, the training difference vector encodings.
[0133] Specifically, the training module is used to obtain a preset loss function, such as the binary cross-entropy (BCE) loss function; use the loss function to calculate the loss function value according to the training difference vector encoding and the actual difference vector encoding. Calculate the gradient based on the calculated loss function value. The gradient is a vector that points in the direction where the loss function value decreases fastest. According to the calculated gradient, use the gradient descent method to update the parameters of the pre-trained speech object recognition model so that the prediction result of the model is closer to the actual result. Repeat the above process to iteratively train the pre-trained speech object recognition model. When the model meets the convergence condition (that is, the loss function value drops to a certain extent or no longer decreases significantly, or the number of iterations reaches the preset upper limit), stop the model training to obtain the trained speech object recognition model.
[0134] In some optional implementation manners of this embodiment, the above training module may include a formatting sub-module, a local extraction sub-module, and a global mapping sub-module. Among them:
[0135] The formatting sub-module is used to format the registered speech according to the input format of the registration embedding network to obtain formatted speech data;
[0136] The local extraction sub-module is used to input the formatted speech data into the registration embedding network, and use the convolutional layer of the registration embedding network to extract features from the formatted speech data to obtain local registration features;
[0137] The global mapping sub-module is used to use the fully connected layer of the registration embedding network to map the local registration features into a preset registration feature space to obtain global registration features, which are used as the registration feature vectors of the training speech object.
[0138] Specifically, the formatting sub-module is used to obtain the input format of the registration embedding network, format the registered speech according to the input format of the registration embedding network, so as to convert the registered speech into a format suitable for input into the registration embedding network to obtain formatted speech data. The local extraction sub-module is used to input the formatted speech data into the registration embedding network, and use multiple convolutional layers in the registration embedding network to perform feature extraction on the formatted speech data through dot product operations to obtain local registration features. Among them, different convolutional layers focus on extracting different types of local registration features. The global mapping sub-module is used to send the extracted local registration features to the fully connected layer of the registration embedding network, use the fully connected layer to map the local registration features into a preset registration feature space, combine to obtain global registration features, and use the global registration features as the registration feature vectors of the training speech object to uniquely identify the training speech object.
[0139] The training module of the embodiment of the present application formats the registered speech according to the input format of the registration embedding network to ensure that the input data meets the processing requirements of the network; through mechanisms such as dot product operations in the convolutional layer of the registration embedding network, it extracts features from the formatted speech data to obtain local registration features, which can capture local details and patterns in the speech signal, helping the model to more accurately distinguish different speech objects. The local registration features are mapped to a preset registration feature space through the fully connected layer to form a global feature vector that can represent the entire training speech object. Based on the embodiment of the present application, by extracting local features through the convolutional layer and performing global mapping through the fully connected layer, the feature information of the speech object can be captured and expressed more comprehensively.
[0140] In some alternative implementation manners of this embodiment, the above training module may include an alignment processing sub-module and an element subtraction sub-module. Among them:
[0141] The alignment processing sub-module is used to align each first element in the test feature vector with each second element in the registration feature vector of the training speech object to obtain an association pair of the first element and the second element;
[0142] The element subtraction sub-module is used to subtract the second element from the first element in the association pair to obtain a training difference vector.
[0143] Specifically, to ensure the compatibility between the registration feature vector and the test feature vector, the registration embedding network and the feature extraction network may adopt the same network dimension structure, so that the test feature vector output by the feature extraction network is consistent with the registration feature vector in dimension. The alignment processing sub-module is used to, when the registration embedding network and the feature extraction network adopt the same network dimension structure, align and associate each first element in the test feature vector with each second element in the registration feature vector of the training speech object one by one to generate several association pairs of the first element and the second element. The element subtraction sub-module is used to subtract the second element from the first element in each association pair in turn to obtain a training difference vector.
[0144] In some alternative implementation manners of this embodiment, the above training module may include a position encoding sub-module, a relationship capture sub-module, and a coding conversion sub-module, where:
[0145] The position encoding sub-module is used to generate a corresponding position encoding according to each third element in the training difference vector, and splice the position encoding to each of the third elements to obtain a training difference vector after position encoding;
[0146] A relationship capturing sub-module, configured to input the training difference vector after position encoding into the encoder of the pre-trained voice object recognition model, and adopt the self-attention mechanism component of the encoder to capture the correlation between each of the third elements in the training difference vector after position encoding, so as to obtain the correlation value between the third elements;
[0147] An encoding conversion sub-module, configured to adopt the feed-forward neural network of the encoder to perform encoding conversion on the correlation value between the third elements, so as to obtain the training difference vector encoding.
[0148] Specifically, a position encoding sub-module is configured to use a combination of sine and cosine functions to generate corresponding position encodings for each of the third elements in the training difference vector, that is, sine-cosine position encodings, and splice the position encodings correspondingly to each of the third elements of the training difference vector to obtain the training difference vector after position encoding. This training difference vector contains both the training difference relationship between voice features and the position relationship between each element in the difference vector. A relationship capturing sub-module is configured to input the training difference vector after position encoding into the encoder of the pre-trained voice object recognition model, and adopt the self-attention mechanism component in the encoder to calculate the correlation value between each of the third elements in the training difference vector after position encoding, so as to capture the correlation between each of the third elements, and obtain the correlation value between the third elements, and this correlation value is represented in the form of a matrix or a vector. An encoding conversion sub-module is configured to deliver the correlation value between the third elements to the feed-forward neural network (Feedforward Neural Network, FFN) of the encoder, and adopt the fully connected layer and non-linear activation function in the feed-forward neural network to perform encoding conversion on the correlation value between the third elements, so as to obtain the training difference vector encoding.
[0149] The training module of the embodiment of the present application generates corresponding position encodings according to each third element in the training difference vector, and splices the position encodings to each third element correspondingly to obtain the training difference vector after position encoding. The training difference vector after position encoding not only reflects the training difference relationship between speech features, but also reflects the position relationship between each element in the difference vector, enabling the model to better capture the temporal features and internal structure of the speech signal, thereby improving the accuracy of speech object recognition. Through the self-attention mechanism component of the encoder, the correlation values between each third element in the training difference vector after position encoding can be calculated to capture the correlation between elements, enabling the model to more flexibly process the dependencies in the input data, especially those dependencies spanning long distances; through the feed-forward neural network (FFN) in the encoder, the correlation values between the third elements can be encoded and transformed to obtain the training difference vector encoding, enabling the model to more efficiently process the input data and generate useful representations. By adding position encoding to the training difference vector and using the encoder of the pre-trained speech object recognition model for processing, a more accurate understanding and representation of the speech signal are achieved, and the performance of the model is improved.
[0150] In some optional implementation manners of this embodiment, the encoding module may include a weight processing sub-module and a score calculation sub-module. Among them:
[0151] The weight processing sub-module is configured to input the target difference vector encoding into a multi-layer perceptron, and process the target difference vector encoding by using the weight matrix set in each layer of the multi-layer perceptron to obtain an output value;
[0152] The score calculation sub-module is configured to calculate the output value by using a preset activation function to obtain a verification score.
[0153] Specifically, the weight processing sub-module is configured to input the target difference vector encoding into a pre-trained multi-layer perceptron (MLP), and perform weighted processing on the target difference vector encoding by using the weight matrix set in each perceptron layer of the multi-layer perceptron to obtain an output value. The score calculation sub-module is configured to obtain a preset activation function, such as the Sigmoid function, and perform a non-linear transformation on the output value by using the activation function to obtain a plurality of verification scores. The verification score reflects the possibility of the target speech object appearing in the speech to be recognized, and its value range is between 0 and 1.
[0154] In some optional implementation manners of this embodiment, the calculation module 404 may include a score comparison sub-module, a first determination sub-module, and a second determination sub-module. Among them:
[0155] A score comparison sub-module, configured to obtain a preset score threshold and compare the verification score with the score threshold;
[0156] A first determination sub-module, configured to determine that the probability of the target voice object appearing in the voice to be recognized is relatively high if the verification score is greater than or equal to the score threshold;
[0157] A second determination sub-module, configured to determine that the probability of the target voice object appearing in the voice to be recognized is relatively low if the verification score is less than the score threshold.
[0158] Specifically, the score comparison sub-module is configured to obtain a preset score threshold, which is used to measure whether the target voice object appears in the voice to be recognized. Compare the verification score with the score threshold. The first determination sub-module is configured to determine that the probability of the target voice object appearing in the voice to be recognized is relatively high when the verification score is greater than or equal to the score threshold; the second determination sub-module is configured to determine that the probability of the target voice object appearing in the voice to be recognized is relatively low when the verification score is less than the score threshold.
[0159] To solve the above technical problems, an embodiment of the present application further provides a computer device. Specifically, please refer to Figure 4 , Figure 4 which is the basic structural block diagram of the computer device in this embodiment.
[0160] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are communicatively connected to each other through a system bus. It should be noted that only the computer device 6 with the memory 61, the processor 62, and the network interface 63 is shown in the figure. However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0161] The computer device may be a desktop computer, a notebook, a palm computer, a cloud server, and other computing devices. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, a voice control device, and other means.
[0162] The memory 61 includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 6. Of course, the memory 61 may also include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is generally used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions of the voice object recognition method. In addition, the memory 61 may also be used to temporarily store various data that have been output or will be output.
[0163] In some embodiments, the processor 62 may be a central processing unit (CPU), controller, microcontroller, microprocessor or other data processing chip. The processor 62 is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to run the computer-readable instructions stored in the memory 61 or process data, such as running the computer-readable instructions of the voice object recognition method.
[0164] The network interface 63 may include a wireless network interface or a wired network interface, and this network interface 63 is generally used to establish a communication connection between the computer device 6 and other electronic devices.
[0165] This application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions, and the computer-readable instructions can be executed by at least one processor, so that the at least one processor executes the steps of the voice object recognition method as described above.
[0166] The computer device, computer-readable storage medium, and computer-readable instructions provided by the embodiments of the present application can accurately capture key information in a speech signal and convert it into a highly discriminative speech feature vector by using a trained speech object recognition model to extract features from the speech to be recognized through a processor. By subtracting the extracted speech feature vector from the registered feature vector in the speech object recognition model, a target difference vector can be obtained, which can highlight the difference between the speech to be recognized and the registered speech, facilitating subsequent recognition of the target speech object. By adding a positional encoding to the target difference vector, the temporal information in the speech signal can be introduced, facilitating subsequent processing of continuous speech signals by an encoder and capturing the relationship between speech features. By inputting the positionally encoded target difference vector into the encoder for processing to further capture the temporal features and internal structure of the speech signal, a target difference vector encoding can be obtained, which can generate a feature representation with stronger expressive power and robustness. By using a multi-layer perceptron and a preset activation function to calculate a verification score based on the target difference vector encoding and determining the probability of the target speech object appearing in the speech to be recognized according to the verification score, the multi-layer perceptron can be used to learn complex non-linear relationships and map the target difference vector encoding to the verification score, thereby achieving an effective judgment on whether the target speech object appears in the speech to be recognized. Based on the solution of the present application, the possibility of the target speech object appearing in the test speech can be more accurately quantified, thus significantly improving the robustness and accuracy of speech object verification.
[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0168] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all of them. The preferred embodiments of this application are shown in the drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments or equivalently replace some of the technical features. Any equivalent structure that makes use of the content of this application's specification and drawings, directly or indirectly applied in other related technical fields, is similarly within the scope of this application's patent protection.
[0169] In the embodiments of this application, the software tools or components that are not of our company are only introduced by way of example and do not represent actual use.
Claims
1. A method for speech object recognition, characterized in that: The steps include: Acquire a speech to be recognized, and use a trained speech object recognition model to perform feature extraction on the speech to be recognized to obtain a speech feature vector; Subtracting the speech feature vector from the registered feature vector in the trained speech object recognition model to obtain a target difference vector, wherein the registered feature vector refers to a feature vector generated by pre-registering and embedding the trained speech object recognition model with the individual speech of the target speech object; Adding a position code to the target difference vector, and inputting the position-coded target difference vector into the encoder of the trained speech object recognition model for processing to obtain a target difference vector code; A multilayer perceptron and a preset activation function are used to calculate a verification score according to the target difference vector encoding, and the probability of the target speech object appearing in the speech to be recognized is determined according to the verification score.
2. The method for speech object recognition according to claim 1, characterized in that: Before the step of obtaining the speech to be recognized, the method further includes: Acquire a pre-collected training sample, wherein the training sample includes a plurality of associated pairs of registration speech and test speech, wherein the registration speech is composed of a separate speech of a training speech object; Inputting the training sample into a pre-trained speech object recognition model, wherein the pre-trained speech object recognition model includes a registration embedding network and a feature extraction network; Using the registration embedding network to perform registration embedding processing according to the registration speech to obtain a registration feature vector of the training speech object; Using the feature extraction network to perform feature extraction based on the test speech to obtain a test feature vector of the test speech; Subtracting the test feature vector from the registered feature vector of the training speech object to obtain a training difference vector; Adding position coding to the training difference vector, inputting the position-coded training difference vector into the encoder of the pre-trained speech object recognition model for processing, and obtaining the training difference vector coding; A loss function value is calculated by using a preset loss function and the training difference vector encoding, a gradient is calculated according to the loss function value, and a parameter of the pre-trained speech object recognition model is updated according to the gradient; The pre-trained speech object recognition model is iteratively trained until a convergence condition is met, and the model training is stopped to obtain a trained speech object recognition model.
3. The method for speech object recognition according to claim 2, characterized in that: The step of using the registration embedding network to perform registration embedding processing according to the registration speech to obtain the registration feature vector of the training speech object includes: Formatting the registered voice according to the input format of the registered embedded network to obtain formatted voice data; Inputting the formatted voice data into the registration embedding network, and extracting features from the formatted voice data using a convolutional layer of the registration embedding network to obtain local registration features; The fully connected layer of the registration embedding network is used to map the local registration features into a preset registration feature space to obtain global registration features as the registration feature vector of the training speech object.
4. The method for speech object recognition according to claim 2, characterized in that: The registered embedding network and the feature extraction network have the same network dimension structure, and the step of subtracting the test feature vector from the registered feature vector of the training speech object to obtain a training difference vector comprises: Aligning each first element in the test feature vector with each second element in the registration feature vector of the training speech object to obtain an associated pair of the first element and the second element; The first element and the second element in the associated pair are subtracted to obtain a training difference vector.
5. The method for speech object recognition according to claim 2, characterized in that: The step of adding position coding to the training difference vector and inputting the position-coded training difference vector into the encoder of the pre-trained speech object recognition model for processing to obtain the training difference vector coding comprises: Generate a corresponding position code according to each third element in the training difference vector, and splice the position code to each third element to obtain a training difference vector after the position code; Inputting the position-encoded training difference vector into the encoder of the pre-trained speech object recognition model, and using the self-attention mechanism component of the encoder to capture the correlation between each of the third elements in the position-encoded training difference vector to obtain the correlation between the third elements; The feedforward neural network of the encoder is used to perform encoding conversion on the correlation between the third elements to obtain the training difference vector encoding.
6. The method for speech object recognition according to claim 1, characterized in that: The step of using a multilayer perceptron and a preset activation function to calculate a verification score according to the target difference vector encoding includes: Inputting the target difference vector code into a multi-layer perceptron, and processing the target difference vector code using a weight matrix set in each layer of the multi-layer perceptron to obtain an output value; The output value is calculated using a preset activation function to obtain a verification score.
7. The method for speech object recognition according to claim 1, characterized in that: The step of determining the probability that the target speech object appears in the speech to be recognized according to the verification score comprises: Obtaining a preset score threshold, and comparing the verification score with the score threshold; If the verification score is greater than or equal to the score threshold, it is determined that the probability that the target voice object appears in the voice to be recognized is high; If the verification score is less than the score threshold, it is determined that the probability that the target speech object appears in the speech to be recognized is low.
8. A speech object recognition device, characterized in that: The device comprises: An acquisition module is used to acquire the speech to be recognized, and use the trained speech object recognition model to extract features of the speech to be recognized to obtain a speech feature vector; A subtraction module, used for subtracting the speech feature vector from the registered feature vector in the trained speech object recognition model to obtain a target difference vector, wherein the registered feature vector refers to a feature vector generated by pre-registering and embedding the trained speech object recognition model with a separate speech of a target speech object; An encoding module, used for adding position coding to the target difference vector, inputting the position-coded target difference vector into the encoder of the trained speech object recognition model for processing, and obtaining the target difference vector coding; A calculation module is used to use a multi-layer perceptron and a preset activation function to calculate a verification score according to the target difference vector encoding, and determine the probability that the target speech object appears in the speech to be recognized according to the verification score.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the speech object recognition method according to any one of claims 1 to 7 when executing the computer-readable instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the voice object recognition method according to any one of claims 1 to 7 are implemented.