A method, apparatus, electronic device, and storage medium for processing voiceprint information
By automatically updating the voiceprint embed code in the voiceprint embed code collection, the inaccurate voice recognition caused by user status changes is solved, the recognition accuracy and user experience are improved, and the steps of manual update are reduced.
Patent Information
- Application Number
- CN202210657997.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-06-10
AI Technical Summary
Due to changes in user status, different physiological development periods and different language types, the voice of the same user changes, and it is difficult for the terminal to accurately identify the user's voice commands in a timely and accurate manner, which affects the user experience. Moreover, users need to frequently manually update the voiceprint embed code to increase the operation burden.
By updating the voiceprint embedding code in the voiceprint embedding code set, the audio information collected by the first terminal and the version identification code of the voiceprint recognition service process are used to calculate the similarity and reliability parameters of the voiceprint embedding code, and automatically update the voiceprint embedding code to ensure its accuracy.
It improves the accuracy of voiceprint recognition, reduces the steps of users to manually update voiceprint embed code, improves user experience, and flexibly selects the usage environment according to different versions of voiceprint recognition service processes.
Smart Images

Figure CN115171660B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to voiceprint information processing technology, and in particular to a method, device, electronic device and storage medium for processing voiceprint information. Background Art
[0002] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction. Among them, voice has become one of the most convenient human-computer interaction methods. However, due to different user states, different physiological development periods, and different language types used, the voice of the same user is often prone to change, and the terminal cannot accurately identify the user's voice command in a timely manner, which affects the user's experience of using voice information. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, device, electronic device and storage medium for processing voiceprint information, which can update the voiceprint embedding codes in the voiceprint embedding code set to ensure the accuracy of the voiceprint embedding codes, improve the accuracy of voiceprint recognition using the voiceprint embedding codes, and at the same time reduce the cumbersome steps of manual updating of the voiceprint embedding codes by users, so that users can obtain a better experience.
[0004] The technical solution of the embodiments of the present invention is implemented as follows:
[0005] Embodiments of the present invention provide a method for processing voiceprint information, including:
[0006] Collecting first audio information of a first target object through a first terminal;
[0007] Parsing the first audio information to obtain first voiceprint information of the first target object;
[0008] Processing the first voiceprint information through a voiceprint recognition service process to obtain a first voiceprint embedding code corresponding to the first voiceprint information;
[0009] Searching for a voiceprint embedding code set that matches the first terminal identification code according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process;
[0010] Calculating the similarity between the first voiceprint embedding code and each voiceprint embedding code in the voiceprint embedding code set;
[0011] When the similarity is greater than or equal to the similarity threshold, calculate the reliability parameter corresponding to the first voiceprint information;
[0012] Update the voiceprint embedding codes in the voiceprint embedding code set according to the reliability parameter corresponding to the first voiceprint information.
[0013] An embodiment of the present invention further provides a voiceprint information processing device, which is characterized in that the device includes:
[0014] An information transmission module, configured to collect first audio information of a first target object through a first terminal;
[0015] An information processing module, configured to parse the first audio information to obtain first voiceprint information of the first target object;
[0016] The information processing module is configured to process the first voiceprint information through a voiceprint recognition service process to obtain a first voiceprint embedding code corresponding to the first voiceprint information;
[0017] The information processing module is configured to find a voiceprint embedding code set that matches the first terminal identification code according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process;
[0018] The information processing module is configured to calculate the similarity between the first voiceprint embedding code and each voiceprint embedding code in the voiceprint embedding code set;
[0019] The information processing module is configured to calculate the reliability parameter corresponding to the first voiceprint information when the similarity is greater than or equal to the similarity threshold;
[0020] The information processing module is configured to update the voiceprint embedding codes in the voiceprint embedding code set according to the reliability parameter corresponding to the first voiceprint information.
[0021] In the above solution,
[0022] The information processing module is configured to configure the first terminal identification code for the first terminal;
[0023] The information processing module is configured to, when collecting audio information of a second target object through the first terminal, configure a second target object identification code for the second target object through the voiceprint recognition service process, and establish a mapping relationship between the second target object identification code and the first terminal identification code;
[0024] The information processing module is configured to obtain at least two pieces of audio information of the second target object through the voiceprint recognition service process;
[0025] The information processing module is configured to calculate second voiceprint embedding codes respectively corresponding to the at least two pieces of audio information through the voiceprint recognition service process;
[0026] The information processing module is configured to calculate the average value of the second voiceprint embedding codes respectively corresponding to the at least two pieces of audio information;
[0027] The information processing module is configured to save the average value of the second voiceprint embedding codes in the voiceprint embedding code set, and mark it with the first terminal identification code, the second target object identification code, and the version identification code of the voiceprint recognition service process.
[0028] In the above solution,
[0029] The information processing module is configured to trigger a voice information recognition model through the voiceprint recognition service process;
[0030] The information processing module is configured to extract the pinyin corresponding to each character in the first voiceprint information and the intonation corresponding to each character in the first voiceprint information through the voice information recognition model according to the recognition environment of the target voice information;
[0031] The information processing module is configured to determine a single-character pronunciation feature vector at each character level in the first voiceprint information according to the pinyin corresponding to each character in the first voiceprint information and the intonation corresponding to each character in the first voiceprint information;
[0032] The information processing module is configured to perform a combination process on the single-character pronunciation feature vectors corresponding to each character in the first voiceprint information through the phonetic encoder network in the voice information recognition model to form a sentence-level pronunciation feature vector;
[0033] The information processing module is configured to use the sentence-level pronunciation feature vector as the first voiceprint embedding code.
[0034] In the above solution,
[0035] The information processing module is configured to perform a channel conversion process on the first audio information to form mono audio data;
[0036] The information processing module is configured to perform a short-time Fourier transform on the mono audio data based on a window function corresponding to the voice information recognition model to form a corresponding Mel spectrogram;
[0037] The information processing module is configured to determine a corresponding input triple sample based on the Mel spectrogram and input the input triple sample into the voice information recognition model;
[0038] The information processing module is configured to process the input triple samples crossly through the convolutional layer and the max pooling layer of the voice information recognition model to obtain the downsampling results of different input triple samples;
[0039] The information processing module is configured to perform normalization processing on the downsampling results of the different input triple samples through the fully connected layer of the voice information recognition model to obtain the first voiceprint embedding code.
[0040] In the above solution,
[0041] The information processing module is configured to obtain the first text recognition result corresponding to the first audio information, the identification information of the first target object, and the similarity corresponding to the first voiceprint information;
[0042] The information processing module is configured to search for the text frequency information list of the first target object according to the identification information of the first target object;
[0043] The information processing module is configured to calculate the display times of the first text recognition result appearing in the text frequency information list;
[0044] The information processing module is configured to calculate the total frequency in the text frequency information list;
[0045] When the similarity is greater than or equal to the similarity threshold, and the display times of the first text recognition result appearing in the text frequency information list is greater than or equal to 1, the information processing module is configured to use the ratio of the display times to the total frequency as the reliability parameter corresponding to the first voiceprint information.
[0046] In the above solution,
[0047] When the reliability parameter corresponding to the first voiceprint information is 0, the information processing module is configured to update the text frequency information list of the first target object and update the display times of the first text recognition result appearing in the text frequency information list.
[0048] In the above solution,
[0049] When the reliability parameter corresponding to the first voiceprint information is greater than or equal to the reliability parameter threshold, and the number of characters of the first text recognition result corresponding to the first audio information is greater than or equal to the character number threshold, the information processing module is configured to obtain the original voiceprint embedding code corresponding to the similarity;
[0050] The information processing module is configured to obtain the first weight parameter of the original voiceprint embedding code and the second weight parameter of the first voiceprint embedding code;
[0051] The information processing module is configured to calculate a weighted average of the original voiceprint embedding code and the first voiceprint embedding code according to the first weight parameter and the second weight parameter, so as to obtain a third voiceprint embedding code;
[0052] The information processing module is configured to update the original voiceprint embedding code in the voiceprint embedding code set through the third voiceprint embedding code.
[0053] In the above solution,
[0054] The information processing module is configured to save the voiceprint embedding code set in a cloud server;
[0055] The information processing module is configured to detect the processing permission of the second terminal when first audio information of a first target object is collected by the second terminal;
[0056] When the second terminal meets the requirements of the processing permission, the voiceprint embedding code set in the cloud server is saved to the second terminal.
[0057] An embodiment of the present invention further provides an electronic device, where the electronic device includes:
[0058] A memory, configured to store executable instructions;
[0059] A processor, configured to implement the foregoing voiceprint information processing method when running the executable instructions stored in the memory.
[0060] An embodiment of the present invention further provides a computer-readable storage medium, storing executable instructions, where the executable instructions implement the foregoing voiceprint information processing method when being executed by a processor.
[0061] The embodiment of the present invention has the following beneficial effects:
[0062] In an embodiment of the present invention, the first terminal collects the first audio information of the first target object; parses the first audio information to obtain the first voiceprint information of the first target object; processes the first voiceprint information through a voiceprint recognition service process to obtain a first voiceprint embedding code corresponding to the first voiceprint information; looks up a set of voiceprint embedding codes matching the first terminal identification code according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process; calculates the similarity between the first voiceprint embedding code and each voiceprint embedding code in the set of voiceprint embedding codes; when the similarity is greater than or equal to a similarity threshold, calculates a reliability parameter corresponding to the first voiceprint information; and updates the voiceprint embedding codes in the set of voiceprint embedding codes according to the reliability parameter corresponding to the first voiceprint information. Thus, by updating the voiceprint embedding codes in the set of voiceprint embedding codes, the accuracy of the voiceprint embedding codes can be ensured, the accuracy of voiceprint recognition using the voiceprint embedding codes can be improved, and at the same time, the cumbersome steps of manually updating the voiceprint embedding codes by the user can be reduced, enabling the user to obtain a better user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a schematic diagram of the usage environment of a voiceprint information processing method provided by an embodiment of the present invention;
[0064] Figure 2 is a schematic diagram of the composition structure of an electronic device provided by an embodiment of the present invention;
[0065] Figure 3 is an optional flowchart of a voiceprint information processing method provided by an embodiment of the present invention;
[0066] Figure 4 is an optional schematic diagram of the structure of a voice information recognition model in an embodiment of the present invention;
[0067] Figure 5 is an optional flowchart of a voiceprint information processing method provided by an embodiment of the present invention;
[0068] Figure 6 is a schematic diagram of the processing process of the voice information recognition model for audio in an embodiment of the present invention;
[0069] Figure 7 is a schematic diagram of the process of updating the voiceprint embedding code in an embodiment of the present invention;
[0070] Figure 8 is a schematic diagram of the usage scenario of a voiceprint information processing method provided by an embodiment of the present invention;
[0071] Figure 9 is an optional flowchart of a voiceprint information processing method provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0073] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0074] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are subject to the following explanations.
[0075] 1) Short-time Fourier transform: The short-time Fourier transform (STFT) is a mathematical transform related to the Fourier transform, used to determine the frequency and phase of the sine wave in the local region of a time-varying signal.
[0076] 2) Mel bank features: Since the obtained spectrogram is large, in order to obtain appropriate-sized voice features, it is usually transformed into Mel bank features through a Mel-scale filter bank.
[0077] 3) Neural network (Neural Network, NN): An artificial neural network (Artificial Neural Network, ANN), abbreviated as neural network or neural-like network, is a mathematical model or computational model that mimics the structure and function of a biological neural network (the central nervous system of an animal, especially the brain) in the fields of machine learning and cognitive science, and is used to estimate or approximate a function.
[0078] 4) Speech recognition (SR Speech Recognition): Also known as automatic speech recognition (ASR Automatic Speech Recognition), computer speech recognition (CSR Computer Speech Recognition) or speech-to-text recognition (STT Speech To Text), its goal is to automatically convert the speech content of humans into corresponding text using a computer.
[0079] 5) Terminals, including but not limited to: general terminals and dedicated terminals, where the general terminals maintain long connections and / or short connections with the sending channel, and the dedicated terminals maintain long connections with the sending channel.
[0080] 6) Clients, which are carriers for implementing specific functions in terminals. For example, a mobile client (APP) is a carrier for specific functions in a mobile terminal, such as implementing a voice wake-up function.
[0081] 7) Mini Programs, which are programs developed based on a front-end-oriented language (such as JavaScript) and implemented services in Hyper Text Markup Language (HTML) pages. They are software downloaded by a client (such as a browser or any client embedded with a browser core) via a network (such as the Internet) and interpreted and executed in the browser environment of the client, saving the steps of installation in the client. For example, by waking up the Mini Program in the terminal through a voice command, Mini Programs for various services such as ticket purchase, task processing and production, and data display can be downloaded and run in a social network client.
[0082] Figure 1 It is a schematic diagram of the usage scenario of the voiceprint information processing method provided by the embodiments of the present invention. Refer to Figure 1 , corresponding clients capable of performing different functions are set on terminals (including terminal 10-1 and terminal 10-2). Among them, through the set clients, the terminals (including terminal 10-1 and terminal 10-2) can obtain different corresponding information from the corresponding server 200 through the network 300 for browsing. The terminals are connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and uses a wireless link to achieve data transmission. Among them, the terminals (including terminal 10-1 and terminal 10-2) can be woken up through the voice commands of users. Specifically, the key technologies of voice technology are automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Among them, voice technology can be applied to electronic devices to achieve the function of waking up electronic devices, that is, voice wake-up technology. Usually, voice wake-up is achieved by setting a fixed wake-up word. After the user says the wake-up word, the voice recognition function on the terminal will be in a working state, otherwise it will be in a sleep state.
[0083] Among them, the intelligent device wake-up method provided by the embodiments of the present application is implemented based on artificial intelligence. Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0084] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0085] In the embodiments of the present application, the artificial intelligence software technologies mainly involved include the above-mentioned speech processing technology and machine learning, etc. For example, it may involve the automatic speech recognition (ASR) technology in speech technology, which includes speech signal preprocessing, speech signal frequency domain analysis, speech signal feature extraction, speech signal feature matching / recognition, speech training, etc.
[0086] For example, it may involve Machine Learning (ML). Machine learning is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning usually includes technologies such as Deep Learning. Deep learning includes artificial neural networks, such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Deep neural network (DNN), etc.
[0087] It can be understood that this method can be applied to intelligent devices. An intelligent device can be any device with a voice wake-up function, such as an intelligent terminal, a smart home device (such as a smart speaker, a smart washing machine, etc.), a smart wearable device (such as a smart watch), an in-vehicle intelligent central control system (waking up small programs that execute different tasks in the terminal through voice commands), or an AI intelligent medical device (triggered by voice commands for wake-up).
[0088] As an example, the terminal (including terminal 10-1 and terminal 10-2) is used to deploy a voiceprint information processing device to implement the voiceprint information processing method provided by the present invention, so as to collect the first audio information of the first target object through the first terminal; parse the first audio information to obtain the first voiceprint information of the first target object; process the first voiceprint information through a voiceprint recognition service process to obtain a first voiceprint embedding code corresponding to the first voiceprint information; according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process, search for a set of voiceprint embedding codes that match the first terminal identification code; calculate the similarity between the first voiceprint embedding code and each voiceprint embedding code in the set of voiceprint embedding codes; when the similarity is greater than or equal to a similarity threshold, calculate a reliability parameter corresponding to the first voiceprint information; update the voiceprint embedding codes in the set of voiceprint embedding codes according to the reliability parameter corresponding to the first voiceprint information, and finally realize the recognition and execution of the voice wake-up instruction.
[0089] The following will elaborate on the structure of the voiceprint information processing device according to the embodiments of the present invention. The voiceprint information processing device can be implemented in various forms, such as a dedicated terminal with voiceprint information processing functions, or a mobile phone or tablet computer equipped with voiceprint information processing functions. For example, the terminal in the preamble Figure 1 in the above. Figure 2 FIG. is a schematic diagram of the composition structure of the voiceprint information processing device provided by the embodiments of the present invention. It can be understood that Figure 2 only the exemplary structure of the voiceprint information processing device is shown, rather than all structures. According to needs, the Figure 2 partial structure or all structures shown can be implemented.
[0090] The voiceprint information processing device provided by the embodiments of the present invention includes: at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. Each component in the voiceprint information processing device is coupled together through a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 205.
[0091] Among them, the user interface 203 may include a display, a keyboard, a mouse, a trackball, a click wheel, a button, a touchpad, or a touch screen, etc.
[0092] It can be understood that the memory 202 can be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The memory 202 in the embodiments of the present invention can store data to support the operation of the terminal (such as 10-1). Examples of these data include: any computer programs for operating on the terminal (such as 10-1), such as an operating system and application programs. Among them, the operating system contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs can include various application programs.
[0093] In some embodiments, the voiceprint information processing device provided by the embodiments of the present invention may be implemented in a combination of software and hardware. As an example, the voice processing model provided by the embodiments of the present invention may be a processor in the form of a hardware decoding processor, which is programmed to execute the semantic processing method of the voice processing model provided by the embodiments of the present invention. For example, a processor in the form of a hardware decoding processor may employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0094] As an example of the voiceprint information processing device provided by the embodiments of the present invention being implemented in a combination of software and hardware, the voiceprint information processing device provided by the embodiments of the present invention may be directly embodied as a combination of software modules executed by the processor 201. The software modules may be located in a storage medium, and the storage medium is located in the memory 202. The processor 201 reads the executable instructions included in the software modules in the memory 202 and, in combination with necessary hardware (for example, including the processor 201 and other components connected to the bus 205), completes the semantic processing method of the voice processing model provided by the embodiments of the present invention.
[0095] As an example, the processor 201 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0096] As an example of the hardware implementation of the voiceprint information processing device provided by the embodiments of the present invention, the device provided by the embodiments of the present invention can be directly implemented by using a processor 201 in the form of a hardware decoding processor. For example, it can be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components to execute the semantic processing method of the voice processing model provided by the embodiments of the present invention.
[0097] The memory 202 in the embodiments of the present invention is used to store various types of data to support the operation of the voiceprint information processing device. Examples of these data include: any executable instructions for operating on the voiceprint information processing device, such as executable instructions. The program for implementing the semantic processing method of the voice processing model according to the embodiments of the present invention can be included in the executable instructions.
[0098] In other embodiments, the voiceprint information processing device provided by the embodiments of the present invention can be implemented in software. Figure 2 Shown is the voiceprint information processing device stored in the memory 202, which can be software in the form of programs and plugins, etc., and includes a series of modules. As an example of the program stored in the memory 202, it can include a voiceprint information processing device, and the following software modules are included in the voiceprint information processing device: an information transmission module 2081, an information processing module 2082. When the software modules in the voiceprint information processing device are read into the RAM by the processor 201 and executed, the semantic processing method of the voice processing model provided by the embodiments of the present invention will be implemented. The functions of each software module in the voiceprint information processing device in the embodiments of the present invention are introduced below, specifically including:
[0099] The information transmission module 2081 is used to collect the first audio information of the first target object through the first terminal.
[0100] The information processing module 2082 is used to parse the first audio information to obtain the first voiceprint information of the first target object.
[0101] The information processing module 2082 is used to process the first voiceprint information through the voiceprint recognition service process to obtain the first voiceprint embedding code corresponding to the first voiceprint information.
[0102] The information processing module 2082 is configured to find a set of voiceprint embedding codes that match the first terminal identification code according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process.
[0103] The information processing module 2082 is configured to calculate the similarity between the first voiceprint embedding code and each voiceprint embedding code in the set of voiceprint embedding codes.
[0104] The information processing module 2082 is configured to calculate a reliability parameter corresponding to the first voiceprint information when the similarity is greater than or equal to a similarity threshold.
[0105] The information processing module 2082 is configured to update the voiceprint embedding codes in the set of voiceprint embedding codes according to the reliability parameter corresponding to the first voiceprint information.
[0106] According to Figure 2 For the electronic device shown, in one aspect of the present application, the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes different embodiments and combinations of embodiments provided in various optional implementation manners of the above voiceprint information processing method.
[0107] Combined with Figure 2 The voiceprint information processing method provided by the embodiments of the present invention will be described with reference to the electronic device 20 shown. Before introducing the voiceprint information processing method provided by the present invention, the defects of the related technology will be introduced first.
[0108] The theoretical basis of voiceprint recognition is that each voice has unique characteristics, and different people's voices can be effectively distinguished through these characteristics. This unique characteristic is mainly determined by two factors. The first is the size of the vocal cavity, specifically including the throat, nasal cavity, oral cavity, etc. The shape, size, and position of these organs determine the magnitude of the vocal cord tension and the range of sound frequencies. Therefore, although different people say the same words, the frequency distribution of their voices is different, and some sound low and some sound loud. Everyone's vocal cavity is different, just like fingerprints, and everyone's voice has unique characteristics.
[0109] The second factor that determines the voice characteristics is the way the vocal organs are manipulated. The vocal organs include the lips, teeth, tongue, soft palate, and palatal muscles, etc. Their interaction will produce clear speech. And the way they cooperate is randomly learned by people through communication with people around them. When people are learning to speak, by imitating the speaking ways of different people around them, they will gradually form their own voiceprint characteristics.
[0110] However, due to different user states (vocal cord damage), different physiological development periods (voice changes during puberty), and different types of languages used (different pronunciations of the same word in different dialects), it is often easy for the voice of the same user to change. The terminal cannot accurately recognize the user's voice command in a timely manner, which affects the user's experience of using voice information. If the user frequently updates the voiceprint embedding code manually, it will increase the user's operation burden.
[0111] To overcome the above defects, refer to Figure 3 , Figure 3 which is an optional flowchart of the voiceprint information processing method provided by an embodiment of the present invention. It can be understood that Figure 3 the steps shown can be executed by various electronic devices running the voiceprint information processing device. For example, it can be a terminal, a server, or a server cluster with voiceprint information processing functions. When the voiceprint information processing device runs in the terminal, it can trigger the instant messaging client in the terminal or a small program in the in-vehicle system to process audio information, so as to improve the speed of voiceprint information processing. The user can also operate the electronic device through the wake-up word in the voice command, and the electronic device executes the task matching the audio information feature; among them, the dedicated device with the voiceprint information processing device can be encapsulated in Figure 1 the terminal shown in Figure 2 to execute the corresponding software module in the voiceprint information processing device shown in the previous Figure 3 The user can obtain the task information and display it through the corresponding client. The following will describe the steps shown in
[0112] Step 301: The voiceprint information processing device collects the first audio information of the first target object through the first terminal.
[0113] In some embodiments of the present invention, before executing step 301, the terminal needs to perform voiceprint information registration on the received voice command, so as to determine which user the voice command comes from after receiving the voice command. Specifically, the voiceprint information registration can be achieved through the following methods:
[0114] Configure a first terminal identification code for the first terminal; when collecting audio information of a second target object through the first terminal, configure a second target object identification code for the second target object through the voiceprint recognition service process, and establish a mapping relationship between the second target object identification code and the first terminal identification code; obtain at least two pieces of audio information of the second target object through the voiceprint recognition service process; calculate second voiceprint embedding codes respectively corresponding to the at least two pieces of audio information through the voiceprint recognition service process; calculate the average value of the second voiceprint embedding codes respectively corresponding to the at least two pieces of audio information; save the average value of the second voiceprint embedding codes in the voiceprint embedding code set, and mark it with the first terminal identification code, the second target object identification code, and the version identification code of the voiceprint recognition service process. For example: taking the first terminal as a smart speaker, for example, users can perform voice control on electronic devices through corresponding voice commands and execute tasks matching the characteristics of audio information to replace traditional manual operations. Specifically, for various operations of different types of electronic devices, corresponding wake-up words can be pre-configured. Users only need to say the wake-up word corresponding to the required task operation through a voice command to control the electronic device to execute the corresponding operation in a voice control manner. For example: when the electronic device is an in-vehicle intelligent central control system, the wake-up word of the electronic device is "play a song". Since the intelligent device can collect audio data at any time, the electronic device can collect the audio data "play music", thereby identifying whether "play music" is a wake-up word and executing a task matching the characteristics of the audio information through the electronic device to realize the electronic device playing a song. At this time, the second target object can be all users using the smart speaker (denoted as speaker users). The speaker users initiate user creation on the first terminal, and the recognition service assigns a unique user identification code u and binds it to the unique first terminal identification code D; then the speaker users perform audio recording. The speaker users record 3 segments of audio of fixed text on the first terminal, save them in the audio storage service, and mark them with "device unique identification code + user unique identification code". The voiceprint recognition service process uses these 3 segments of recorded audio to calculate three voiceprint embedding codes through the voice information recognition model, and calculates the average value of the 3 voiceprint embedding codes as the final voiceprint embedding code e of the speaker user, saves it in the embedding code storage service, and marks it with "device unique identification code D + user unique identification code u + voiceprint recognition service version number vn" to complete the voiceprint registration process of the speaker user.
[0115] Step 302: The voiceprint information processing device analyzes the first audio information to obtain the first voiceprint information of the first target object.
[0116] Step 303: The voiceprint information processing device processes the first voiceprint information through the voiceprint recognition service process to obtain a first voiceprint embedding code corresponding to the first voiceprint information.
[0117] Continue to refer to Figure 4 , Figure 4 FIG. 1 is an optional structural schematic diagram of a voice information recognition model in an embodiment of the present invention. Among them, the Encoder includes: N = 6 identical layers, and each layer contains two sub-layers. The first sub-layer is the multi-head attention layer, and then a simple fully connected layer. A residual connection and normalization are added to each sub-layer.
[0118] Combined with Figure 4 the model structure shown in FIG. 2, refer to Figure 5 , Figure 5 FIG. 3 is an optional flowchart of a voice information recognition method provided by an embodiment of the present invention. It can be understood that Figure 5 the steps shown in FIG. 3 can be executed by various electronic devices running the voice information recognition device to obtain the phonetic feature vector and the glyph feature vector corresponding to the voice information to be recognized, specifically including the following steps:
[0119] Step 501: According to the recognition environment of the target voice information, through the phonetic encoder network in the voice information recognition model, extract the pinyin corresponding to each character in the voice information to be recognized, and the intonation corresponding to each character in the voice information to be recognized.
[0120] Step 502: Determine the single-character pronunciation feature vector at the character level in the voice information to be recognized according to the pinyin corresponding to each character in the voice information to be recognized and the intonation corresponding to each character in the voice information to be recognized.
[0121] Step 503: Through the phonetic encoder network in the voice information recognition model, perform a combination process on the single-character pronunciation feature vector corresponding to each character in the voice information to be recognized to form a pronunciation feature vector at the sentence level.
[0122] Step 504: Use the pronunciation feature vector at the sentence level as the first voiceprint embedding code.
[0123] In some embodiments of the present invention, when performing phonetic recognition processing, a 4-layer Transformer model is used for sentence-level phonetic encoding, and the input is the output of the word-level phonetic encoder. It should be noted that the Gated Recurrent Unit (GRU) is a model with fewer parameters than LSTM that can handle sequence information well. Next, the fused features will be input into a feed-forward neural network to process the effective information of other features. The recognition of incorrect characters is regarded as a problem of predicting the occurrence probability, and the sigmoid function (logical function) is used as the output layer. The loss function is the standard cross-entropy loss. Reference can be made to Formula 1:
[0124]
[0125] Among them, the GRU layer is used for deep feature extraction. The GRU layer can also be omitted and replaced by multiple concatenated feed-forward neural network layers, which can also effectively process and fuse features.
[0126] In some embodiments of the present invention, the method for obtaining the first voiceprint embedding code further includes:
[0127] Performing channel conversion processing on the first audio information to form mono audio data; performing short-time Fourier transform on the mono audio data based on a window function corresponding to the speech information recognition model to form a corresponding Mel spectrogram; determining a corresponding input triple sample based on the Mel spectrogram, and inputting the input triple sample into the speech information recognition model; processing the input triple sample crosswise through the convolutional layer and the max-pooling layer of the speech information recognition model to obtain the downsampling results of different input triple samples; performing normalization processing on the downsampling results of different input triple samples through the fully connected layer of the speech information recognition model to obtain the first voiceprint embedding code. Figure 6This is a schematic diagram of the processing process of the voice information recognition model for audio in the embodiments of the present invention. Feature extraction can be performed through the VGGish network. Among them, the feature extraction of the voice information recognition model can be implemented through the Visual Geometry Group (VGGish). For example, for the audio information in the user video detection information, the audio file can be extracted to obtain the audio file. For the audio file, the corresponding Mel spectrogram is obtained. Then, for the Mel spectrogram, audio features are extracted through the Vggish network. The extracted vectors are clustered and encoded through the NetVlad (Net Vector of Locally Aggregated Descriptors) to obtain the audio feature vector. NetVlad can save the distance between each feature point and the nearest clustering center and use it as a new feature. Since the obtained audio information may have noise in the environment, in order to better calculate the first voiceprint embedding code, the audio can be first resampled to 16KHz single-channel audio; then a 25ms Hann window is used with a frame shift of 10ms, and the audio is subjected to short-time Fourier transform with a periodic Hann window to obtain the corresponding spectrogram; the spectrogram is mapped into a 64-order mel filter bank to calculate the mel spectrogram, where the range of mel bins is 125 - 7500Hz; log(mel-spectrum + 0.01) is calculated to obtain a stable mel spectrogram. The added bias of 0.01 is to avoid taking the logarithm of 0; the obtained features are framed in units of 0.96s without frame overlap. Each frame contains 64 mel bands and has a duration of 10ms (a total of 96 frames), thereby realizing the extraction of the corresponding Mel spectrogram. Finally, through the processing of the Mel spectrogram, a clear first voiceprint embedding code is obtained, ensuring the accuracy of the first voiceprint embedding code.
[0128] Step 304: The voiceprint information processing device searches for a set of voiceprint embedding codes that match the first terminal identification code according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process.
[0129] In some embodiments of the present invention, the version identification code of the voiceprint recognition service process can be used to indicate the storage location of the voiceprint embedding code set. For example, when the electronic device is an in-vehicle intelligent central control system, when the version identification code of the voiceprint recognition service process is version 1.1 (domestic version) or version 19.0.1 (overseas version), the voiceprint embedding code set is only stored in the storage device of the in-vehicle intelligent central control system for performing the voiceprint information processing method provided in this application. When the version identification code of the voiceprint recognition service process is version 3.1 (domestic version) or version 21.0.1 (overseas version), the voiceprint embedding code set is not only stored in the storage device of the in-vehicle intelligent central control system, but can also be stored in the corresponding cloud network (cloud server cluster) for performing the voiceprint information processing method provided in this application. When the user replaces the vehicle, the stored voiceprint embedding code set can be obtained from the cloud network and applied to the in-vehicle intelligent central control system, avoiding the defect of the user manually copying the voiceprint embedding code set and making it more convenient for the user to use.
[0130] In some embodiments of the present invention, when the user needs to use the cloud network to facilitate the storage of the voiceprint embedding code set, the version identification code of the voiceprint recognition service process can be adjusted by purchasing the voiceprint recognition service process. For example, upgrading version 1.1 (domestic version) to version 3.1 and above (domestic version); or upgrading version 19.0.1 (overseas version) to version 21.0.1 and above (overseas version) to meet the user's usage requirements. At the same time, the user can flexibly select the domestic version voiceprint recognition service process or the overseas version voiceprint recognition service process according to the different usage regions to meet the legal requirements for voiceprint information collection in the usage region.
[0131] Step 305: The voiceprint information processing device calculates the similarity between the first voiceprint embedding code and each voiceprint embedding code in the voiceprint embedding code set.
[0132] Still taking the smart speaker in the previous embodiment as an example, speaker user A verifies audio recording. The speaker user records an audio on the first terminal and carries the device unique identification code D A and the user unique identification code u A , and initiates a voiceprint verification to the voiceprint service. Then the voiceprint embedding code is calculated. The voiceprint recognition service process calculates the voiceprint embedding code e A of the audio through the speech recognition model, and at the same time, according to the unique identification code D A of the first terminal and the version number v n of the voiceprint recognition service process, obtains the embedding codes E = {e1, e2,..., en} of all speaker users on the first terminal. The voiceprint recognition service process calculates the embedding code e AThe cosine similarity with all the voiceprint embedding codes E of the speakers of the first terminal is taken, and the maximum similarity value C is selected, corresponding to the voiceprint embedding code e i .
[0133] In some embodiments of the present invention, the similarity threshold T is set to 0.6. If C is greater than or equal to the threshold T = 0.6, the verification is successful, and the voiceprint recognition service process returns e i The corresponding speaker user identification code u i And the similarity value C to the first terminal; if C is less than the threshold T, the verification fails, and the voiceprint recognition service process returns a verification failure prompt to the first terminal, and the smart speaker user can be prompted to trigger other verification methods to change the voiceprint embedding code.
[0134] Step 306: When the similarity is greater than or equal to the similarity threshold, the voiceprint information processing device calculates the reliability parameter corresponding to the first voiceprint information.
[0135] In some embodiments of the present invention, due to the change of voiceprint information caused by the physiological development of the user, when calculating the reliability parameter corresponding to the first voiceprint information, when the similarity of each voiceprint embedding code in the obtained voiceprint embedding code set is greater than or equal to the similarity threshold, and the number of similarities is greater than or equal to 2, in order to ensure the security of the voiceprint embedding code change, different verification strategies can be triggered according to the version identification code of the voiceprint recognition service process. When the voiceprint recognition service process is version 3.1 and above (domestic version), or version 21.0.1 and above (overseas version), the face recognition information of the user can be sent to the cloud server network. When the face recognition information passes the verification of the cloud server network, the reliability parameter corresponding to the first voiceprint information is calculated to update the voiceprint embedding code. When the voiceprint recognition service process is version 3.1 and below (domestic version), or version 21.0.1 and below (overseas version), since the cloud server network cannot be used, only the local voiceprint embedding code can be updated. Therefore, in order to ensure the security of the voiceprint embedding code change, it is necessary to re-collect the first audio information and calculate the similarity between the first voiceprint embedding code and each voiceprint embedding code in the voiceprint embedding code set.
[0136] In some embodiments of the present invention, the reliability parameter corresponding to the first voiceprint information can be calculated in the following way:
[0137] Obtain the first text recognition result corresponding to the first audio information, the identification information of the first target object, and the similarity corresponding to the first voiceprint information; according to the identification information of the first target object, search for the text frequency information list of the first target object; calculate the number of display times that the first text recognition result appears in the text frequency information list; calculate the total frequency in the text frequency information list; when the similarity is greater than or equal to the similarity threshold, and the number of display times that the first text recognition result appears in the text frequency information list is greater than or equal to 1, use the ratio of the number of display times to the total frequency as the reliability parameter corresponding to the first voiceprint information. Combined with the foregoing Figure 4 As shown in the structure of the speech recognition model, the audio information can be text-recognized through the speech recognition model, and the audio information can also be converted through the text-to-speech conversion server to obtain the corresponding audio information feature set; the audio information feature set is processed by the first neural network to determine a feature vector with the same number of test feature frames as that of the test feature frames, and the feature vector is averaged to extract the corresponding wake-up word feature. Among them, when the user adds a new wake-up word, the user can only input audio information, and the audio information is converted through the text-to-speech conversion server deployed in the cloud. Among them, the embodiments of the present invention can be implemented in combination with cloud technology or blockchain network technology. Cloud technology is a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to realize the calculation, storage, processing, and sharing of data. It can also be understood as the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. The background services of the technology network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. Therefore, cloud technology needs to be supported by cloud computing.
[0138] It should be noted that cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to the user to be infinitely expandable, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool platform will be established, abbreviated as a cloud platform, generally referred to as Infrastructure as a Service (IaaS). Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (which can be virtual machines, including operating systems), storage devices, and network devices.
[0139] In some embodiments of the present invention, the TTS server in the cloud can generate N different wake-up word voices (pronunciations) using the wake-up text, forming feature vectors with different frame lengths. For example, the user can arbitrarily modify the audio information according to different usage scenarios. The TTS server converts each character included in the audio information into a syllable identifier according to the pronunciation dictionary to extract the corresponding wake-up word features.
[0140] Step 307: The voiceprint information processing device updates the voiceprint embedding codes in the voiceprint embedding code set according to the reliability parameter corresponding to the first voiceprint information.
[0141] Reference Figure 7 , Figure 7 is a schematic diagram of the process of updating the voiceprint embedding code in the embodiment of the present invention, which specifically includes the following steps:
[0142] Step 701: When the reliability parameter corresponding to the first voiceprint information is greater than or equal to the reliability parameter threshold, and the number of characters in the first text recognition result corresponding to the first audio information is greater than or equal to the character number threshold, obtain the original voiceprint embedding code corresponding to the similarity.
[0143] Step 702: Obtain the first weight parameter of the original voiceprint embedding code and the second weight parameter of the first voiceprint embedding code.
[0144] Step 703: Calculate the weighted average of the original voiceprint embedding code and the first voiceprint embedding code according to the first weight parameter and the second weight parameter to obtain the third voiceprint embedding code.
[0145] Step 704: Update the original voiceprint embedding code in the voiceprint embedding code set through the third voiceprint embedding code.
[0146] Combined with the previous embodiment, the first terminal calculates the text recognition result T of the audio information, the identifier u of the target object i and the similarity value C to obtain a voiceprint recognition result reliability value R, 0 < R < 1. Send R and T to the voiceprint recognition service process, where the calculation method of R is: R = m / N, where m is the occurrence frequency of T in L, L is the text frequency information list of the first target object, and N is the total frequency of the text frequency information list.
[0147] The voiceprint recognition service process determines whether to update the voiceprint embedding code according to the obtained reliability value R and the verified audio text recognition result T. In some embodiments of the present invention, the similarity threshold R can be 0.7, that is: when R > 0.7 and the number of words in T is greater than 10, the embedding code is updated; otherwise, it is not updated.
[0148] In some embodiments of the present invention, the first weight parameter is 0.9 and the second weight parameter is 0.1. A weighted average is performed to obtain a new embedded code ei` = 0.1 * er + 0.9 * ei, and then ei` is used to replace ei stored in the embedded code storage service.
[0149] Step 705: When the reliability parameter corresponding to the first voiceprint information is 0, update the text frequency information list of the first target object, and update the display times of the first text recognition result in the text frequency information list.
[0150] Taking the wake-up process of an in-vehicle system in an in-vehicle usage environment as an example below, the voiceprint information processing method provided by the present application will be described. Figure 8 It is a schematic diagram of the usage scenario of the voiceprint information processing method provided by the embodiments of the present invention. The voiceprint information processing method provided by the present invention can serve various types of customers in the form of cloud services (for example: encapsulated in an in-vehicle terminal or encapsulated in different mobile electronic devices). Among them, the user interface includes a perspective view of observing the task information processing environment in the instant client from the first-person perspective of different types of users. The user interface also includes a task control component and an information display component; through the user interface, the task matching the wake-up voice feature and the corresponding wake-up word are displayed by using the information display component; based on the result of the wake-up decision, the task processing result matching the wake-up voice feature is displayed by using the information display component through the user interface, so as to realize the information interaction between the electronic device and the user. For example, the user can use the wake-up word through a voice command to trigger the in-vehicle system to execute the music playback function or wake up the map applet in the in-vehicle WeChat for use.
[0151] Specifically, referring to Figure 9 , Figure 9 It is an optional flowchart of the voiceprint information processing method provided by the embodiments of the present invention, which specifically includes:
[0152] Step 901: Collect the audio information of the user through the in-vehicle terminal, and parse the audio information to obtain the voiceprint information of the user.
[0153] Step 902: Process the voiceprint information through the voiceprint recognition service process to obtain the embedded code corresponding to the voiceprint information, and find the set of voiceprint embedded codes matching the first terminal identification code according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process.
[0154] Step 903: Calculate the similarity between the voiceprint embedded code and each voiceprint embedded code in the set of voiceprint embedded codes.
[0155] Step 904: When the similarity is greater than or equal to the similarity threshold, calculate the reliability parameter corresponding to the voiceprint information, and update the voiceprint embedding codes in the voiceprint embedding code set.
[0156] Step 905: Obtain the corresponding voice command judgment result through the updated voiceprint embedding code to determine whether to wake up the in-vehicle terminal.
[0157] Beneficial technical effects:
[0158] In the embodiment of the present invention, the first audio information of the first target object is collected by the first terminal; the first audio information is analyzed to obtain the first voiceprint information of the first target object; the first voiceprint information is processed by the voiceprint recognition service process to obtain the first voiceprint embedding code corresponding to the first voiceprint information; according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process, a voiceprint embedding code set matching the first terminal identification code is searched; the similarity between the first voiceprint embedding code and each voiceprint embedding code in the voiceprint embedding code set is calculated; when the similarity is greater than or equal to the similarity threshold, the reliability parameter corresponding to the first voiceprint information is calculated; according to the reliability parameter corresponding to the first voiceprint information, the voiceprint embedding codes in the voiceprint embedding code set are updated. Thus, by updating the voiceprint embedding codes in the voiceprint embedding code set, the accuracy of the voiceprint embedding codes can be ensured, the accuracy of voiceprint recognition using the voiceprint embedding codes can be improved, and at the same time, the cumbersome steps of manual updating of the voiceprint embedding codes by the user can be reduced, enabling the user to obtain a better user experience. At the same time, different versions of the voiceprint recognition service process can be flexibly selected according to different usage environments, enhancing the user experience.
[0159] As mentioned above, the above are only embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A voiceprint information processing method, characterized in that, The method includes: Collecting first audio information of a first target object through a first terminal; Parsing the first audio information to obtain first voiceprint information of the first target object; Processing the first voiceprint information through a voiceprint recognition service process to obtain a first voiceprint embedding code corresponding to the first voiceprint information; Searching for a set of voiceprint embedding codes matching the first terminal identification code according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process; Calculating the similarity between the first voiceprint embedding code and each voiceprint embedding code in the set of voiceprint embedding codes; Searching for a text frequency information list of the first target object according to the identification information of the first target object; When the similarity is greater than or equal to a similarity threshold, and the display times of the first text recognition result corresponding to the first audio information in the text frequency information list is greater than or equal to 1, using the ratio of the display times to the total frequency in the text frequency information list as a reliability parameter corresponding to the first voiceprint information; Updating the voiceprint embedding codes in the set of voiceprint embedding codes according to the reliability parameter corresponding to the first voiceprint information.
2. The method according to claim 1, wherein The method further includes: Configuring the first terminal identification code for the first terminal; When collecting audio information of a second target object through the first terminal, configuring a second target object identification code for the second target object through the voiceprint recognition service process, and establishing a mapping relationship between the second target object identification code and the first terminal identification code; Obtaining at least two pieces of audio information of the second target object through the voiceprint recognition service process; Calculating second voiceprint embedding codes respectively corresponding to the at least two pieces of audio information through the voiceprint recognition service process; Calculating the average value of the second voiceprint embedding codes respectively corresponding to the at least two pieces of audio information; Saving the average value of the second voiceprint embedding codes in the set of voiceprint embedding codes, and marking it with the first terminal identification code, the second target object identification code, and the version identification code of the voiceprint recognition service process.
3. The method according to claim 1, wherein The step of processing the first voiceprint information through the voiceprint recognition service process to obtain a first voiceprint embedding code corresponding to the first voiceprint information includes: Triggering a voice information recognition model through the voiceprint recognition service process; Extracting the pinyin corresponding to each character in the first voiceprint information and the intonation corresponding to each character in the first voiceprint information through the voice information recognition model according to the recognition environment of the target voice information; Determining a single-character pronunciation feature vector at the character level in the first voiceprint information according to the pinyin corresponding to each character in the first voiceprint information and the intonation corresponding to each character in the first voiceprint information; Performing a combination process on the single-character pronunciation feature vectors corresponding to each character in the first voiceprint information through a character-to-phoneme encoder network in the voice information recognition model to form a sentence-level pronunciation feature vector; Using the sentence-level pronunciation feature vector as the first voiceprint embedding code.
4. The method according to claim 1, characterized in that The method further includes: Perform channel conversion processing on the first audio information to form mono audio data; Based on a window function corresponding to a speech information recognition model, perform short-time Fourier transform on the mono audio data to form a corresponding Mel spectrogram; Based on the Mel spectrogram, determine a corresponding input triple sample, and input the input triple sample into the speech information recognition model; Process the input triple sample crosswise through the convolutional layer and the max pooling layer of the speech information recognition model to obtain downsampling results of different input triple samples; Through the fully connected layer of the speech information recognition model, perform normalization processing on the downsampling results of different input triple samples to obtain the first voiceprint embedding code.
5. The method according to claim 1, characterized in that The method further includes: When the reliability parameter corresponding to the first voiceprint information is 0, update the text frequency information list of the first target object, and update the display times of the first text recognition result in the text frequency information list.
6. The method according to claim 1, wherein The updating of the voiceprint embedding codes in the voiceprint embedding code set according to the reliability parameter corresponding to the first voiceprint information includes: When the reliability parameter corresponding to the first voiceprint information is greater than or equal to the reliability parameter threshold, and the number of characters of the first text recognition result corresponding to the first audio information is greater than or equal to the character number threshold, obtain the original voiceprint embedding code corresponding to the similarity; Obtain the first weight parameter of the original voiceprint embedding code and the second weight parameter of the first voiceprint embedding code; According to the first weight parameter and the second weight parameter, calculate the weighted average of the original voiceprint embedding code and the first voiceprint embedding code to obtain a third voiceprint embedding code; Update the original voiceprint embedding code in the voiceprint embedding code set through the third voiceprint embedding code.
7. The method according to claim 1, wherein The method further includes: Save the voiceprint embedding code set in a cloud server; When collecting the first audio information of the first target object through a second terminal, detect the processing permission of the second terminal; When the second terminal meets the requirements of the processing permission, save the voiceprint embedding code set in the cloud server to the second terminal.
8. A voiceprint information processing device, characterized in that, The device includes: An information transmission module for collecting the first audio information of the first target object through a first terminal; An information processing module for parsing the first audio information to obtain the first voiceprint information of the first target object; The information processing module for processing the first voiceprint information through a voiceprint recognition service process to obtain a first voiceprint embedding code corresponding to the first voiceprint information; The information processing module for searching for a voiceprint embedding code set matching the first terminal identification code according to the first terminal identification code of the first terminal and the version identification code of the voiceprint recognition service process; The information processing module for calculating the similarity between the first voiceprint embedding code and each voiceprint embedding code in the voiceprint embedding code set; The information processing module is configured to find the text frequency information list of the first target object according to the identification information of the first target object; when the similarity is greater than or equal to the similarity threshold, and the display times of the first text recognition result corresponding to the first audio information in the text frequency information list is greater than or equal to 1, use the ratio of the display times to the total frequency in the text frequency information list as the reliability parameter corresponding to the first voiceprint information; The information processing module is configured to update the voiceprint embedding codes in the voiceprint embedding code set according to the reliability parameter corresponding to the first voiceprint information.
9. An electronic device, characterized in that, The electronic device includes: a memory for storing executable instructions; a processor, when running the executable instructions stored in the memory, implements the voiceprint information processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing executable instructions, characterized in that, When the executable instructions are executed by the processor, the voiceprint information processing method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, the voiceprint information processing method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Personalized voice control system
CN111968645A
Multimedia information processing method and device, electronic equipment and storage medium
CN112104892A
Voiceprint data processing method and device, electronic equipment and storage medium
CN112328994A
Voice information recognition method and device, electronic equipment and storage medium
CN113555006A