Speech recognition method, electronic device, and computer-readable storage medium
By recognizing and displaying spoken content and distinguishing speakers while recording audio, the system solves the problem of users needing to listen to audio sentence by sentence to ensure the integrity and accuracy of the content, thus improving the efficiency and accuracy of recording spoken content.
Patent Information
- Application Number
- CN202311376389.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-20
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2043-10-20
AI Technical Summary
When recording audio to capture spoken content, users need to listen to the audio sentence by sentence repeatedly to ensure the completeness and accuracy of the content, which results in excessive time consumption.
Electronic devices can identify and display spoken content while recording audio, and can distinguish speakers, identify registered speakers through voiceprint features or identify unregistered speakers through clustering, and provide speech recognition results.
It reduces the time users spend manually recording and correcting, and improves the efficiency and accuracy of recording spoken content.
Smart Images

Figure CN119905096B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the terminal field, and in particular to a speech recognition method, an electronic device and a computer readable storage medium. BACKGROUND
[0002] In some speech scenarios, a user usually needs to record the speech content of one or more speakers in a visualized form such as text on an electronic device. The user sometimes records audio by means of the electronic device and stores an audio file, so that the user can open the audio file by means of the electronic device to listen to the audio and then correct the recorded speech content based on the audio. In order to ensure the integrity and accuracy of the recorded speech content, the user may need to listen to the audio repeatedly and sentence by sentence, which makes the user spend a lot of time in recording the speech content on the electronic device. SUMMARY
[0003] The present application provides a speech recognition method, an electronic device and a computer readable storage medium. The electronic device can recognize and display the speech content in the audio while recording the audio. In addition, the electronic device can also distinguish the speaker of each speech content, so that the user can obtain the text record of the speech content and the speaker identification in the audio by means of the electronic device without manual recording.
[0004] In a first aspect, the present application provides a speech recognition method, which is applied to an electronic device, the electronic device storing a first identifier and a first voiceprint feature of a first speaker, and the method comprising: the electronic device receiving a first operation of a user selecting a speech recognition mode, the speech recognition mode comprising a first recognition mode and a second recognition mode; the electronic device obtaining a first audio segment, the first audio segment containing a first sub-segment corresponding to a first speech content; in the case that the first operation selects the first recognition mode, the electronic device determining that the voiceprint feature of the first sub-segment matches the first voiceprint feature, and the electronic device displaying the first speech content and the first identifier, the first identifier indicating that the first speech content is spoken by the first speaker; in the case that the first operation selects the second recognition mode, the electronic device determining that the first speech content is spoken by one speaker, and the electronic device displaying the first speech content and a second identifier, the second identifier being used to identify the speaker of the first speech content.
[0005] The electronic device can recognize the speaking content in the audio and display the speaking content. The electronic device can also recognize the speaker based on a voice recognition method selected by the user. The electronic device can recognize a registered speaker in the audio based on a voiceprint feature, or recognize an unregistered speaker based on a clustering method. Registering a speaker means that the electronic device stores the identity and voiceprint feature of the speaker. When the user selects the voiceprint recognition method, the electronic device can recognize the speaker in the audio based on the registered speaker. When the user selects the unregistered speaker recognition method, the electronic device can cluster the audio segments to distinguish the speakers in the audio segments. In this way, the electronic device can provide the user with voice recognition results containing speaking content and corresponding speakers, regardless of whether the user has registered the speaker in advance.
[0006] In combination with the first aspect, in some embodiments, after the electronic device obtains the first audio segment, the method further includes: the electronic device divides the first audio segment into a plurality of sub-segments, the plurality of sub-segments including the first sub-segment, and the plurality of sub-segments having different semantics.
[0007] The first audio segment can be a valid audio segment. The electronic device can divide the valid audio segment based on semantics to obtain a plurality of sub-segments, each sub-segment having different semantics of speaking content. In this way, the electronic device can improve the probability that each sub-segment contains only one speaker, thereby improving the accuracy of subsequent recognition of the speaker in the sub-segment.
[0008] In combination with the first aspect, in some embodiments, the first audio segment includes a second sub-segment, the second sub-segment including second speaking content of a second speaker, and the electronic device does not store a second voiceprint feature of the second speaker. The method further includes: in a case where the first operation selects a first recognition method, the electronic device determines that the second speaking content is spoken by one speaker, and the electronic device displays the second speaking content and a third identity, the third identity identifying the speaker of the second speaking content.
[0009] The first recognition method can be a voiceprint recognition method. When the electronic device recognizes the speaker by voiceprint, it can occur that the voiceprint feature in some sub-segments cannot be matched with the voiceprint feature stored by the electronic device. This can be because the electronic device does not store the voiceprint feature of the speaker in the sub-segment. The electronic device can use a clustering recognition method to determine to cluster the sub-segment, thereby distinguishing the speaker of the speaking content. The third identity can be different from the type of the first identity. For example, the third identity can be "Speaker A", "Speaker B", etc., and the first identity can be "Zhang San", "Li Si", etc.
[0010] In some embodiments, the first audio segment includes a second sub-segment, the second sub-segment includes second speech content of a second speaker, the electronic device stores a second voiceprint feature of the second speaker and a fourth identifier, and the method further includes: in a case where the first operation selects a first recognition manner, the electronic device determines that a voiceprint feature of the second sub-segment matches the second voiceprint feature, and the electronic device displays the second speech content and the fourth identifier, the fourth identifier indicating that the second speech content is spoken by the second speaker.
[0011] The first recognition manner can be a voiceprint recognition manner. The electronic device can extract a voiceprint feature from the second sub-segment, and then match the voiceprint feature with the stored voiceprint feature. Finally, the electronic device can determine that the voiceprint feature of the second sub-segment matches the second voiceprint feature, and then display the fourth identifier corresponding to the second voiceprint feature. That is, the electronic device can label the speaker of the speech content based on the stored voiceprint feature and the identifier corresponding thereto, wherein the speaker identifier of each voiceprint feature stored by the electronic device can be input by the user, so that the user can clearly understand which speaker spoke each speech content through the speaker identifier.
[0012] In some embodiments, after the electronic device displays the second speech content and the fourth identifier, the method further includes: the electronic device receives a second operation for filtering the first speaker, and in response to the second operation, the electronic device hides the display of the second speech content and the fourth identifier.
[0013] That is, the electronic device can respond to the user operation to filter the speech content of one or more speakers, and the electronic device can hide the display of the speech content of the unfiltered speaker and only display the speech content of the filtered speaker. In this way, the user can filter the speech content of the speaker of interest from the speech content of many speakers for viewing, thereby improving the user experience.
[0014] In some embodiments, the first audio segment further includes a third sub-segment, the third sub-segment includes third speech content, the speaker of the third speech content is different from the speaker of the first speech content, and the method further includes: in a case where the first operation selects a second recognition manner, the electronic device determines that the third speech content is spoken by a speaker, and the electronic device displays the third speech content and a fifth identifier, the fifth identifier indicating that the speaker of the third speech content is different from the speaker of the first speech content.
[0015] The second identification manner can be a clustering identification manner. The electronic device can determine, by using the clustering identification module, that the speaker in the third speech content is different from the speaker of the first speech content, and then the electronic device can assign a speaker identifier different from the second identifier to the third speech content, so as to distinguish the speaker of the first speech content. That is, when the speech recognition method is a clustering identification method, the electronic device can assign different identifiers to different speakers, which is convenient for the user to understand.
[0016] In combination with the first aspect, in some embodiments, the first audio segment further includes a fourth sub-segment, and the fourth sub-segment includes fourth speech content. The method further includes: in a case where the first operation selects the second identification manner, the electronic device determines that the fourth speech content is spoken by one speaker, and the speaker of the fourth speech content is the same as the speaker of the first speech content. The electronic device displays the fourth speech content, and displays the second identifier again at the first position associated with the fourth speech content. The second identifier displayed at the first position is used to indicate the speaker of the fourth speech content.
[0017] The electronic device can determine, by using the clustering identification module, that the speaker in the fourth speech content is the same as the speaker of the first speech content, and then the electronic device can assign the second identifier to the fourth speech content, which is used to indicate that the speakers of the first speech content and the fourth speech content are the same. That is, when the speech recognition method is a clustering identification method, the electronic device can assign the same identifier to the same speaker, which is convenient for the user to understand.
[0018] In combination with the first aspect, in some embodiments, before the electronic device receives the first operation of the user selecting the speech recognition manner, the method further includes: the electronic device receives a second audio segment and the first identifier input by the user, and the speaker of the second audio segment is the first speaker; the electronic device extracts a first voiceprint feature of the first speaker from the second audio segment; and the electronic device stores the first voiceprint feature and the first identifier, where the first voiceprint feature and the first identifier are stored in association.
[0019] The electronic device can receive an audio segment and a speaker identifier input by the user to register a speaker. The audio segment input by the user can be an audio segment containing only one speaker. Then the electronic device can extract a voiceprint feature of the speaker from the audio segment and store the voiceprint feature in association with the speaker. In this way, the electronic device can identify the speaker of the audio segment to be recognized based on the stored voiceprint feature and the speaker identifier provided by the user during the speech recognition process, and the electronic device can use the pre-stored speaker identifier provided by the user to label the speech content of the audio segment to be recognized, so that the user can accurately and clearly understand the speaker of the speech content.
[0020] In combination with the first aspect, in some embodiments, the method further includes: the electronic device collecting a second audio segment through the microphone during display of the first speech content; and the electronic device identifying the speech content and the speaker in the second audio segment based on the voice recognition manner selected by the first operation.
[0021] That is, the electronic device can perform real-time voice recognition while collecting audio. The electronic device can continue to collect audio while displaying the speech content. Then the electronic device can continue to identify the speech content and the speaker in the newly collected audio. In this way, the user can obtain real-time recording of the speech content and the corresponding speaker in the speech scenario through the electronic device, thereby improving the user experience.
[0022] In a second aspect, the present application provides an electronic device, which includes a storage module configured to store a first identifier and a first voiceprint feature of a first speaker; a user interaction module configured to receive a first operation of a user selecting a voice recognition manner, the voice recognition manner including a first recognition manner and a second recognition manner; an audio collection module configured to obtain a first audio segment, the first audio segment including a first sub-segment corresponding to first speech content; a voiceprint recognition module configured to determine, in a case where the first operation selects the first recognition manner, that a voiceprint feature of the first sub-segment matches the first voiceprint feature; a display module configured to display, in the case where the first operation selects the first recognition manner, the first speech content and the first identifier, the first identifier indicating that the first speech content is spoken by the first speaker; and a clustering recognition module configured to determine, in a case where the first operation selects the second recognition manner, that the first speech content is spoken by one speaker based on an acoustic feature of the first audio segment; and the display module is further configured to display, in the case where the first operation selects the second recognition manner, the first speech content and a second identifier, the second identifier being used to identify the speaker of the first speech content.
[0023] In combination with the second aspect, in some embodiments, the electronic device further includes a voice recognition module and a voice segmentation module; the content recognition module is configured to identify speech content included in the first audio segment; and the voice segmentation module is configured to divide the first audio segment into a plurality of sub-segments based on the speech content included in the first audio segment, the plurality of sub-segments including the first sub-segment, and the plurality of sub-segments having different semantics.
[0024] In combination with the second aspect, in some embodiments, the first audio segment includes a second sub-segment, the second sub-segment including second speech content of a second speaker, and the storage module does not store a second voiceprint feature of the second speaker; in the case where the first operation selects the second recognition manner, the clustering recognition module is further configured to determine that the second speech content is spoken by one speaker; and the display module is further configured to display the second speech content and a third identifier, the third identifier being used to identify the speaker of the second speech content.
[0025] With reference to the second aspect, in some embodiments, the first audio segment further includes a second sub-segment, the second sub-segment includes second speech content of a second speaker, the second voiceprint feature of the second speaker and the fourth identifier are stored in the storage module; in a case where the first operation selects the first recognition mode, the voiceprint recognition module is further configured to determine that the voiceprint feature of the second sub-segment matches the second voiceprint feature; and the display module is further configured to display the second speech content and the fourth identifier, the fourth identifier indicating that the second speech content is spoken by the second speaker.
[0026] With reference to the second aspect, in some embodiments, the first audio segment further includes a third sub-segment, the third sub-segment includes third speech content, and a speaker of the third speech content is different from a speaker of the first speech content; in a case where the first operation selects the second recognition mode, the clustering recognition module is further configured to determine that the third speech content is spoken by one speaker; and the display module is further configured to display the third speech content and a fifth identifier, the fifth identifier indicating that the speaker of the third speech content is different from the speaker of the first speech content.
[0027] With reference to the second aspect, in some embodiments, the first audio segment further includes a fourth sub-segment, the fourth sub-segment includes fourth speech content; in a case where the first operation selects the second recognition mode, the clustering recognition module is further configured to determine that the fourth speech content is spoken by one speaker, and the speaker of the fourth speech content is the same as the speaker of the first speech content; and the display module is further configured to display the fourth speech content and display the second identifier again at a first position associated with the fourth speech content, the second identifier displayed at the first position indicating the speaker of the fourth speech content.
[0028] With reference to the second aspect, in some embodiments, the electronic device further includes a voiceprint registration module, the voiceprint recognition module is configured to receive a second audio segment and a first identifier input by a user, a speaker of the second audio segment being a first speaker; the voiceprint registration module is further configured to send the second audio segment to the voiceprint recognition module; the voiceprint recognition module is further configured to extract a first voiceprint feature of the first speaker from the second audio segment, and send the first voiceprint feature to the voiceprint registration module; the voiceprint registration module is further configured to receive the first voiceprint feature, and send the first voiceprint feature and the first identifier to the storage module; and the storage module is further configured to store the first voiceprint feature and the first identifier, wherein the first voiceprint feature and the first identifier are stored in association.
[0029] In a third aspect, the present application provides an electronic device, which includes a display screen, a storage, and a processor coupled to the storage; the display screen is configured to display an interface, the storage stores a computer program, and the processor executes the computer program to enable the electronic device to implement the method of any one of the first aspect.
[0030] In a fourth aspect, the present application provides a computer readable storage medium storing a computer program or computer instructions, and the computer program or computer instructions are executed by a processor to implement the method of any one of the first aspect.
[0031] In a fifth aspect, the present application provides a computer program product, and when the computer program product is executed by a processor, the method of any one of the first aspect will be implemented.
[0032] In a sixth aspect, the present application provides a chip, which includes a processor and a memory, wherein the memory is used to store a computer program or computer instructions, and the processor is used to execute the computer program or computer instructions stored in the memory, so that the chip executes the method of any one of the first aspect.
[0033] The solutions provided in the second aspect to the sixth aspect are used to implement or cooperate to implement the method provided in the first aspect, and thus can achieve the same or corresponding beneficial effects as the corresponding method in the first aspect, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0035] Figure 2 FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0036] Figure 3A FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0037] Figure 3B FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0038] Figures 4A-4W FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application;
[0039] Figure 5 FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application; DETAILED DESCRIPTION
[0040] The terminology used in the following description of the embodiments herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used in the description of the embodiments and the appended claims herein, the singular forms "a", "an" and "the" are intended to include both singular and plural forms, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0041] Hereinafter, the terms "first", "second", etc. are used only for the purpose of description, and should not be understood as implying or indicating relative importance or implying the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise specified.
[0042] In some speaking scenarios (such as conference scenarios), a user often needs to record the speaking content of one or more speakers in the form of visualization such as text. The speaking content can include the words, phrases, sentences spoken by the speaker, and the tone of the speaker, etc. During the speaking of the speaker, the user usually records the speaking content in the form of text into an electronic device in real time. Sometimes, in order to ensure the completeness and accuracy of the speaking content recorded by the user, the user will record the audio during the speaking of the speaker, and then store the recorded audio in the electronic device. In this way, the user can open the audio file to listen to the audio at any time, and modify or supplement the speaking content recorded by the user. In this process, the user may need to listen to the audio repeatedly and sentence by sentence, which makes the user spend a lot of time in recording the speaking content on the electronic device.
[0043] In order to save the time spent by the user in recording the speaking content, the embodiments of the present application provide a speech recognition method, an electronic device and a computer readable storage medium. In the speech recognition method, the electronic device can recognize the speaking content contained in the audio while recording the audio, and then the electronic device can display the speaking content contained in the audio. Moreover, the electronic device can distinguish the speakers in the audio, so that the user knows which speaker speaks each sentence, each word, etc. speaking content in the audio. That is to say, the electronic device can generate and display the speaking record containing the speaking content and the speaker identification in real time when the speaker speaks, so that the user can modify the speaking record in time, thereby improving the work efficiency of the user.
[0044] Next, an exemplary electronic device 100 provided by the embodiments of the present application is introduced.
[0045] Figure 1FIG. 1 is a structural schematic diagram of an electronic device 100 provided by an embodiment of the present application.
[0046] The embodiments will be described below with the electronic device 100 as an example. It should be understood that the electronic device 100 can have more or fewer components than those shown in FIG. 1, can combine two or more components, or can have a different configuration of components. Figure 1 The various components shown in FIG. 1 can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits. Figure 1 The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a display screen 190, and the like. The sensor module 180 can include a pressure sensor 180A, a touch sensor 180B, and the like.
[0047] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), and the like. Different processing units can be independent devices or integrated into one or more processors.
[0048] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals to complete the control of fetching and executing instructions.
[0049]
[0050] The processor 110 can also have a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has just used or is using repeatedly. If the processor 110 needs to use the instructions or data again, it can call them directly from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thus improving the efficiency of the system.
[0051] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc. The components in the electronic device 100 can be connected through the above interfaces.
[0052] The charging management module 140 is configured to receive a charging input from a charger. The charger can be a wireless charger or a wired charger.
[0053] The power management module 141 is configured to connect the battery 142 and the charging management module 140. The power management module 141 receives the input of the battery 142 and / or the charging management module 140, and supplies power to the processor 110, the internal memory 121, the external memory, the display 190, etc.
[0054] The modem processor can include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to a loudspeaker 170A, a receiver 170B, etc.), or displays an image or a video through the display screen 190. In some embodiments, the modem processor can be a separate device.
[0055] The electronic device 100 implements a display function through a GPU, the display screen 190, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 190 and the application processor. The GPU is configured to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.
[0056] The display screen 190 is configured to display an image, a video, etc. The display screen 190 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include one or N display screens 190, where N is a positive integer greater than 1.
[0057] The digital signal processor is configured to process digital signals, in addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is configured to perform Fourier transform on the frequency point energy, etc.
[0058] The NPU is a neural-network (NN) computing processor, which is configured to process input information quickly by referring to the structure of a biological neural network, for example, by referring to the transmission mode between human brain neurons, and can also constantly self-learn.
[0059] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to extend the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, files such as music and videos are stored in the external memory card.
[0060] The internal memory 121 can be used to store computer executable program codes including instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required by a function (such as a face recognition function, a fingerprint recognition function, a mobile payment function, etc.), and the like. The data storage area can store data created during use of the electronic device 100 (such as face information template data, fingerprint information template, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.
[0061] The electronic device 100 can implement an audio function through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, the application processor, and the like. For example, music playing, recording, and the like.
[0062] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode an audio signal. In some embodiments, the audio module 170 can be disposed in the processor 110, or part of the functions of the audio module 170 can be disposed in the processor 110.
[0063] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into an acoustic signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0064] The receiver 170B, also known as a "earpiece", is used to convert an audio electrical signal into an acoustic signal. When the electronic device 100 answers a call or a voice message, the receiver 170B can be held close to the ear to listen to the voice.
[0065] Microphone 170C, also called "microphone", "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can make a sound by putting his mouth close to the microphone 170C, inputting the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, in addition to collecting sound signals, it can also realize the function of noise reduction. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C, which can realize the functions of collecting sound signals, noise reduction, identifying sound sources, realizing directional recording, etc.
[0066] The earphone interface 170D is used to connect the wired earphone. The earphone interface 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0067] The pressure sensor 180A is used to sense the pressure signal, which can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be provided on the display screen 190. There are many types of pressure sensors 180A, such as resistance pressure sensors, inductance pressure sensors, and capacitance pressure sensors. The capacitance pressure sensor can include at least two parallel plates made of conductive material. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure according to the change of the capacitance. When a touch operation acts on the display screen 190, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A.
[0068] The touch sensor 180B, also called "touch panel". The touch sensor 180B can be provided on the display screen 190, and the touch sensor 180B and the display screen 190 form a touch screen, also called "touch screen". The touch sensor 180B is used to detect the touch operation acting on or near it. The touch sensor can pass the detected touch operation to the application processor to determine the touch event type. The visual output related to the touch operation can be provided through the display screen 190. In other embodiments, the touch sensor 180B can also be provided on the surface of the electronic device 100, which is different from the position of the display screen 190.
[0069] In some embodiments, the electronic device can receive audio containing speech content through the microphone 170C, the audio module 170 can convert the audio data in the form of an analog signal output by the microphone 170C into a digital signal, and the DSP can perform resampling, denoising, etc. on the audio in the form of a digital signal output by the audio module 170, and then send the audio data to the NPU. Further, the NPU can recognize the speech content in the audio, that is, convert the speech content contained in the audio into visual (such as text, pattern, etc.) speech content.
[0070] In some embodiments, after the NPU converts the speech content in the audio into text form, it can also segment the audio data based on the text. For example, the NPU can divide the audio into multiple segments based on the semantics of the speech content, and each segment of audio can correspond to a complete phrase or sentence, and then the NPU can identify the speaker corresponding to each segment of audio.
[0071] In some embodiments, the NPU can identify the speaker corresponding to each segment of audio based on a clustering algorithm, and in other embodiments, the NPU can identify the speaker corresponding to each segment of audio based on a voiceprint. The voiceprint is a biometric feature composed of multiple dimensional features such as wavelength, frequency, intensity, etc., which is used to uniquely identify a speaker. The voiceprint can contain a mathematical vector obtained by embedding multiple dimensional features. Optionally, the electronic device can store the voiceprint of one or more speakers in the internal memory 121, for example, the electronic device can store the voiceprint in the read-only memory (ROM).
[0072] The electronic device 100 can also register a speaker, which means that the electronic device stores the voiceprint of the speaker to be registered in the storage or a cloud server based on the voiceprint extracted from the audio data of the speaker to be registered. The electronic device can receive the voice audio of the speaker to be registered through the microphone 170C, and after the audio data is processed by the audio module 170 and the DSP, it can be sent to the NPU, and the NPU can extract the voiceprint of the speaker to be registered from the audio data. Alternatively, the electronic device can also send the audio data of the speaker to be registered stored in the external memory or the cloud to the NPU, and then the NPU can extract the voiceprint of the speaker to be registered from the audio data. Optionally, the NPU can extract the voiceprint of the speaker to be registered based on the D-Vector or X-Vector method, and the embodiments of the present application do not limit the method of the NPU extracting the voiceprint of the speaker to be registered from the audio data.
[0073] Figure 2 is a software structure block diagram of an electronic device 100 provided by an embodiment of the present application.
[0074] The layered architecture divides software into several layers, each of which has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the system of the electronic device 100 can include five layers, namely an application layer, an application framework layer, a runtime and system library, a kernel layer, and an artificial intelligence (AI) framework engine.
[0075] The application layer can include a series of application packages. As shown in Figure 2 , the application packages can include note, recorder, and other applications (also referred to as applications).
[0076] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications of the application layer. The application framework layer includes some pre-defined functions. For example, the application framework layer can include an audio system and a view system.
[0077] The audio system is responsible for the comprehensive management of the transmission and control of audio streams. Among them, the audio system can provide an interface for the upper layer application to call the audio control. Among them, the audio stream is continuous audio data transmitted using streaming technology. Streaming technology can be real-time streaming or progressive streaming.
[0078] The view system includes visual controls, such as controls that display text, controls that display pictures, and the like. The view system can be used to build applications. The display interface can be composed of one or more views. For example, the view can include a view that displays the content of a speech in a note application.
[0079] Without limitation, the application framework layer can also include a content provider (not shown in Figure 2 ) and a window manager (not shown in Figure 2 ). Among them, the content provider is used to store and obtain data, and makes these data accessible to applications. The above-mentioned data can include videos, images, audio, dialed and answered calls, notes, phone books, and the like. The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and the like.
[0080] The runtime includes core libraries and virtual machines. The runtime is responsible for the scheduling and management of the system.
[0081] The core library includes two parts: one part is a function function that the programming language (for example, the java language) needs to call, and the other part is the core library of the system.
[0082] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the programming files (for example, the java files) of the application layer and the application framework layer into binary files. The virtual machine is used to perform functions such as management of the object life cycle, stack management, thread management, management of security and exceptions, and garbage collection.
[0083] The system library can include a plurality of function modules. For example: an audio stream processing library, a three-dimensional graphics processing library (for example: OpenGL ES), a two-dimensional graphics engine (for example: SGL), a surface manager, media libraries, and the like.
[0084] The audio stream processing library is used to implement the transmission and control of the audio stream. For example, after the note application receives a user operation for starting the audio recording function, the interface of the audio system can be called, the audio stream processing library can be called to transmit the request for starting the audio recording to the corresponding module, and the function of recording the audio is realized.
[0085] The three-dimensional graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing.
[0086] The two-dimensional graphics engine is a drawing engine for 2D drawing.
[0087] The surface manager is used to manage the display subsystem, and provides the fusion of two-dimensional (2-Dimensional, 2D) and three-dimensional (3-Dimensional, 3D) layers for a plurality of applications.
[0088] The media library supports a plurality of commonly used audio, video format playback and recording, and static image files. The media library can support a plurality of audio and video coding formats, for example: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, and the like.
[0089] The kernel layer is the layer between the hardware and the software. The kernel layer can include a display driver and an audio driver. The display driver can be used to drive the display screen 190 to display the picture, and the audio driver can be used to drive the microphone 170C to record the voice audio.
[0090] The AI framework engine can be used to process artificial intelligence tasks. Among them, the AI framework engine can include a content recognition module, a speech segmentation module, a voiceprint registration module, a voiceprint recognition module, and a clustering recognition module. The application program can call the software modules in the AI framework engine to realize the corresponding functions. The introduction of each module in the AI framework engine can refer to the introduction of the subsequent embodiments, which will not be expanded here.
[0091] Figure 2 Each software module in the software structure block diagram shown is only an example and is not a limitation of the embodiments of the present application. The software structure of the electronic device 100 can have more or fewer software modules than those shown, two or more software modules can be combined, and different software module configurations can be used. Figure 2 The software structure shown can have more or fewer software modules than those shown, two or more software modules can be combined, and different software module configurations can be used.
[0092] For example, in some embodiments, the electronic device 100 can further include a storage module, a user interaction module, an audio acquisition module, and a display module. The storage module is used to store the speaker identification and its corresponding voiceprint features. The user interaction module is used to receive user operations and determine the operation content of the user. The audio acquisition module is used to acquire the audio stream through the microphone. The display module is used to display the user interface.
[0093] The structure of an AI framework engine provided by an embodiment of the present application is introduced below.
[0094] Figure 3A is a structural block diagram of an AI framework engine of an electronic device 100 provided by an embodiment of the present application.
[0095] As Figure 3A shown, the AI framework engine can include a content recognition module, a speech segmentation module, a voiceprint registration module, a voiceprint recognition module, and a clustering recognition module.
[0096] The electronic device 100 first acquires an audio stream and segments the audio stream to obtain an effective audio segment. Among them, the effective audio refers to the audio that may contain human voice. The audio stream is continuous audio data transmitted using streaming technology. The electronic device 100 can filter out the effective audio in the audio stream by calculating the short-time energy value or short-time zero-crossing rate of the audio segment in the audio stream. For example, the electronic device 100 can calculate whether the short-time energy value in each window of the audio stream exceeds the threshold value by using a sliding window. When the short-time energy value in a certain window of the audio stream is lower than the preset threshold value, the electronic device can select the starting time point of the window to cut the audio stream. In this way, the electronic device 100 filters out an effective audio segment, and the short-time energy value in each window of the effective audio segment is greater than the preset threshold value. Not limited to this method, the electronic device 100 can also filter out the effective audio segment by other methods, which are not limited by the embodiments of the present application.
[0097] The content recognition module can be used to recognize the speaking content contained in the audio. The electronic device 100 can input the valid audio segments filtered from the audio stream into the content recognition module. One valid audio segment can refer to an audio segment with continuous voice. Here, the continuous voice can refer to the voice part in the audio being continuously present or appearing, the voice part being uninterrupted or the duration of interruption being no more than a preset duration, that is, the duration of the speaker's pause in the audio being no more than a preset duration. The electronic device 100 can determine whether the voice in the audio stream is continuous by using a short-time energy value or a short-time zero-crossing rate. Each valid audio segment is composed of one or more frames of audio, and the electronic device 100 can use streaming technology to input the data of one or more frames of audio in the audio segment into the content recognition module. The electronic device 100 can take audio every 20 milliseconds as one frame of audio, which is not limited to 20 milliseconds, and one frame of audio can also be audio of other lengths of time, for example, one frame of audio data can also be 10 milliseconds of audio, or 5 milliseconds of audio, and the like, which is not limited in the embodiments of the present application. The electronic device 100 can process each frame of audio by the content recognition module frame by frame to recognize the speech text content in each frame of audio.
[0098] Exemplarily, the content recognition module can include an encoder, an embedding layer, a language prediction model, and a joint neural network (Joint NN) model. The encoder can encode the audio data to obtain an acoustic encoding sequence (or referred to as acoustic encoding features). The embedding layer can receive the text that has been recognized from the current valid audio segment, and then encode the text to obtain a text encoding sequence (or referred to as text encoding features). Further, the language prediction model can predict subsequent text based on the text encoding sequence, and then generate a new text encoding sequence. The joint neural network model can splice the text encoding sequence and the acoustic encoding sequence to form a new encoding sequence, and then the joint neural network model can obtain a probability distribution matrix based on the spliced encoding sequence to output the text of the speaking content contained in the audio segment, wherein the probability distribution matrix can be a four-dimensional tensor, used to represent the probability of different words in the vocabulary corresponding to the audio input into the content recognition module this time.
[0099] Taking the content recognition module recognizing Chinese as an example, after the electronic device 100 cuts out an effective audio segment containing human voice, the content recognition module can be called to process the data of the effective audio segment. The first frame of audio data in the data of the effective audio segment can contain the pronunciation (wo) of the word “wo”, and the second frame of audio data can contain the pronunciation (“shi”) of the word “shi”. The content recognition module first uses the encoder and the joint neural network model to recognize the speaking content contained in the first frame of audio data in the effective audio segment, and outputs the speaking content text “wo” contained in the first frame of audio data. Then the content recognition module uses the encoder to obtain the acoustic encoding sequence of the second frame of audio data, and at the same time, the content recognition module can input the speaking content text “wo” recognized according to the first frame of audio data into the embedding layer. The embedding layer will generate a text encoding sequence based on the text “wo”, and then the language prediction model can predict the subsequent text based on the text encoding sequence to generate a new text encoding sequence, which can be used to strengthen the possible text. For example, “wo” can be followed by “shi”, and the language prediction model can strengthen “shi” through the new text encoding sequence. The joint neural network model can splice the text encoding sequence and the acoustic encoding sequence to obtain a probability distribution matrix. The pronunciation “shi” of the second frame of audio data can be recognized as “si”, “shi”, “shi”, etc. Since the language prediction model has strengthened “shi” based on “wo”, the joint neural network model can determine that the probability of the speaking content contained in the second frame of audio data being the word “shi” is higher, and thus the speaking content corresponding to the second frame of audio data is recognized as “shi”. In this way, the accuracy of the content recognition module in recognizing the speaking content text is improved.
[0100] The method for the content recognition module to recognize the content of other languages from the voice can refer to the above method for recognizing the content of Chinese from the voice, which will not be described here. Without being limited to the above method, the content recognition module can also use other methods to recognize the speaking content contained in the audio segment. The present embodiment does not limit the method for the content recognition module to recognize the speaking content contained in the audio segment and then output the speaking content in the form of text.
[0101] The electronic device can input the speaking content text of the valid audio segment into the speech segmentation module. Among them, the content recognition module can input some intermediate results into the speech segmentation module in the process of identifying the speaking content of each valid audio segment. The above-mentioned intermediate results refer to the speaking content text output by the content recognition module when it has not completely identified all the speaking content of the audio segment. For example, the speaking content of an audio segment includes: I plan to go to the seaside, and the intermediate results identified by the content recognition module can include "I", "I plan", "I plan to go", then the electronic device 100 can input the above-mentioned intermediate results into the speech segmentation module as soon as they are identified. And / or, the content recognition module can input all the speaking content text contained in each valid audio segment into the speech segmentation module after identifying all the speaking content of the audio segment, for example, the content recognition module inputs "I plan to go to the seaside" into the speech segmentation module after identifying all the speaking content "I plan to go to the seaside" of the audio segment. The speech segmentation module can receive the speaking content text of the valid audio segment and the valid audio segment, and then cut the valid audio segment based on the semantic features of the text and the audio features in the audio, and divide the valid audio segment into one or more sub-segments.
[0102] In some embodiments, the speech segmentation module can include a multi-modal speech segmentation model. Among them, the multi-modal speech segmentation model can receive the speaking content text processed by the electronic device 100 and identified by the content recognition module from the valid audio segment, the multi-modal speech segmentation model can obtain the audio features from the valid audio segment, and obtain the semantic features from the speaking content text, and then segment the valid audio segment based on the audio features and the semantic features, thereby outputting one or more audio sub-segments. Among them, the multi-modal speech segmentation model can cut the audio when the duration of the pause of the voice in the audio exceeds the preset duration, or when a semantically complete sentence ends. The multi-modal speech segmentation model can be a neural network model, and the multi-modal speech segmentation model can include one or more of the following: a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, a long short-term memory (LSTM) model, a deep neural network (DNN) model, a Transformer model, etc.
[0103] The electronic device 100 can further divide the one or more valid audio segments obtained by the preliminary division of the audio stream. For example, some of the valid audio segments can include speech of a single speaker that lasts for a long time. If the electronic device directly inputs an entire valid audio segment to a subsequent voiceprint recognition module, the recognition efficiency can be poor. Alternatively, some of the valid audio segments can include speech of multiple speakers that speak alternately. If the electronic device inputs a valid audio segment that includes speech of multiple speakers to a subsequent voiceprint recognition module or a clustering recognition module, the accuracy of speaker recognition can be affected. Therefore, the speech division module can further divide the audio segments to obtain one or more audio sub-segments.
[0104] For example, the speech content of a valid audio segment input to the speech division module by the electronic device can be "How about playing football tomorrow if it is sunny? OK, I think so." Here, "How about playing football tomorrow if it is sunny?" can be the speech content of a first speaker, and "OK, I think so" can be the speech content of a second speaker. The electronic device 100 can divide the valid audio segment into two audio sub-segments by the speech division module, and the speech content of the first audio sub-segment can be "How about playing football tomorrow if it is sunny?", and the speech content of the second audio sub-segment can be "OK, I think so." The speech division module can then input the data of the two audio sub-segments to the subsequent voiceprint recognition module or the clustering recognition module in sequence for processing. Since the electronic device 100 uses the speech division module to divide the audio sub-segments, the probability that each audio sub-segment includes only a single speaker can be increased, which can improve the efficiency and accuracy of speaker recognition by the subsequent voiceprint recognition module or the clustering recognition module.
[0105] The voiceprint recognition module can receive the audio sub-segments output by the speech division module, and then extract the voiceprints included in the audio sub-segments, and further recognize the speakers based on the voiceprints. The voiceprint is a biometric feature composed of multiple dimensions of characteristics of sound, such as wavelength, frequency, and intensity, and is used to uniquely identify a speaker. The voiceprint can include a mathematical vector obtained by vectorizing the multiple dimensions of characteristics. The voiceprint recognition module can include an audio segmentation module, a voiceprint feature extraction module, a speaker classification module, and an activation function.
[0106] The audio segmentation module can segment the audio sub-segment obtained by the speech segmentation module according to a preset time length to obtain a plurality of audio short segments with shorter time. The voiceprint feature extraction module can receive data of the plurality of segmented audio short segments, and then extract voiceprint features. The voiceprint feature extraction module can input data of each audio short segment into a DNN model, and then obtain a plurality of feature vectors. Further, the voiceprint feature extraction module can calculate the average of the plurality of feature vectors, and output the calculation result to the speaker classification module. The calculation result can also be referred to as the voiceprint feature of the audio sub-segment. The speaker classification module can match the voiceprint feature of the audio sub-segment with the voiceprint features stored by the electronic device, wherein each voiceprint feature can correspond to a speaker. Further, the speaker classification module can input the matching result of the voiceprint feature of the audio sub-segment and the voiceprint features of one or more speakers stored by the electronic device into an activation function (for example, which can be a softmax function). The activation function can normalize the above-mentioned matching result to obtain the probability of matching the voiceprint feature of the audio sub-segment with each voiceprint feature stored by the electronic device. Finally, the voiceprint recognition module can output the identification of the speaker corresponding to the voiceprint feature with the maximum probability. The voiceprint recognition method based on the voiceprint recognition speaker of the voiceprint recognition module can also be referred to as a D-Vector-based voiceprint recognition method, which is not limited to the D-Vector-based voiceprint recognition method. The voiceprint recognition module can also extract the voiceprint of the audio segment by other methods, and then identify the speaker. The embodiments of the present application do not limit the method of extracting the voiceprint of the audio segment by the voiceprint recognition module, and then identifying the speaker based on the voiceprint.
[0107] The voiceprint features stored by the electronic device can be obtained by the voiceprint registration module after registering the speaker. The voiceprint registration module can input the speaker identification and the to-be-identified audio segment spoken by the speaker into the voiceprint recognition module. After the audio segmentation module segments the to-be-identified audio segment into a plurality of audio short segments, the voiceprint feature extraction module inputs data of the plurality of audio short segments into the voiceprint feature extraction module. The voiceprint feature extraction module can extract the voiceprint of the entire audio segment based on the data of the plurality of audio short segments. Further, the voiceprint registration module can store the voiceprint extracted based on the to-be-identified audio segment and the speaker identification corresponding to the to-be-identified audio segment in the storage. In this way, the voiceprint recognition module can identify the speaker based on the voiceprint stored in the storage. Each speaker identification can correspond to the voiceprint of the speaker, wherein the identification of the speaker can be input by the user, for example, the identification of the speaker can be "Zhang San", "Li Si", etc.
[0108] The clustering identification module can include an audio segmentation module, a number of persons identification module, and a clustering module. The audio segmentation module is configured to segment the audio sub-segments segmented by the speech segmentation module into a plurality of audio short segments according to a preset time length. The audio short segments segmented by the audio segmentation module in the clustering identification module can have a different time length from the audio short segments segmented by the audio segmentation module in the voiceprint identification module. The number of persons identification module can identify the number of speakers included in the audio sub-segments based on the segmented audio short segments. The number of persons identification module can extract audio features in the audio short segments, calculate the similarity of the audio features of each audio short segment and other audio short segments in the audio sub-segment, and then determine the number of speakers included in the audio sub-segment. The audio features can be, for example, mel-frequency cepstral coefficients (MFCC) or voiceprint features, which are not limited in the embodiments of the present application. Then, the number of persons identification module can output the number of speakers included in the data of the whole audio sub-segment output by the speech segmentation module to the clustering module. The clustering module can receive the number of speakers output by the number of persons identification module and the data of the plurality of audio sub-segments output by the audio segmentation module. Then, the clustering module can construct a similarity matrix and a Laplacian matrix, and cluster the data of the audio sub-segments based on the Laplacian matrix. Finally, the clustering module can output one or more clusters (also referred to as categories), each cluster indicating that one or more audio sub-segment data is produced by the same speaker. For example, when the number of persons identification module identifies that only one person speaks in the audio sub-segment, the electronic device 100 can cluster the audio short segments included in the audio sub-segment into the same cluster. The multiple speakers included in the same audio sub-segment can mean that multiple speakers in the audio sub-segment speak alternately and / or multiple speakers in the audio sub-segment speak simultaneously. The clustering identification module can distinguish each cluster using different speaker roles. The result output by the clustering identification module can include the identification of the audio sub-segment corresponding to each cluster and the speaker role identification corresponding to each cluster. For example, the clustering identification module divides five audio short segments in an audio sub-segment into two clusters, the first cluster corresponds to the first audio short segment to the second audio short segment, and the second cluster corresponds to the third audio short segment to the fifth audio short segment. The result output by the clustering identification module can be used to indicate that the speaker corresponding to the first audio short segment and the second audio short segment is “role A”, and the speaker corresponding to the third audio short segment to the fifth audio short segment is “role B”. Here, the speaker role output by the clustering identification module can only distinguish different speakers, and is not associated with specific speakers.That is to say, the same speaker role can be used on different speakers, for example, in one speech recognition result, "role A" is used to identify the speech content of speaker A, and in another speech recognition result, "role A" can be used to identify the speech content of speaker B. The speaker identification output by the voiceprint recognition module can be associated with a specific speaker, and the user can determine a unique speaker according to the speaker identification. For example, the electronic device can store the identification "Zhang San" of speaker A and the voiceprint features of speaker A. Then, when the electronic device uses the voiceprint recognition module to identify the speaker, as long as the audio segment contains the speech content of speaker A, the electronic device uses "Zhang San" to identify the speech content of speaker A. In this way, the user only needs to see the speaker identification "Zhang San" in the speech recognition result to know that the speech content corresponding to the speaker identification is spoken by speaker A.
[0109] In some embodiments, the clustering module can store the clustering results of each audio sub-segment output by the speech segmentation module when clustering the audio sub-segment. In this way, the clustering module can determine whether the data of an audio short segment can be classified into a previously classified cluster when clustering a subsequent audio sub-segment. For example, the clustering module obtains two clusters after clustering the first audio sub-segment, and the two clusters are distinguished by "role A" and "role B". When the clustering module clusters the second audio sub-segment, it can be determined whether there is an audio short segment in the second audio sub-segment that can be merged into the cluster corresponding to "role A" or "role B".
[0110] That is to say, the voiceprint recognition module and the clustering recognition module can both distinguish speakers. The voiceprint recognition module needs to register the speaker in advance when identifying the speaker. Here, the speaker registration refers to extracting the voiceprint of a speaker from the audio data of the speaker, and then storing the voiceprint of the speaker in the storage (such as ROM) or a cloud server. The voiceprint recognition module can only identify the voiceprints of registered speakers. The clustering recognition module can distinguish the roles of speakers without prior registration of the speakers. The electronic device can input the audio sub-segments output by the speech segmentation module into different recognition modules to meet different needs of the user. For example, the electronic device can input the audio sub-segments to be identified into the voiceprint recognition module to identify the registered speakers in the audio segments. Or, the electronic device can input the audio sub-segments to be identified into the clustering recognition module to distinguish different speakers in the audio sub-segments without registration of the speakers. In this way, whether the user registers the speakers or not, the electronic device can distinguish the speakers of each speech content while identifying the speech content in the audio stream, improving the user experience. When the speaker corresponding to the speech in the audio stream is registered, the electronic device 100 provides the speaker identifier of the speaker at the time of registration in the speech recognition content, facilitating the user to clearly distinguish which registered speaker speaks the speech content in the audio stream.
[0111] Figure 3A The method for the clustering recognition module to cluster the audio sub-segments and distinguish the roles of speakers is only an example, and the clustering recognition module can also use other clustering algorithms to distinguish the speakers. Meanwhile, the clustering recognition module can also include more or fewer modules than those shown in the figure, or combine some modules, and the embodiments of the present application do not limit this. In some embodiments, the clustering recognition module can not include the number of people identification module.
[0112] Figure 3B FIG. 1 is a flowchart of the process of the electronic device 100 performing speech recognition. Figure 3B In FIG. 1, the rectangle represents an audio stream. As shown in FIG. 1, the electronic device 100 performs speech recognition mainly includes the following steps: Figure 3B
[0113] 1. The electronic device collects audio through a microphone.
[0114] The electronic device 100 can collect audio through a microphone, and then the electronic device 100 can generate audio data in the form of a digital signal based on the audio collected by the microphone in the form of an analog signal. Multiple audio data are transmitted in the form of an audio stream by using streaming technology.
[0115] 2. The electronic device uses endpoint detection to truncate the audio stream.
[0116] The electronic device 100 can truncate the audio segment in which there is no human voice in the audio stream through endpoint detection. Wherein, the electronic device 100 can determine whether there is human voice through whether the short-time energy value of the audio segment of each time window in the audio stream exceeds the threshold value. Then the valid audio segment can be the audio segment in which the short-time energy value of the audio data in each time window is greater than the preset threshold value.
[0117] 3. The electronic device calls a content recognition module to recognize the speaking content in the valid audio segment.
[0118] The electronic device 100 can call a content recognition module in the artificial intelligence framework engine to recognize the speaking content in the valid audio segment.
[0119] 4. The electronic device calls a speech segmentation module to segment the valid audio segment.
[0120] Wherein, the speech segmentation module can segment the valid audio segment based on semantic features and audio features.
[0121] For example, after the electronic device 100 calls the content recognition module to recognize two sentences: statement 1 and statement 2 from a piece of valid audio segment, the electronic device 100 can input the speaking content 1, the speaking content 2 and the valid audio segment corresponding to the speaking content 1 and the speaking content 2 to the speech segmentation module. Then the speech segmentation module can segment two audio sub-segments: audio sub-segment 1 and audio sub-segment 2 from the piece of valid audio segment. Wherein, the speaking content of the audio sub-segment 1 is the speaking content 1, and the speaking content of the audio sub-segment 2 is the speaking content 2.
[0122] 5. The electronic device calls a voiceprint recognition module to recognize the speaker of the audio sub-segment.
[0123] In some embodiments, the electronic device 100 can call the voiceprint recognition module to recognize the voiceprint of the audio sub-segment segmented by the speech segmentation module, and then recognize the speaker. For example, the electronic device 100 recognizes the voiceprint of the speaker A from the audio sub-segment 1, and then can determine that the speaker of the audio sub-segment 1 is the speaker A, that is, the speaking content 1 is spoken by the speaker A; Similarly, the electronic device 100 recognizes the voiceprint of the speaker B from the audio sub-segment 2, and then can determine that the speaker of the audio sub-segment 2 is the speaker B, that is, the speaking content 2 is spoken by the speaker B.
[0124] 5', the electronic device calls a clustering recognition module to distinguish the speaker of the audio sub-segment.
[0125] In other embodiments, the electronic device 100 can also invoke a clustering recognition module to identify the voiceprints of the audio segments segmented by the speech segmentation module, thereby distinguishing the speakers. For example, the electronic device 100 can use the clustering recognition module to divide audio segment 1 into multiple audio short segments of a fixed length, then extract the audio features of each audio short segment, and then cluster the audio short segments based on the audio features of each audio short segment. Finally, the electronic device can cluster the multiple audio short segments in audio segment 1 into category 1; the electronic device 100 can cluster the audio data in audio segment 2, and then cluster the multiple audio short segments in audio segment 2 into category 2. The electronic device can use different identifiers to distinguish between the two categories. For example, the electronic device can use "speaker A" to identify category 1 and "speaker B" to identify category 2, where "speaker A" and "speaker B" represent two different speakers. This indicates that the speech content 1 in audio segment 1 was spoken by speaker A, and the speech content 2 in audio segment 2 was spoken by speaker B.
[0126] In some embodiments, one or more of the content recognition module, speech segmentation module, voiceprint recognition module, voiceprint registration module, and clustering recognition module in the AI framework engine can also be deployed on a cloud server. The electronic device 100 can segment the valid audio in the audio stream and send the segmented audio data to the cloud server. The cloud server then identifies the spoken content and corresponding speaker contained in the audio data and returns the recognition result to the electronic device 100. The method by which the cloud server identifies the spoken content and corresponding speaker contained in the audio data can refer to the method by which the electronic device identifies the spoken content and corresponding speaker contained in the audio data, and will not be elaborated here.
[0127] The following describes the application scenarios of the speech recognition method provided in the embodiments of this application.
[0128] Figures 4A-4W This is a schematic diagram of user interface (UI) involved in the speech recognition method provided in the embodiments of this application.
[0129] This section first introduces a scenario where electronic devices perform speech recognition through a clustering recognition module when the speaker is not registered.
[0130] like Figure 4A As shown, Figure 4AAn exemplary illustration shows a home screen interface 400 on an electronic device 100. The home screen interface 400 may include desktop icons for one or more applications, including a desktop icon 401 for a note-taking application. The note-taking application is used to create and edit notes. Notes may include titles and content, and the content of the notes may include, but is not limited to, text, images, audio, etc. The electronic device can store the data of the user-created notes in its memory.
[0131] The electronic device 100 can detect user actions, such as clicks, applied to desktop icons 401. In response to this action, the electronic device 100 can display, for example... Figure 4B The main interface 410 of the note-taking application is shown. The main interface 410 of the note-taking application may include a note display area 411 and a note adding button 412.
[0132] The note display area 411 may include one or more note list items, including note list item 411A. Note list item 411A may correspond to a note, providing users with an entry point to edit the note title or content, etc.
[0133] For example, such as Figure 4B As shown, note list item 411A can display the note's title (such as "Welcome to Notes"), the time the note was created (such as "10 minutes ago"), and a preview of the note's content (such as "You can create notes here;"), helping users quickly understand the relevant information about the note.
[0134] The note-adding button 412 is used to create a new note.
[0135] The electronic device 100 can receive the user's click on the note-adding button 412. In response to this operation, the electronic device 100 can create a new note and display it as shown below. Figure 4C The note editing interface 420 shown is used to edit the title or content of notes. The note editing interface 420 may include: a title editing box 421, a note content display area 422, a toolbar 423, a soft keyboard 424, a save button 425, and a return button 426.
[0136] The title editing box 421 can display the title of the current note. Figure 4C The default title (“Title”) of the note can be displayed in the title editing box 421.
[0137] The note content display area 422 is used to display the content of the current note.
[0138] The toolbar 423 can include one or more tool buttons, each of which is used to implement different functions. The toolbar 423 can include a recording button 423A, which is used to start a voice recognition function.
[0139] The soft keyboard 424 is used to edit the title text. The soft keyboard 424 can include a zoom button 424A, which is used to hide the soft keyboard 424 displayed in the note editing interface 420.
[0140] The save button 425 is used to store the data of the current note, which can include the title, content, creation time, and the like of the note.
[0141] The back button 426 is used to return to the main interface 410 of the note application.
[0142] The electronic device 100 can receive an operation of the user editing the title text through the soft keyboard 424, and in response to the operation, the electronic device 100 can synchronize the text entered by the user to be displayed in the title editing box 421. As shown in Figure 4D , the text displayed by the electronic device 100 in the title editing box 421 is changed from "title" to "lunch place". Further, the electronic device 100 can receive an operation of the user clicking the zoom button 424A, and in response to the operation, the electronic device can hide the soft keyboard 424, as shown in Figure 4E .
[0143] Further, the electronic device 100 can start the voice recognition function through the recording button 423A in the toolbar 423. As shown in Figure 4E , the electronic device 100 can receive an operation of the user clicking the recording button 423A, and in response to the operation, the electronic device 100 can display a menu bar 430 as shown in Figure 4F .
[0144] The menu bar 430 can include multiple options for selecting a mode of voice recognition, or changing parameters when voice recognition, and the like. The multiple options can include a "recognize only voice content" option, a "register-free speaker distinction" option, a "voiceprint recognition speaker" option, and a "voiceprint management" option. The "recognize only voice content" option can be used to instruct the electronic device to only recognize the speaking content in the audio, without recognizing the speaker. The "register-free speaker distinction" option is used to instruct the electronic device to distinguish the speaker using a register-free speaker manner, and the electronic device 100 can use a clustering recognition module to distinguish the speaker in the audio. The "voiceprint recognition speaker" option is used to instruct the electronic device to recognize the speaker based on the voiceprint, and the electronic device 100 can use a voiceprint recognition module to recognize the speaker in the audio. The "voiceprint management" option is used to provide an entry for the user to manage the voiceprint.
[0145] The electronic device can receive a user operation of tapping the "unregistered speaker distinction" option in the menu bar 430, and in response to the operation, the electronic device 100 can collect audio through the microphone and display a voice recognition frame 431 for displaying a voice recognition result. Figure 4G The voice recognition frame 431 is illustrated as a frame for displaying a voice recognition result. The voice recognition frame 431 can include an audio total duration 431A, a waveform display area 431B, a speaking content display area 431C, and a stop button 431D.
[0146] The audio total duration 431A is used to display the time length of the audio collected by the electronic device in the current note.
[0147] The waveform display area 431B is used to display the waveform of the audio collected by the electronic device.
[0148] The speaking content display area 431C is used to display the speaking content recognized by the electronic device from the audio, and the speaking content display area 431C can distinguish the speaker using different identifiers.
[0149] The stop button 431D is used to stop recording the audio.
[0150] The electronic device 100 can obtain an audio stream through the microphone, and then transmit the audio data in the audio stream to the content recognition module. At the same time, the electronic device 100 performs endpoint detection on the audio stream, and when the electronic device 100 detects that a certain frame of audio data in the audio stream does not contain human voice, the electronic device 100 can cut off the audio stream.
[0151] As illustrated in FIG. 4B, as the electronic device 100 collects audio through the microphone, the audio total duration changes from 00:00:00 to 00:00:13, which can indicate that the electronic device 100 collects 13 seconds of audio through the microphone. At the same time, the waveform display area 431B can display the waveform of the 13 seconds of audio. The note application can call the artificial intelligence framework engine when the microphone collects the audio, and then the artificial intelligence framework engine can process the audio stream in real time and feed back the voice recognition result to the note application. Figure 4H Figure 4G As illustrated in FIG. 4B, as the electronic device 100 collects audio through the microphone, the audio total duration changes from 00:00:00 to 00:00:13, which can indicate that the electronic device 100 collects 13 seconds of audio through the microphone. At the same time, the waveform display area 431B can display the waveform of the 13 seconds of audio. The note application can call the artificial intelligence framework engine when the microphone collects the audio, and then the artificial intelligence framework engine can process the audio stream in real time and feed back the voice recognition result to the note application.
[0152] The speaking content display area 431C can display the result of the artificial intelligence framework engine recognizing the audio stream in real time. Figure 4H In the illustrated embodiment, the speaking content display area 431C can display a voice recognition record, which can include an identifier ("Recognizing…"), speaking content ("What do we eat for lunch today? How about eating grilled fish, or we can also eat hot pot."), and the starting time of the above-mentioned speaking content in the audio ("00:00:03").
[0153] AsFigure 4I As shown, the total length of the audio recorded by the electronic device 100 becomes 00:00:18, Figure 4I The user interface displayed by the electronic device 100 in the embodiment shown can be regarded as Figure 4H The electronic device 100 in the embodiment shown obtains the speech content display area 431C after collecting 5 seconds of audio again. The speech content display area 431C can include three speech recognition records, wherein the first speech recognition record can include an identity identifier (“Speaker A”), speech content (“Where shall we go for lunch today?”), and the start time and end time of the speech content in the audio (“00:00:03-00:00:07”). The speech recognition record is obtained by the electronic device after recognizing the speech from the third second to the seventh second. The second speech recognition record of the speech content display area 431C can include an identity identifier (“Speaker B”), speech content (“How about going to eat grilled fish, or going to eat hot pot?”), and the start time and end time of the speech content in the audio (“00:00:07-00:00:13”). The third speech recognition record can include an identity identifier (“Recognizing…”), speech content (“Let's go to eat hot pot.”), and the start time of the speech content in the audio (“00:00:15”). Here, the first and second speech recognition records can be regarded as Figure 4H The final recognition result of the speech recognition record in the speech content display area 431C shown. When the electronic device 100 displays the speech recognition record in the speech content display area 431C, the speech content of different speakers can be segmented, so that the reading effect of the speech content display area 431C can be avoided.
[0154] Here, the process of the electronic device 100 performing speech recognition in the scenario shown is described. In the scenario shown, the electronic device 100 can perform speech recognition on the audio recorded by the electronic device 100. Figures 4G-4I The process of the electronic device 100 performing speech recognition in the scenario shown is described. In the scenario shown, the electronic device 100 can perform speech recognition on the audio recorded by the electronic device 100. Figure 4GIn the illustrated embodiment, the electronic device starts the microphone to obtain audio, and then the electronic device 100 can determine that the audio from 00:00:00-00:00:03 does not contain human voice by an endpoint detection method (for example, by calculating whether the short-time energy value of the audio exceeds a threshold value), and the electronic device 100 can not input the audio data in the time period from 00:00:00-00:00:03 of the audio stream into the content recognition module. When the time at which the electronic device records the audio reaches 00:00:03, the electronic device 100 detects that human voice appears in the audio stream, and then the electronic device 100 can start to input the audio data in the audio stream into the content recognition module. The content recognition module can process the audio data in the audio stream frame by frame. For example, taking the speaking content of a frame of audio data as “today” as an example, the content recognition module first identifies “today” from the audio stream, and the electronic device can display “today” to the speaking content display area 431C. Since the speaking content in the valid audio segment has not been completely recognized, the electronic device 100 can display the identity as “being recognized”. Then the electronic device 100 can input the next frame of audio data into the content recognition module, and input “today” into the embedding layer. The content recognition module can process the next frame of audio data by the encoder, process “we” by using the embedding layer and the language prediction model, and then the content recognition module can recognize the next frame of audio data. It should be noted that the speaking content that can be contained in a frame of audio data is only for ease of understanding, and in fact the content recognition module can identify “today” from multiple frames of audio data respectively. The content recognition module can process the audio data in the audio stream frame by frame until the truncation position of the audio stream is determined. Finally, the content recognition module can identify “we go to eat where for lunch today? How about eating grilled fish, or we can also eat hot pot.” Although “we go to eat where for lunch today?” and “how about eating grilled fish, or we can also eat hot pot.” are spoken by two speakers, but the two sentences may be separated by a very short interval due to partial overlap, and the electronic device does not separate the two sentences in the preliminary endpoint detection, and the electronic device 100 detects that the audio does not contain the speaker when the sentence “or we can also eat hot pot.” ends. At this time, the electronic device 100 can truncate the position (i.e., at 00:00:13) in the audio stream that is detected to not contain the speaker. The electronic device can input the speaking content text “we go to eat where for lunch today? How about eating grilled fish, or we can also eat hot pot.” and the audio segment from 00:00:03-00:00:13 that is recognized to the speech segmentation module. Here, the audio segment from 00:00:03-00:00:13 can also be referred to as a valid audio segment.The speech segmentation module then divides the 00:00:03-00:00:13 segment into two audio sub-segments based on semantic and audio features. One audio sub-segment is from 00:00:03 to 00:00:07, containing the statement: "Where should we eat lunch today?" The other audio sub-segment is from 00:00:07 to 00:00:13, containing the statement: "How about we go eat grilled fish, or we can go eat hot pot." In this way, the speech segmentation module separates the long speech from two speakers. The speech segmentation module then sequentially inputs the two audio sub-segments into the clustering recognition module. The clustering recognition module can cluster the two audio sub-segments into different clusters. The electronic device can then use "Speaker A" to identify the first audio sub-segment and "Speaker B" to identify the second audio sub-segment. Thus, the electronic device 100 can... Figure 4H The speech recognition record shown is segmented into two parts. Figure 4I The first and second voice recognition records are shown.
[0155] After the electronic device 100 collects another 2 seconds of audio, it can determine that the audio stream does not contain human voice after 00:00:18, and therefore, the electronic device 100 will truncate the audio stream at 00:00:18. The clustering and recognition module can then distinguish the speaker in the audio sub-segment from 00:00:15 to 00:00:18. Finally, the clustering and recognition module can determine that the audio sub-segment from 00:00:15 to 00:00:18 and the audio sub-segment from 00:00:03 to 00:00:07 can be clustered into the same cluster. This indicates that the two audio sub-segments contain the same speaker's voice content. Therefore, the electronic device 100 can also use "Speaker A" to label the audio sub-segment from 00:00:15 to 00:00:18. Figure 4J As shown, the total duration of the audio recorded by electronic device 100 becomes 00:00:20. Here, the third voice recognition record can be the state of recognition completion. For example, the third voice recognition record may include: identity identifier ("speaker A"), speech content ("Let's go eat hot pot.") and the start and end times of the speech content in the audio ("00:00:15-00:00:18").
[0156] like Figure 4J As shown, the electronic device 100 receives a user's click on the stop button 431D. In response to this operation, the electronic device 100 can stop recording audio and simultaneously stop inputting the audio stream to the artificial intelligence framework engine for speech recognition. The electronic device 100 can then display... Figure 4KThe play button 431E and slider 431F are shown, while the stop button 431D is transformed into a record button 431G. The play button 431E is used to play the recorded audio, the slider 431F and its slider are used to view and adjust the audio playback progress, and the record button 431G can be used to instruct the electronic device to continue recording audio and to identify spoken content in the audio. Figure 4K As shown, the electronic device 100 can receive a user's click operation on the save button 425. In response to this operation, the electronic device 100 can store the note title and note content in the note editing interface 420. The note content includes both the audio of the current note recorded by the electronic device 100 in the note editing interface 420, and the speech content and identity identifiers recognized by the electronic device 100 while recording the audio. Figure 4L As shown, the electronic device 100 can receive a user's click on the back button 426, and in response to this operation, the electronic device 100 can return to the previous state. Figure 4M The main interface 410 of the note-taking application is shown. The note display area 411 includes a note list item 411B, which provides access to the "Lunch Location" note (i.e.,...). Figures 4C-4L The illustrated embodiment shows the entry point for notes created and edited by the electronic device. When the user clicks on note list item 411B, the electronic device can display the following: Figure 4M The note editing interface 420 is shown. Note list item 411B can display the note's title ("Lunch Location") and a partial preview of the note's content (e.g., a spoken text "Where are we going for lunch today?").
[0157] The following describes a scenario where electronic devices use a voiceprint recognition module to perform speech recognition on registered speakers.
[0158] like Figure 4N As shown, the electronic device 100 can receive the user's action of clicking the note-adding button 412 on the main interface 410 of the note-taking application. In response to this action, the electronic device 100 can create a new note and display it. Figure 4O The note editing interface 420 is shown. In this interface, the electronic device 100 can receive and respond to the user's operation to edit the note title to "Meeting Time", and then the electronic device 100 can display "Meeting Time" in the title editing box 421.
[0159] like Figure 4O As shown, the electronic device 100 can receive a user's click on the record button 423A, and in response to this operation, the electronic device 100 can display as shown in the image. Figure 4P The menu bar shown is 430C. For more information about the menu bar 430C, please refer to [link / reference needed]. Figure 4FThe description of the illustrated embodiment will not be repeated here. The electronic device 100 can receive a user's click on the "Voiceprint Management" option in the menu bar 430C. In response to this operation, the electronic device 100 can display a voiceprint management interface 440, which is used to manage voiceprints. The voiceprint management interface 440 may include a speaker registration button 441 and a voiceprint list 442. The speaker registration button 441 is used to register a speaker's voiceprint. The voiceprint list 442 can be used to view and edit voiceprints. The voiceprint list 442 may include one or more voiceprint options, where each voiceprint option corresponds to a speaker's voiceprint, providing the user with an entry point to edit that speaker's voiceprint. The user can edit the voiceprint of the speaker corresponding to the voiceprint option through the voiceprint option, for example, by re-extracting the voiceprint, deleting the voiceprint, etc. For example, the voiceprint list 442 may include the "Zhang San" option and the "Li Si" option, indicating that the current electronic device 100 stores the voiceprints of two speakers, one of which corresponds to the speaker's identifier "Zhang San" and the other to the speaker's identifier "Li Si".
[0160] like Figure 4Q As shown, the electronic device 100 can receive a user's click operation on the speaker registration button 441. In response to this operation, the electronic device 100 can register the speaker's voiceprint. The electronic device 100 can send the audio data of a speaker's speech and the speaker's identifier to the voiceprint registration module. The voiceprint registration module can extract the voiceprint from the audio data through the audio segmentation module and the voiceprint feature extraction module. Then, the voiceprint registration module can store the speaker's identifier and the speaker's voiceprint in the memory. Correspondingly, the electronic device 100 can display the voiceprint options corresponding to the newly registered speaker in the voiceprint list 442, such as... Figure 4R As shown, a newly registered speaker option, namely the "Wang Wu" option, has been added to the voiceprint list 442. This indicates that the electronic device 100 has added a speaker's voiceprint, and the speaker's identifier is "Wang Wu".
[0161] Electronic device 100 can receive a user's click on the back button 426. In response to this operation, electronic device 100 can exit the voiceprint management interface 440 and redisplay the following: Figure 4S The note-editing interface shown is 420. (As shown in the image...) Figure 4S As shown, electronic device 100 can receive and respond to the user's click on the "Voiceprint Recognition Speaker" option in menu bar 430C, and then electronic device 100 can start recording audio and display as shown. Figure 4TThe voice recognition box 431 is shown. Meanwhile, the electronic device 100 can process the audio stream collected by the microphone through the artificial intelligence framework engine, wherein the electronic device 100 can extract a voiceprint from the audio stream through a voiceprint recognition module in the artificial intelligence framework engine, and match the extracted voiceprint with a pre-stored voiceprint. Further, the electronic device 100 can display the result of voice recognition on the display screen. Optionally, when no voiceprint of any speaker is stored in the electronic device 100, the electronic device 100 can receive an operation of the user clicking the "voiceprint recognition speaker" option in the menu bar 430C, and in response to the operation, the electronic device 100 can prompt the user to register the voiceprint of the speaker on the electronic device, so that the electronic device can identify the speaker using the voiceprint recognition method.
[0162] Figure 4U is the state of the electronic device 100 after collecting 13 seconds of audio. Figure 4T The speech content display area 431C can display the result of voice recognition by the electronic device 100 based on the 13 seconds of audio. The speech content display area 431C can display two voice recognition records, wherein the first voice recognition record can include: the speaker identification ("Zhang San"), the speech content ("What time is the meeting today?"), and the start time and end time of the speech content in the audio ("00:00:02-00:00:05"). The second voice recognition record can include: the speaker identification ("Wang Wu"), the speech content ("Hmm, it should be 2 o'clock in the afternoon, right."), and the start time and end time of the speech content in the audio ("00:00:07-00:00:10"). As shown in Figure 4V The electronic device 100 can receive an operation of the user clicking the stop button 431D, and the electronic device 100 can stop recording the audio and display the play button 431E and the slide bar 431F as shown in Figure 4W The play button 431E and the slide bar 431F are described above, and will not be described here. The electronic device can store the recorded audio and the recognition result displayed in the speech content display area 431C into the memory, so that the user can check at any time.
[0163] In some embodiments, the electronic device 100 can also receive an operation of the user filtering the speech content of one or more speakers, and in response to the operation, the electronic device 100 can hide the speech content and identification of other speakers. For example, the user can click the "filter" button 431G, and in response to the operation, the electronic device 100 can display a filter menu 431H as shown in Figure 4WThe illustrated embodiment is an example, the electronic device 100 can receive the operation of the user filtering the speech content of "Zhang San", in response to the operation, the electronic device 100 can hide the speech recognition record (for example, the speech recognition record of "Wang Wu") other than "Zhang San", and only display the speech recognition record of Zhang San. In this way, the user can only view the speech content of the filtered speaker, improving the user's use experience.
[0164] In some embodiments, when the electronic device identifies the speaker using the voiceprint recognition module in response to the user's operation of selecting the "voiceprint recognition speaker" option, it may occur that the voiceprint of part of the audio sub-fragments cannot be matched with the voiceprint of the speaker already stored in the electronic device or the cloud server. This may be due to the fact that the voiceprint of the speaker in the audio sub-fragment is not stored in the electronic device or the cloud server. The electronic device 100 can invoke the clustering recognition module to cluster the audio features in the audio sub-fragment, thereby distinguishing the speaker in the audio sub-fragment. Finally, the electronic device 100 can identify the speaker in the audio sub-fragment using the speaker role identifier (such as "Speaker A" and "Speaker B"). In this way, even if an unregistered speaker appears in the audio when the user selects the voiceprint recognition speaker, the electronic device 100 can distinguish the unregistered speaker in the audio.
[0165] Compared with Figures 4F-4L The illustrated embodiment, Figures 4P-4U The identification of the speaker identified by the electronic device 100 in the illustrated embodiment can refer to a specific speaker, because the voiceprint of the speaker is stored in the memory of the electronic device 100 or the cloud server. Therefore Figures 4P-4U The illustrated embodiment compared with Figures 4F-4L The illustrated embodiment has more accurate identification results.
[0166] When the electronic device 100 performs speech recognition on the registered speaker, the method by which the electronic device processes the audio stream through the content recognition module and the speech segmentation module can refer to the method by which the electronic device performs speech recognition on the unregistered speaker in the foregoing embodiments, which will not be described here. When the electronic device 100 performs speech recognition on the registered speaker, the electronic device 100 can invoke the voiceprint recognition module to identify the voiceprint in the audio sub-fragment after processing the audio stream using the content recognition module and the speech segmentation module, thereby determining the speaker of the audio sub-fragment.
[0167] In other words, when a user has not registered as a speaker, they can choose the speaker role recognition method without registration. In this case, the electronic device 100 will distinguish different speaker roles through a clustering recognition module. When a user has already registered as a speaker, they can use voiceprint recognition to identify the speaker. In this way, the electronic device can identify the specific speaker through the voiceprint recognition module. Thus, regardless of whether the user has registered as a speaker, the electronic device can distinguish different speakers during speech recognition, thereby improving the user experience.
[0168] Figure 5 This is a flowchart illustrating a speech recognition method provided in an embodiment of this application.
[0169] like Figure 5 As shown, this speech recognition method may include, but is not limited to, the following steps:
[0170] S501. The electronic device receives a first operation from the user to select a voice recognition method, wherein the voice recognition method includes a first recognition method and a second recognition method.
[0171] by Figure 4F Taking the note-editing interface 420 as an example, the first operation can be the user clicking the "Recognize only speech content" option, the "Distinguish speakers without registration" option, or the "Voiceprint recognition speaker" option in the menu bar 430. The first recognition method can refer to the electronic device using voiceprint recognition to identify the speaker when recognizing the speaker in the speech. When the first operation selects the "Voiceprint recognition speaker" option, it indicates that the user has selected the voiceprint recognition method to identify the speaker, and the electronic device can call the voiceprint recognition module to identify the speaker. The second recognition method can refer to the electronic device using cluster recognition to identify the speaker when recognizing the speaker in the speech. When the first operation selects the "Distinguish speakers without registration" option, it indicates that the user has selected the cluster recognition method to identify the speaker, and the electronic device can call the cluster recognition module to identify the speaker.
[0172] S502, The electronic device acquires a first audio segment, the first audio segment including a first sub-segment corresponding to the first spoken content.
[0173] The first audio segment can be a valid audio segment, and the first audio segment can contain a first sub-segment, the content of which is the first spoken content. Figures 4H-4IThe embodiment shown is an example, and the first audio segment can be a valid audio segment of 00:00:03-00:00:13 in the audio stream. The valid audio segment can refer to an audio segment in which the sound is continuous. Here, the sound continuity can refer to the fact that the human voice part in the audio is continuously present or appears, and the human voice part is not interrupted or the interruption duration does not exceed a preset duration, that is, the duration of the speaker's pause in the audio does not exceed the preset duration. The electronic device 100 can determine whether the sound in the audio stream is continuous by calculating the short-time energy value or the short-time zero-crossing rate of the audio in the sliding window.
[0174] S503, the electronic device determines the voice recognition mode selected by the first operation.
[0175] The electronic device can determine the voice recognition mode based on the first operation of the user. The method in which the electronic device determines the voice recognition mode selected by the first operation can refer to the description in step S501, which will not be repeated here. The electronic device can call the content recognition module to recognize the speaking content of the first audio segment, and then the electronic device can input the first audio segment and the speaking content of the first audio segment into the voice segmentation module, and then the voice segmentation module can segment the first audio segment based on the semantic text features of the speaking content and the audio features of the first audio segment, etc. to obtain the first sub-segment. Then the electronic device can input the first sub-segment into the voiceprint recognition module or the clustering recognition module based on the first operation.
[0176] When the electronic device determines that the first operation selects the first recognition mode, step S504 can be performed, and when the electronic device determines that the first operation selects the second recognition mode, step S505 can be performed.
[0177] S504, the electronic device determines that the voiceprint feature of the first sub-segment matches the first voiceprint feature, and the electronic device displays the first speaking content and the first identifier, the first identifier indicating that the first speaking content is spoken by the first speaker.
[0178] In the case where the first operation selects the first recognition mode, the electronic device can call the voiceprint recognition module to process the first sub-segment. The voiceprint recognition module can determine that the voiceprint feature of the first sub-segment matches the first voiceprint feature, and then the electronic device can display the first speaking content and the first identifier. The electronic device can pre-store the first identifier and the first voiceprint feature of the first speaker, so that the speaker can be matched based on the stored first voiceprint feature.
[0179] In the case where the first operation selects the first recognition mode, the electronic device can call the voiceprint recognition module to process the first sub-segment. The voiceprint recognition module can determine that the voiceprint feature of the first sub-segment matches the first voiceprint feature, and then the electronic device can display the first speaking content and the first identifier. The electronic device can pre-store the first identifier and the first voiceprint feature of the first speaker, so that the speaker can be matched based on the stored first voiceprint feature. Figure 4RThe illustrated embodiment is an example, and the electronic device 100 stores the voiceprints of Zhang San, Li Si, and Wang Wu. Here, the first voiceprint feature can refer to the voiceprint feature of Zhang San, and the first speaker can refer to Zhang San. In this way, after the voiceprint recognition module obtains the audio sub-fragment, it can extract the voiceprint feature in the audio sub-fragment and match it with the voiceprints of Zhang San, Li Si, and Wang Wu. For example, Figure 4U As shown, the first sub-fragment can be an audio sub-fragment of 00:00:02-00:00:05 in the audio stream. The audio sub-fragment of 00:00:02-00:00:05 can be obtained by the speech segmentation module based on the first speech content ("What time is the meeting today?") to segment the valid audio fragment. The electronic device can display the first speech content and the first identifier (i.e., the "Zhang San" identifier in Figure 4U ).
[0180] S505, the electronic device determines that the first speech content is spoken by one speaker, and the electronic device displays the first speech content and a second identifier, where the second identifier is used to identify the speaker of the first speech content.
[0181] In the case where the first operation selects the second identification mode, the electronic device can call the clustering identification module to process the first sub-fragment. The clustering identification module can determine that the first speech content is spoken by one speaker, and the clustering identification module can display the first speech content and a second identifier, where the second identifier is used to identify the speaker of the first speech content.
[0182] For example, in the illustrated embodiment, the first sub-fragment can be, for example, an audio sub-fragment of 00:00:03-00:00:07 in the audio stream. This audio sub-fragment can be obtained by the speech segmentation module to segment the valid audio fragment. Here, the first speech content can be, for example, "Where do we eat today?". The clustering identification module can then determine that this audio sub-fragment contains only one speaker, and the clustering identification module can then cluster the audio features of each audio short fragment in this audio sub-fragment into the same cluster (or class), and assign a second identifier to this class. Here, the second identifier can be, for example, "Speaker A". Figure 4I That is, the electronic device can provide an entry for pre-registration and registration-free speakers while identifying the speech content. In this way, the electronic device can distinguish the speaker of each speech content regardless of whether the user registers the speaker, and display the result of distinguishing the speaker on the display screen. When the user selects a voice recognition mode that requires pre-registration of the speaker, the electronic device can provide a more accurate speaker identification (such as the name of the speaker) for the user, so that the user can determine which speaker is speaking based on the speaker identification.
[0183]
[0184] In some embodiments, after the electronic device acquires the first audio segment, the method further includes: the electronic device splits the first audio segment into a plurality of sub-segments, the plurality of sub-segments include the first sub-segment, and the plurality of sub-segments have different semantics. The electronic device can split the first audio segment by using a speech segmentation module to obtain the plurality of sub-segments. In the embodiments of the present application, the sub-segment can also be referred to as an audio sub-segment. The speech segmentation module can split the first audio segment based on the text semantic features of the speech content contained in the first audio segment and the audio features of the first audio segment, and finally the speech segmentation module can split a plurality of sub-segments with different semantics. For example, as shown in the embodiments, the first audio segment can include "Where do we go for lunch today? How about eating grilled fish, or we can also eat hot pot", and the speech segmentation module can split two sub-segments, wherein the first sub-segment has speech content "Where do we go for lunch today?", and the other sub-segment has speech content "How about eating grilled fish, or we can also eat hot pot". The semantics of the speech content of the two sub-segments are different. Figures 4H-4I
[0185] In some embodiments, the first audio segment includes a second sub-segment, the second sub-segment includes second speech content of a second speaker, and the electronic device does not store a second voiceprint feature of the second speaker. In this case, the method further includes: in a case where the first operation selects the second recognition mode, the electronic device determines that the second speech content is spoken by one speaker, and the electronic device displays the second speech content and a third identifier, the third identifier being used to identify the speaker of the second speech content.
[0186] In a case where the first recognition mode (i.e., voiceprint recognition speaker) is selected, the electronic device calls a voiceprint recognition module to determine whether the voiceprint in each sub-segment matches the pre-stored voiceprint feature. When the electronic device does not store the second voiceprint feature of the second speaker, the electronic device cannot identify the second speaker by calling the voiceprint recognition module. The electronic device can call a clustering recognition module to process the second sub-segment, and then use a role identifier (e.g., speaker A) to distinguish the speaker of the second sub-segment. In this way, in a case where the user selects voiceprint recognition, the electronic device can also distinguish the speech content of an unregistered speaker.
[0187] In some embodiments, the first audio segment includes a second sub-segment, the second sub-segment includes second speech content of a second speaker, and the electronic device stores a second voiceprint feature of the second speaker and a fourth identifier. In this case, the method further includes: in a case where the first operation selects the first recognition mode, the electronic device determines that the voiceprint feature of the second sub-segment matches the second voiceprint feature, and the electronic device displays the second speech content and the fourth identifier, the fourth identifier indicating that the second speech content is spoken by the second speaker.
[0188] by Figure 4U Taking the illustrated embodiment as an example, the first sub-segment can be an audio sub-segment from 00:00:02 to 00:00:05, and the second sub-segment can be an audio sub-segment from 00:00:07 to 00:00:10. The electronic device stores the voiceprint features of Zhang San and Wang Wu, as well as their identifiers. Here, Zhang San can be the first speaker, his voiceprint feature is the first voiceprint feature, and his name "Zhang San" is the first identifier; Wang Wu can be the second speaker, his voiceprint feature can be called the second voiceprint feature, and his name "Wang Wu" can be called the fourth identifier. Therefore, if voiceprint recognition is selected in the first operation, the voiceprint recognition module can identify that the voiceprint of the first sub-segment matches Zhang San's first voiceprint feature, and the electronic device can then display Zhang San's first identifier next to the first spoken content. Similarly, the voiceprint recognition module can identify that the voiceprint of the second sub-segment matches Wang Wu's second voiceprint feature, and the electronic device can then display Wang Wu's fourth identifier next to the second spoken content.
[0189] In some embodiments, after the electronic device displays the second speaking content and the fourth identifier, the method further includes: the electronic device receiving a second operation, the second operation being used to filter the first speaker; in response to the second operation, the electronic device hiding the display of the second speaking content and the fourth identifier.
[0190] The second operation can be to filter the speech content of a certain speaker. In response to this operation, the electronic device can hide the speaker's identifier and speech content, and only display the speech content of the filtered speaker.
[0191] In some embodiments, the first audio segment further includes a third sub-segment, the third sub-segment including third speaking content, the speaker of the third speaking content being different from the speaker of the first speaking content, the method further includes: when the first operation selects a second recognition method, the electronic device determines that the third speaking content is spoken by a speaker, the electronic device displays the third speaking content and a fifth identifier, the fifth identifier being used to indicate that the speaker of the third speaking content is different from the speaker of the first speaking content.
[0192] That is to say, in the case that the first operation selects the second recognition manner, the electronic device can call the clustering recognition module to determine the speaker in the third sub-fragment. The clustering recognition module determines that the third sub-fragment contains one speaker. After the clustering recognition module clusters the audio features of the audio short fragments in the third sub-fragment, it can be determined that the audio features contained in the third sub-fragment and the audio features contained in the first sub-fragment do not belong to the same category, and then the electronic device can use the fifth identifier to identify the speaker of the third sub-fragment, where the fifth identifier is different from the second identifier. For example, the fifth identifier can be "Speaker B", and the second identifier can be "Speaker A", both of which are used to indicate that the speaker of the third speaking content is different from the speaker of the first speaking content.
[0193] In some embodiments, the first audio segment further includes a fourth sub-fragment, and the fourth sub-fragment includes fourth speaking content. The method further includes: in the case that the first operation selects the second recognition manner, the electronic device determines that the fourth speaking content is spoken by one speaker, and the speaker of the fourth speaking content is the same as the speaker of the first speaking content. The electronic device displays the fourth speaking content and displays the second identifier again at the first position associated with the fourth speaking content. The second identifier displayed at the first position is used to indicate the speaker of the fourth speaking content.
[0194] That is to say, in the case that the first operation selects the second recognition manner, the electronic device determines that the speaker contained in the fourth sub-fragment is the same as the speaker of the first speaking content through the clustering recognition module, that is, the audio features of the audio short fragments contained in the fourth sub-fragment can be classified into the same category as the audio features of the audio short fragments contained in the first sub-fragment. In this way, the electronic device can assign the same identifier (i.e., the second identifier) to the fourth sub-fragment as the first sub-fragment, indicating that the speaker contained in the fourth sub-fragment is the same as the first sub-fragment. The electronic device can display the second identifier again at the first position. Here, displaying the second identifier again at the first position can mean that the electronic device displays the second identifier at the first position in addition to displaying the second identifier at the second position. The second identifier displayed at the second position is used to indicate the speaker of the first speaking content, and the second identifier displayed at the first position is used to indicate the speaker of the fourth speaking content. The first position is different from the second position.
[0195] In some embodiments, before the electronic device receives the first operation of the user selecting the voice recognition manner, the method further includes: the electronic device receives a second audio segment and a first identifier input by the user, the speaker of the second audio segment being a first speaker; the electronic device extracts a first voiceprint feature of the first speaker from the second audio segment; and the electronic device stores the first voiceprint feature and the first identifier, where the first voiceprint feature and the first identifier are stored in association.
[0196] The electronic device can receive an operation of the user registering a speaker (for example,Figure 4Q The electronic device receives an operation of a user clicking a speaker registration button 441, and then the electronic device can receive a second audio segment input by the user and a first label. The voiceprint registration module can input the second audio segment and the first label into the voiceprint identification module, and the voiceprint identification module can extract a first voiceprint feature of the first speaker from the second audio segment. Then the electronic device can store the first voiceprint feature and the first label. For example, as shown in the following table, the electronic device can store the first voiceprint feature of the speaker "Wang Wu" and the first label "Wang Wu". Figure 4Q The electronic device receives an operation of a user clicking a speaker registration button 441, and then the electronic device can receive a second audio segment input by the user and a first label. The voiceprint registration module can input the second audio segment and the first label into the voiceprint identification module, and the voiceprint identification module can extract a first voiceprint feature of the first speaker from the second audio segment. Then the electronic device can store the first voiceprint feature and the first label. For example, as shown in the following table, the electronic device can store the first voiceprint feature of the speaker "Wang Wu" and the first label "Wang Wu".
[0197] In some embodiments, the method further includes: the electronic device acquires a second audio segment through the microphone during the display of the first speaking content; and the electronic device identifies the speaking content and the speaker of the second audio segment based on the voice recognition mode selected by the first operation.
[0198] That is, the electronic device performs real-time voice recognition while recording the audio. The electronic device can also acquire a second audio segment during the display of the first speaking content. Then the electronic device can identify the speaking content and the speaker of the second audio segment after displaying the first speaking content.
[0199] The above-described embodiments are merely intended for describing and illustrating, but not limiting the technical solutions of the present application; even though the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0200] In the above embodiments, according to the context, the term "when" can be interpreted as meaning "if" or "after" or "in response to determining" or "in response to detecting". Similarly, according to the context, the phrase "on determining" or "if detecting (the stated condition or event)" can be interpreted as meaning "if determining" or "in response to determining" or "on detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)".
[0201] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk) and the like.
[0202] Those of ordinary skill in the art understand that all or part of the processes in the above embodiments can be implemented by a computer program to instruct the relevant hardware, which can be stored in a computer readable storage medium. The program can include the processes of the above method embodiments when executed. The aforementioned storage medium includes ROM or random access memory (RAM), magnetic disk or optical disk, and various media that can store program codes.
Claims
1. A voice recognition method, the method being applied to an electronic device, characterized by, The electronic device stores a first identifier and a first voiceprint feature of the first speaker, and the method includes: The electronic device receives a first operation from the user to select a voice recognition method, wherein the voice recognition method includes a first recognition method and a second recognition method; The electronic device acquires a first audio segment, the first audio segment containing a first sub-segment corresponding to the first spoken content; When the first identification method is selected in the first operation, the electronic device determines that the voiceprint feature of the first sub-segment matches the first voiceprint feature, and the electronic device displays the first speaking content and the first identifier, wherein the first identifier indicates that the first speaking content was spoken by the first speaker; When the second identification method is selected in the first operation, the electronic device determines that the first speech content is spoken by a speaker, and the electronic device displays the first speech content and a second identifier, the second identifier being used to identify the speaker of the first speech content.
2. The method of claim 1, wherein, After the electronic device acquires the first audio segment, the method further includes: The electronic device divides the first audio segment into multiple sub-segments, each of which includes the first sub-segment, and the semantics of the multiple sub-segments are different.
3. The method according to claim 1 or 2, characterized in that, The first audio segment includes a second sub-segment, the second sub-segment includes a second speaker's second speech content, the electronic device does not store the second speaker's second voiceprint feature, and the method further includes: When the first identification method is selected in the first operation, the electronic device determines that the second speech content is spoken by a speaker, and the electronic device displays the second speech content and a third identifier, the third identifier being used to identify the speaker of the second speech content.
4. The method according to claim 1 or 2, characterized in that, The first audio segment includes a second sub-segment, the second sub-segment includes a second speaker's second speech content, the electronic device stores a second speaker's second voiceprint feature and a fourth identifier, and the method further includes: When the first identification method is selected in the first operation, the electronic device determines that the voiceprint feature of the second sub-segment matches the second voiceprint feature, and the electronic device displays the second speaking content and the fourth identifier, the fourth identifier indicating that the second speaking content was spoken by the second speaker.
5. The method according to claim 4, characterized in that, After the electronic device displays the second spoken content and the fourth identifier, the method further includes: The electronic device receives a second operation, the second operation being used to filter the first speaker; In response to the second operation, the electronic device hides the second spoken content and the fourth identifier.
6. The method according to claim 1 or 2, characterized in that, The first audio segment further includes a third sub-segment, which includes third spoken content, the speaker of which is different from the speaker of the first spoken content. The method further includes: When the second identification method is selected in the first operation, the electronic device determines that the third speech content is spoken by a speaker, and the electronic device displays the third speech content and a fifth identifier, the fifth identifier being used to indicate that the speaker of the third speech content is different from the speaker of the first speech content.
7. The method according to claim 1 or 2, characterized in that, The first audio segment further includes a fourth sub-segment, which includes a fourth spoken content, and the method further includes: When the second identification method is selected in the first operation, the electronic device determines that the fourth speech content is spoken by a speaker, and the speaker of the fourth speech content is the same as the speaker of the first speech content. The electronic device displays the fourth speech content and displays the second identifier again at a first position associated with the fourth speech content. The second identifier displayed at the first position is used to indicate the speaker of the fourth speech content.
8. The method according to claim 1 or 2, characterized in that, Before the electronic device receives the first operation of the user selecting a voice recognition method, the method further includes: The electronic device receives a second audio segment input by the user and the first identifier, wherein the speaker of the second audio segment is the first speaker; The electronic device extracts the first voiceprint feature of the first speaker from the second audio segment; The electronic device stores the first voiceprint feature and the first identifier, wherein the first voiceprint feature and the first identifier are stored together.
9. The method according to claim 1 or 2, characterized in that, The method further includes: The electronic device captures a second audio segment through a microphone while displaying the first spoken content; The electronic device identifies the speech content and speaker of the second audio segment based on the speech recognition method selected in the first operation.
10. An electronic device, characterized in that, The electronic device includes: The storage module is used to store the first identifier and first voiceprint features of the first speaker; The user interaction module is used to receive a first operation from the user to select a voice recognition method, wherein the voice recognition method includes a first recognition method and a second recognition method. An audio acquisition module is used to acquire a first audio segment, the first audio segment containing a first sub-segment corresponding to the first spoken content; The voiceprint recognition module is used to determine that the voiceprint features of the first sub-segment match the first voiceprint features when the first recognition method is selected in the first operation. The display module is used to display the first spoken content and the first identifier when the first recognition method is selected in the first operation, wherein the first identifier indicates that the first spoken content was spoken by the first speaker; The clustering recognition module is used to determine, based on the acoustic features of the first audio segment, that the first speech content was spoken by a speaker when the second recognition method is selected in the first operation; The display module is further configured to display the first spoken content and the second identifier when the second recognition method is selected in the first operation, wherein the second identifier is used to identify the speaker of the first spoken content.
11. The electronic device according to claim 10, characterized in that, The electronic device further includes a content recognition module and a speech segmentation module; the content recognition module is used to recognize the spoken content contained in the first audio segment; the speech segmentation module is used to divide the first audio segment into multiple sub-segments based on the spoken content contained in the first audio segment, wherein the multiple sub-segments include the first sub-segment, and the semantics of the multiple sub-segments are different.
12. The electronic device according to claim 10 or 11, characterized in that, The first audio segment includes a second sub-segment, the second sub-segment includes a second speaker's second speech content, and the storage module does not store the second speaker's second voiceprint features; When the second identification method is selected in the first operation, the clustering identification module is further used to determine that the second speaking content is spoken by a speaker; the display module is further used to display the second speaking content and a third identifier, the third identifier being used to identify the speaker of the second speaking content.
13. The electronic device according to claim 10 or 11, characterized in that, The first audio segment includes a second sub-segment, the second sub-segment includes a second speaker's second speech content, and the storage module stores the second speaker's second voiceprint feature and a fourth identifier; When the first identification method is selected in the first operation, the voiceprint recognition module is further configured to determine that the voiceprint features of the second sub-segment match the second voiceprint features; the display module is further configured to display the second speaking content and the fourth identifier, the fourth identifier indicating that the second speaking content was spoken by the second speaker.
14. The electronic device according to claim 10 or 11, characterized in that, The first audio segment also includes a third sub-segment, which includes a third speech content, and the speaker of the third speech content is different from the speaker of the first speech content; When the second identification method is selected in the first operation, the clustering identification module is further used to determine that the third speaking content is spoken by a speaker; the display module is further used to display the third speaking content and a fifth identifier, the fifth identifier being used to indicate that the speaker of the third speaking content is different from the speaker of the first speaking content.
15. The electronic device according to claim 10 or 11, characterized in that, The first audio segment also includes a fourth sub-segment, which includes a fourth spoken content; When the second identification method is selected in the first operation, the clustering identification module is further configured to determine that the fourth speaking content is spoken by a speaker, and the speaker of the fourth speaking content is the same as the speaker of the first speaking content; the display module is further configured to display the fourth speaking content, and display the second identifier again at a first position associated with the fourth speaking content, the second identifier displayed at the first position being used to indicate the speaker of the fourth speaking content.
16. The electronic device according to claim 10 or 11, characterized in that, The electronic device further includes a voiceprint registration module, which is used to receive a second audio segment input by the user and the first identifier, wherein the speaker of the second audio segment is the first speaker; The voiceprint registration module is also used to send the second audio segment to the voiceprint recognition module; The voiceprint recognition module is also used to extract the first voiceprint feature of the first speaker from the second audio segment and send the first voiceprint feature to the voiceprint registration module; The voiceprint registration module is also used to receive the first voiceprint feature and send the first voiceprint feature and the first identifier to the storage module; The storage module is also used to store the first voiceprint feature and the first identifier, wherein the first voiceprint feature and the first identifier are stored together.
17. An electronic device, characterized in that, The electronic device includes: a display screen, a memory, and a processor coupled to the memory; the display screen is used to display a user interface, the memory stores a computer program, and the processor executes the computer program to cause the electronic device to perform the method as described in any one of claims 1 to 9.
18. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed on the electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Interface circuit of mobile microgrid system
CN107947168A
Speech translation method, device and equipment and translation machine
CN112309370A