Speaker separation method and device for FunASR monaural audio, equipment and medium
By modifying the FunASR model and combining it with a speaker separation module, parallel processing of mono audio was achieved, solving the problem that FunASR cannot process in parallel in existing technologies. This enabled parallel processing of audio transcription and speaker separation, improving processing efficiency and stability.
Patent Information
- Application Number
- CN202511068600.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies using FunASR for speaker separation in mono audio have low processing efficiency and cannot achieve parallel processing of audio transcription and speaker separation.
By modifying the FunASR model, removing the punctuation prediction model and the text inverse normalization model, and combining it with the speaker separation module, the speech activity detection model and the speaker separation model are used for parallel processing to ensure the consistency of timestamps.
It achieves parallel processing of audio transcription text and speaker separation, which improves processing efficiency, ensures the consistency of timestamps, and enhances the overall stability and efficiency of processing.
Smart Images

Figure CN120977330A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to a speaker separation method and device for FunASR monaural audio, equipment and medium. BACKGROUND
[0002] FunASR is an open-source automatic speech recognition (ASR) toolkit implemented based on PyTorch, supporting various speech recognition models such as Squeezeformer, Conformer, etc. It is mainly used to convert speech signals into text, and is suitable for scenarios such as voice assistants, voice input, voice search, etc. Monaural audio refers to all audio signals transmitted through a single channel without directional information (such as the difference in left and right positions of the sound source, the difference in phase, etc.). In this scenario, it refers to the audio signals of multiple speakers transmitted through a single channel. When using FunASR, monaural audio is a very important input requirement, and most ASR models only support monaural audio input. Speaker separation, also known as speaker recognition, refers to a technology that separates the voices of different speakers according to the speaker's identity.
[0003] According to the business needs, the monaural audio of the agent is converted into text by FunASR, but FunASR cannot directly perform speaker separation, so the speaker separation spkrec-ecapa-voxceleb model is used to obtain the speaker separation result. The prior art can only be processed in series, and the efficiency is low. SUMMARY
[0004] The present disclosure aims to at least solve one of the technical problems in the related art to some extent.
[0005] To this end, the purpose of the present disclosure is to propose a speaker separation method and device for FunASR monaural audio, computer equipment and storage medium, so as to ensure the consistency of the timestamps output by the modified FunASR model and the speaker separation module, realize parallel processing of audio transcription text and speaker separation, and effectively improve the processing efficiency.
[0006] To achieve the above purpose, the first aspect of the present disclosure proposes a speaker separation method for FunASR monaural audio, comprising:
[0007] inputting the monaural audio into the modified FunASR model to obtain at least one audio character and a character timestamp corresponding to each audio character, wherein the modified FunASR model does not configure a punctuation prediction model and a text inverse normalization model;
[0008] The mono audio is input into the speaker separation module to obtain at least one speaker tag and a speaker timestamp corresponding to each speaker tag;
[0009] Matching is performed based on the character timestamp and the speaker timestamp to obtain a speaker separation result, wherein the speaker separation result is used to indicate the speaker timestamp and the audio character corresponding to each speaker tag in the mono audio.
[0010] To achieve the above objectives, the speaker separation device for FunASR mono audio provided in the second aspect of this disclosure includes:
[0011] The first processing module is used to input mono audio into the modified FunASR model to obtain at least one audio character and a character timestamp corresponding to each audio character, wherein the modified FunASR model is not configured with a punctuation prediction model and a text inverse normalization model.
[0012] The second processing module is used to input the mono audio into the speaker separation module to obtain at least one speaker tag and a speaker timestamp corresponding to each speaker tag;
[0013] The third processing module is used to perform matching processing based on the character timestamp and the speaker timestamp to obtain a speaker separation result, wherein the speaker separation result is used to indicate the speaker timestamp and the audio character corresponding to each speaker tag in the mono audio.
[0014] The computer device proposed in the third aspect of this disclosure includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the speaker separation method for FunASR mono audio as proposed in the first aspect of this disclosure.
[0015] A fourth aspect of this disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the speaker separation method for FunASR mono audio as proposed in a first aspect of this disclosure.
[0016] A fifth aspect of this disclosure provides a computer program product that, when executed by a processor, performs a speaker separation method for FunASR mono audio as described in a first aspect of this disclosure.
[0017] This disclosure provides a speaker separation method, apparatus, computer device, and storage medium for FunASR mono audio. The method involves inputting mono audio into a modified FunASR model to obtain at least one audio character and a corresponding character timestamp. The modified FunASR model does not include a punctuation prediction model or a text inverse normalization model. The mono audio is then input into a speaker separation module to obtain at least one speaker tag and a corresponding speaker timestamp. Matching is performed based on the character timestamps and speaker timestamps to obtain the speaker separation result. This result indicates the speaker timestamp and audio character corresponding to each speaker tag in the mono audio. This ensures consistency between the timestamps output by the modified FunASR model and the speaker separation module, enabling parallel processing of audio-to-text transcription and speaker separation, thereby effectively improving processing efficiency.
[0018] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description
[0019] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which:
[0020] Figure 1 This is a schematic flowchart of a speaker separation method for FunASR mono audio proposed in an embodiment of this disclosure;
[0021] Figure 2 This is a flowchart of a method for speaker separation in mono audio for FunASR based on the present disclosure;
[0022] Figure 3 This is a schematic diagram of the speaker separation device for FunASR mono audio according to an embodiment of the present disclosure;
[0023] Figure 4 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0024] Embodiments of this disclosure are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are used only to explain this disclosure, and should not be construed as limiting this disclosure. Rather, embodiments of this disclosure include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.
[0025] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0026] Figure 1 This is a schematic flowchart of a speaker separation method for FunASR mono audio proposed in one embodiment of this disclosure.
[0027] It should be noted that the speaker separation method for FunASR mono audio in this embodiment is implemented by a speaker separation device for FunASR mono audio. This device can be implemented by software and / or hardware. The device can be configured in a computer device, which may include, but is not limited to, a terminal, a server, etc. For example, the terminal may be a mobile phone, a PDA, etc.
[0028] like Figure 1 As shown, the speaker separation method for FunASR mono audio includes:
[0029] S101: Input mono audio into the modified FunASR model to obtain at least one audio character and a character timestamp corresponding to each audio character. The modified FunASR model does not have a punctuation prediction model or a text inverse normalization model configured.
[0030] The modified FunASR model can refer to the FunASR model obtained by removing the punctuation prediction model and text inverse normalization model from the conventional FunASR model.
[0031] Audio characters can refer to text characters obtained by transcribing mono audio.
[0032] Among them, the character timestamp can be used to indicate the time point when the corresponding audio character appears in the mono audio.
[0033] Among them, the punctuation prediction model is mainly used for punctuation recovery tasks. This model is specifically designed to add appropriate punctuation marks to text without punctuation in order to improve the readability of the text and the accuracy of subsequent processing.
[0034] Among them, the text inverse normalization model is used to implement the inverse text normalization process based on finite state converters. It transforms text through a series of predefined states and transformation rules, thereby achieving a mapping from one form to another, such as converting standardized expressions such as numbers, dates, and currencies into more colloquial or specific formats.
[0035] It is understandable that the punctuation prediction model and text denormalization model in the conventional FunASR model may modify the timestamp information in the transcription results, resulting in inconsistencies between the output results and the timestamps in the speaker separation module, affecting the accuracy of subsequent matching processes. Therefore, in this embodiment of the disclosure, the punctuation prediction model and text denormalization model in the conventional FunASR model can be removed before being used in the technical solution of this disclosure.
[0036] S102: Input mono audio into the speaker separation module to obtain at least one speaker tag and a speaker timestamp corresponding to each speaker tag.
[0037] The speaker separation module can refer to the module used in this embodiment to distinguish different speakers in mono audio.
[0038] Speaker tags are used to identify different speakers in mono audio.
[0039] Among them, the speaker timestamp can be used to refer to the time point of the speaker's speech activity in mono audio as indicated by the corresponding speaker tag.
[0040] In other words, in this embodiment of the present disclosure, a speaker separation module can be pre-configured to synchronously realize the speaker separation process for mono audio.
[0041] S103: Perform matching processing based on character timestamps and speaker timestamps to obtain speaker separation results, wherein the speaker separation results are used to indicate the speaker timestamp and audio character corresponding to each speaker tag in the mono audio.
[0042] For example, in this embodiment of the present disclosure, when performing matching processing based on character timestamps and speaker timestamps to obtain speaker separation results, the inclusion relationship between character timestamps and speaker timestamps may be determined, and then the audio characters corresponding to the character timestamps included in the speaker timestamp range may be associated with speaker tags.
[0043] In this embodiment, mono audio is input into the modified FunASR model to obtain at least one audio character and a corresponding character timestamp. The modified FunASR model does not include a punctuation prediction model or a text inverse normalization model. The mono audio is then input into a speaker separation module to obtain at least one speaker tag and a corresponding speaker timestamp. Matching is performed based on the character timestamps and speaker timestamps to obtain the speaker separation result. This result indicates the speaker timestamp and audio character corresponding to each speaker tag in the mono audio. This ensures consistency between the timestamps output by the modified FunASR model and the speaker separation module, enabling parallel processing of audio-to-text transcription and speaker separation, thereby effectively improving processing efficiency.
[0044] Optionally, in some embodiments, the speaker separation module includes a speech activity detection model, which is consistent with the speech activity detection model configured in the modified FunASR model. The speech activity detection model in the speaker separation module is used to detect active speech in mono audio and output a timestamp of the active speech. This ensures that the timestamp output by the speaker separation module is consistent with the timestamp output by the modified FunASR model.
[0045] In other words, in this embodiment of the present disclosure, the speech activity detection model configured in the speaker separation module should be consistent with the speech activity detection model in the modified FunASR model, so as to ensure the consistency of the output timestamps of the two.
[0046] Optionally, in some embodiments, the speaker separation module further includes a speaker separation model and a clustering module, wherein the speaker separation model is used to extract features from the active speech; and the clustering module is used to cluster the active speech according to the feature extraction results to determine at least one speaker label and a speaker timestamp corresponding to each speaker label.
[0047] For example, in the embodiments of this disclosure, when extracting features from active speech, feature extraction can be performed based on the speaker's voiceprint, speech rate, and other sound information.
[0048] Optionally, in some embodiments, the modified FunASR model and speaker separation module are scheduled in parallel based on the Celery distributed task queue framework. Thus, while the mono audio is being transcribed using FunASR, speaker separation is also being performed, allowing the entire process to be processed in parallel, resulting in high efficiency and stability.
[0049] Optionally, in some embodiments, after obtaining the speaker separation result, punctuation marks in the audio text from the speaker separation result can be recovered based on a preset punctuation prediction model, wherein the audio text consists of multiple audio characters. This can improve the readability of the text and the accuracy of subsequent processing.
[0050] Optionally, in some embodiments, after recovering the punctuation in the audio text from the speaker separation result based on a preset punctuation prediction model, the audio text can also be denormalized based on a preset text denormalization model. This can effectively improve the normalization of the audio text.
[0051] Based on the above embodiments, this disclosure provides a method for speaker separation in FunASR mono audio, which can embed speaker information into the transcribed text to meet business needs. The entire process is parallelized, stable, and efficient. The program can be developed using Python, involving the Celery distributed task queue framework, a recompiled FunASR model, a speech activity detection model, a speaker separation model, a punctuation prediction model, and a text inverse normalization model. Figure 2 As shown, Figure 2 This is a flowchart of a method for speaker separation in mono audio for FunASR based on the present disclosure.
[0052] The details are as follows:
[0053] 1. Celery distributed task queue framework:
[0054] The mono audio is scheduled by Celery, calling the recompiled FunASR model and the speaker separation scheme respectively. In terms of scheduling, since calling the FunASR model via API is a non-CPU-intensive operation, mainly waiting for the FunASR return result, it consumes relatively few CPU resources. Therefore, when setting up workers (each worker can only call one CPU core), a high-concurrency scheduling scheme is adopted. The speaker separation scheme, on the other hand, is a CPU-intensive operation, so when setting up workers, a low-concurrency scheme with multiple workers is adopted, allowing each core to fully handle the appropriate number of recordings.
[0055] 2. Speaker separation scheme:
[0056] 2.1 Voice Activity Detection Model:
[0057] The purpose of the speech activity detection model is to detect the audio of the activity (i.e., the audio portion of the speaker's speech). In order to align with FunASR, this scheme adopts the speech activity detection model speech_paraformer_vad_zh-cn-16kcommon-onnx, which is consistent with FunASR.
[0058] Model input: audio.
[0059] Model processing: Detecting active speech in audio.
[0060] Model output: timestamps of active speech, in a format such as (in milliseconds): [[100,300],[300,500]].
[0061] 2.2 Speaker Separation Model:
[0062] The speaker separation model used in this scheme is the spkrec-ecapa-voxceleb model, which performs well under CPU-based schemes, with high separation efficiency and good results.
[0063] Model input: audio.
[0064] Model processing: Feature extraction is performed on the input audio, mainly based on the speaker's voiceprint, speech rate and other audio information.
[0065] Model output: Results of feature extraction.
[0066] 2.3 K-means clustering algorithm:
[0067] K-means clustering is a classic unsupervised learning method, mainly used to divide a dataset into K clusters, such that data points within the same cluster have high similarity, while data points between different clusters have low similarity.
[0068] Input: Identification and timestamps of audio from multiple activities.
[0069] Model processing: Calculate the similarity between features of each active audio and perform clustering (2 clusters, corresponding to 2 speakers).
[0070] Output: Timestamps of the active audio for the two speakers.
[0071] 2.4 Speaker separation process in the overall workflow:
[0072] For the input audio file, the speech activity detection model obtains the timestamps of multiple active speech segments. The audio is then segmented according to the timestamps to obtain multiple active audio segments. Each of the multiple audio segments is then subjected to a speaker separation model to obtain the features of the multiple active audio segments. Finally, the k-means clustering algorithm is used to obtain the activity timestamps of the two speakers.
[0073] The format is: timestamp + speaker information, format: [{"timestamp":"[100,200]","speaker":"0"},{"timestamp":"[250,350]","speaker":"1"}]
[0074] 3. Recompiled FunASR
[0075] 3.1 Original FunASR
[0076] FunASR is a highly integrated speech recognition toolkit that includes speech activity detection models, speech recognition models, punctuation prediction models, text inverse normalization models, hot word enhancement models, and more.
[0077] Model input: audio.
[0078] Model processing: The input audio is format converted, and the speech activity detection model, speech recognition model, punctuation prediction model, and text denormalization model are called in a pipeline manner.
[0079] Model output: Transcribed text and timestamp, formatted as follows (unit: milliseconds, only important content is displayed): "stamp_sents":[{"text":"Hello. XXXX","timestamp":"[[380,980],[1340,1500]"}).
[0080] 3.2. Recompiled FunASR after modification
[0081] Because the original FunASR model went through a speech activity detection model, a punctuation prediction model, and a text denormalization model, the output timestamps were inconsistent with the results obtained by the speaker separation scheme. The main reason is that FunASR transcribes the active speech obtained by the speech activity detection model word by word. At this point, the timestamps are roughly consistent with those of the speaker separation scheme. However, after going through the punctuation prediction model and the text denormalization model, FunASR will modify the timestamps again, making the timestamps inconsistent and unable to embed speaker information. Therefore, this proposal modifies the source code of FunASR, removes unnecessary models (punctuation prediction model and text denormalization model), and recompiles it.
[0082] Input to the modified model: audio.
[0083] The modified model processes the input audio in a pipeline manner, including converting the audio format and calling the speech activity detection model and speech recognition model. It does not include the punctuation prediction model or the text denormalization model.
[0084] Model output: Transcribed text and timestamp, in a format such as (unit: milliseconds; since part of the model has been removed, the output is now a single character and its timestamp without punctuation): "stamp_sents":[{"text":"Hello XXXX","timestamp":"[[380,480],[500,680],…,[1340,1500]"}).
[0085] 4. Technical processing
[0086] The recompiled FunASR outputs word-by-word text and timestamps, while the speaker separation scheme provides timestamps of an active audio segment and speaker information. Analysis reveals that the timestamp range obtained by the speaker separation scheme is larger than the timestamps of a single character obtained by FunASR (i.e., inclusion relationship). This proposal extracts and integrates the FunASR results based on the timestamps obtained by the speaker separation scheme, ultimately yielding the timestamps of an active audio segment and the corresponding text and speaker information.
[0087] 5. Punctuation Prediction Model
[0088] The punctuation prediction model `punc_ct-transformer_cn-en-common` is primarily used for punctuation recovery tasks. This model is specifically designed to add appropriate punctuation marks to text lacking them, thereby improving readability and the accuracy of subsequent processing. Particularly for plain text outputs from Automatic Speech Recognition (ASR), this model can significantly improve their structure.
[0089] The text, after technical processing, lacks punctuation. Therefore, this proposal uses the original FunASR punctuation prediction model for punctuation recovery.
[0090] 6. Text Inverse Normalization Model
[0091] This proposal utilizes the fst_itn text inverse normalization model, primarily for implementing an inverse text normalization process based on finite state converters. It transforms text through a series of predefined states and transformation rules, thereby achieving a mapping from one form to another, such as converting standardized expressions like numbers, dates, and currencies into more colloquial or specific formats.
[0092] In addition to punctuation restoration, the text also needs to be denormalized to make the entire text more standardized. Therefore, the original FunASR text denormalization model was called to obtain text data that meets the business scenario. This data includes text, timestamps, and speaker information.
[0093] Figure 3 This is a schematic diagram of the speaker separation device for FunASR mono audio according to an embodiment of the present disclosure.
[0094] like Figure 3 As shown, the speaker separation device 30 for FunASR mono audio includes:
[0095] The first processing module 301 is used to input mono audio into the modified FunASR model to obtain at least one audio character and a character timestamp corresponding to each audio character. The modified FunASR model does not have a punctuation prediction model and a text inverse normalization model configured.
[0096] The second processing module 302 is used to input mono audio into the speaker separation module to obtain at least one speaker tag and a speaker timestamp corresponding to each speaker tag;
[0097] The third processing module 303 is used to perform matching processing based on the character timestamp and the speaker timestamp to obtain the speaker separation result, wherein the speaker separation result is used to indicate the speaker timestamp and audio character corresponding to each speaker tag in the mono audio.
[0098] It should be noted that the foregoing explanation of the speaker separation method for FunASR mono audio also applies to the speaker separation device for FunASR mono audio in this embodiment, and will not be repeated here.
[0099] In this embodiment, mono audio is input into the modified FunASR model to obtain at least one audio character and a corresponding character timestamp. The modified FunASR model does not include a punctuation prediction model or a text inverse normalization model. The mono audio is then input into a speaker separation module to obtain at least one speaker tag and a corresponding speaker timestamp. Matching is performed based on the character timestamps and speaker timestamps to obtain the speaker separation result. This result indicates the speaker timestamp and audio character corresponding to each speaker tag in the mono audio. This ensures consistency between the timestamps output by the modified FunASR model and the speaker separation module, enabling parallel processing of audio-to-text transcription and speaker separation, thereby effectively improving processing efficiency.
[0100] Figure 4 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Figure 4 The computer device 12 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0101] like Figure 4 As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0102] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0103] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0104] Memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 4 Not shown; usually referred to as a "hard drive".
[0105] although Figure 4Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a Compact Disc Read-Only Memory (CD-ROM), a Digital Video Disc Read-Only Memory (DVD-ROM), or other optical media). In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.
[0106] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this disclosure.
[0107] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable human interaction with the computer device 12, and / or with any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0108] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the speaker separation method for FunASR mono audio mentioned in the foregoing embodiments.
[0109] To implement the above embodiments, this disclosure also proposes a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the speaker separation method for FunASR mono audio as proposed in the foregoing embodiments of this disclosure.
[0110] To implement the above embodiments, this disclosure also proposes a computer program product that, when executed by an instruction processor, performs a speaker separation method for FunASR mono audio as proposed in the foregoing embodiments of this disclosure.
[0111] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this disclosure all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0112] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0113] This disclosure is intended to provide implementation schemes for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.
[0114] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0115] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0116] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.
[0117] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0118] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0119] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0120] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0121] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A speaker separation method for FunASR mono audio, characterized in that, include: Mono audio is input into the modified FunASR model to obtain at least one audio character and a character timestamp corresponding to each audio character. The modified FunASR model does not have a punctuation prediction model or a text inverse normalization model configured. The mono audio is input into the speaker separation module to obtain at least one speaker tag and a speaker timestamp corresponding to each speaker tag; Matching is performed based on the character timestamp and the speaker timestamp to obtain a speaker separation result, wherein the speaker separation result is used to indicate the speaker timestamp and the audio character corresponding to each speaker tag in the mono audio.
2. The method as described in claim 1, characterized in that, The speaker separation module includes a speech activity detection model, and the speech activity detection model in the speaker separation module is consistent with the speech activity detection model configured in the modified FunASR model; wherein, The speech activity detection model in the speaker separation module is used to detect active speech in the mono audio and output the timestamp of the active speech.
3. The method as described in claim 2, characterized in that, The speaker separation module further includes a speaker separation model and a clustering module, wherein, The speaker separation model is used for feature extraction of the active speech; The clustering module is used to perform clustering processing on the active speech based on the feature extraction results, so as to determine at least one speaker label and a speaker timestamp corresponding to each speaker label.
4. The method as described in claim 1, characterized in that, The modified FunASR model and the speaker separation module are scheduled in parallel based on the Celery distributed task queue framework.
5. The method as described in claim 1, characterized in that, The method further includes: Punctuation marks in the audio text of the speaker separation result are recovered based on a preset punctuation prediction model, wherein the audio text is composed of multiple audio characters.
6. The method as described in claim 5, characterized in that, The method further includes: The audio text is denormalized based on a preset text denormalization model.
7. A speaker separation device for FunASR mono audio, characterized in that, include: The first processing module is used to input mono audio into the modified FunASR model to obtain at least one audio character and a character timestamp corresponding to each audio character, wherein the modified FunASR model is not configured with a punctuation prediction model and a text inverse normalization model. The second processing module is used to input the mono audio into the speaker separation module to obtain at least one speaker tag and a speaker timestamp corresponding to each speaker tag; The third processing module is used to perform matching processing based on the character timestamp and the speaker timestamp to obtain a speaker separation result, wherein the speaker separation result is used to indicate the speaker timestamp and the audio character corresponding to each speaker tag in the mono audio.
8. A computer device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, in, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-6.