Voice data processing method and device, electronic equipment and storage medium
By extracting host keywords and segmenting and separating the voice data, the problem of voice overlap in the voice data is solved, and the quality and accuracy of the voice data are improved.
Patent Information
- Application Number
- CN202510323856.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
AI Technical Summary
Voice overlap often occurs in voice data, resulting in poor quality of voice data.
By obtaining the original mixed speech, text conversion process is performed to extract the host keywords, and the speech is segmented and separated based on these keywords. The speech of each speech object is further separated through speech characteristics classification and splicing.
It effectively reduces the phenomenon of voice overlap in voice data, improves the quality of voice data, and enables the voice of each speaker to be recognized and processed more accurately.
Smart Images

Figure CN120183395A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and is applicable to the fields of fintech and healthcare. In particular, it relates to a method, apparatus, electronic device, and storage medium for processing voice data. Background Art
[0002] A voice dataset is a dataset used to train and test voice processing models. The quality of the voice dataset directly affects the performance of the voice processing model. In the field of natural language processing, it is often necessary to collect high-quality voice data to construct a voice dataset for training a voice processing model. In the field of fintech, for example, in the conversation business of the insurance scenario, it usually involves multi-party interactions among speakers such as a host, an insurance service provider, and a user. Under the management of the host, the service provider and the user conduct exchanges related to insurance products. During this exchange process, conversation voices can be collected, and then an insurance conversation voice dataset can be constructed. In the field of healthcare, for example, in the scenario of a medical academic conference, it usually involves multi-party interactions among speakers such as a conference host and multiple medical experts. The medical experts can conduct medical academic discussions under the coordination of the host. During the conference, conference voices can be collected, and then a medical conference voice dataset can be constructed.
[0003] However, since the voice data is collected in the presence of multiple speakers, multiple speakers may speak simultaneously, so there is often a significant voice overlap phenomenon in the voice data, which makes the quality of the voice data poor.
[0004] Therefore, how to improve the quality of voice data has become a technical problem to be solved urgently. Summary of the Invention
[0005] The main purpose of the embodiments of this application is to propose a method, apparatus, electronic device, and storage medium for processing voice data, aiming to reduce the voice overlap phenomenon in voice data, thereby improving the quality of voice data.
[0006] To achieve the above object, a first aspect of the embodiments of this application proposes a method for processing voice data, the method includes:
[0007] Obtain original mixed voice; wherein, the original mixed voice is the voice collected from a host object and at least two target speaker objects;
[0008] Perform text conversion processing on the original mixed voice to obtain a transcribed text;
[0009] Extract the host keyword from the transcribed text according to the preset candidate host keywords to obtain the target host keyword; wherein, the target host keyword is used to represent the keyword for the host object to indicate the start or end of the speech of the target speaker.
[0010] Segment the original mixed speech according to the time points of the target host keyword in the original mixed speech to obtain at least two mixed speech segments.
[0011] Separate the speech of each mixed speech segment according to the speaker to obtain single-speaker speech sub-segments.
[0012] Classify the at least two single-speaker speech sub-segments according to the speech features of at least two target speakers to obtain at least two target single-speaker speech classification results; wherein, each target single-speaker speech classification result includes the single-speaker speech sub-segments of the same target speaker.
[0013] Stitch the single-speaker speech sub-segments of each target single-speaker speech classification result to obtain the target single-speaker speech of each target speaker.
[0014] In some embodiments, the segmenting the original mixed speech according to the time points of the target host keyword in the original mixed speech to obtain at least two mixed speech segments includes:
[0015] Perform time division on the original mixed speech according to the time points of the target host keyword in the original mixed speech to obtain at least two initial time ranges.
[0016] Update the start time or end time of the initial time range according to the start time point or end time point of the speech of any one of the target speakers to obtain the target time range.
[0017] Intercept the original mixed speech according to the target time range to obtain mixed speech segments.
[0018] In some embodiments, before the updating the start time or end time of the initial time range according to the start time point or end time point of the speech of any one of the target speakers to obtain the target time range, the method further includes:
[0019] Extract the end speech keyword from the original mixed speech according to the preset candidate end speech keywords to obtain the target end speech keyword.
[0020] Determine the time point of the target end-of-speech keyword in the original mixed speech as the end-of-speech time point.
[0021] In some embodiments, at least two of the target host keywords include a host start keyword and a host end keyword;
[0022] The time division of the original mixed speech according to the time points of the target host keywords in the original mixed speech to obtain at least two initial time ranges includes:
[0023] Determine the time range between the time point of the host end keyword in the original mixed speech and the time point of the host start keyword in the original mixed speech as the host time range;
[0024] Intercept the time range outside the host time range from the original mixed speech to obtain at least two of the initial time ranges.
[0025] In some embodiments, at least two of the target speakers include a first speaker and a second speaker;
[0026] The classification of the single-speaker voice sub-fragments of at least two target speakers according to the voice characteristics of at least two target speakers to obtain at least two target single-speaker voice classification results includes:
[0027] Obtain the voice characteristics of the first speaker to obtain the first voice characteristics;
[0028] Obtain the voice characteristics of the second speaker to obtain the second voice characteristics;
[0029] According to the first voice characteristics and the second voice characteristics, perform model training on a pre-constructed initial clustering model to obtain a target clustering model;
[0030] Perform clustering processing on the single-speaker voice sub-fragments of at least two target speakers through the target clustering model to obtain a first single-speaker voice classification result and a second single-speaker voice classification result; wherein, the first single-speaker voice classification result includes the single-speaker voice sub-fragments of the first speaker; the second single-speaker voice classification result includes the single-speaker voice sub-fragments of the second speaker.
[0031] In some embodiments, before splicing the single-speaker voice sub-fragments of each target single-speaker voice classification result to obtain the target single-speaker voice of each target speaker, the method further includes:
[0032] Perform time-frequency conversion on each of the single-speaker object voice sub-fragments to obtain voice spectrum data;
[0033] Extract low-frequency components from the voice spectrum data to obtain fundamental frequency data;
[0034] Perform energy peak analysis on the fundamental frequency data to obtain the fundamental frequency energy peak;
[0035] If the fundamental frequency energy peak is greater than or equal to a preset energy peak threshold, determine the voice segment corresponding to the fundamental frequency energy peak in the single-speaker object voice sub-fragment as the overlapping voice segment;
[0036] Perform voice deletion processing on the overlapping voice segment of each of the single-speaker object voice sub-fragments to obtain the remaining voice segment, and determine the remaining voice segment as the single-speaker object voice sub-fragment.
[0037] In some embodiments, the performing time-frequency conversion on each of the single-speaker object voice sub-fragments to obtain voice spectrum data includes:
[0038] Select the starting voice segment for each of the single-speaker object voice sub-fragments according to the starting time of each of the single-speaker object voice sub-fragments and a preset overlapping segment duration threshold to obtain the starting voice segment;
[0039] Select the ending voice segment for each of the single-speaker object voice sub-fragments according to the ending time of each of the single-speaker object voice sub-fragments and the overlapping segment duration threshold to obtain the ending voice segment;
[0040] Perform first time-frequency conversion processing on the starting voice segment to obtain the starting voice spectrum;
[0041] Perform second time-frequency conversion processing on the ending voice segment to obtain the ending voice spectrum;
[0042] Determine the starting voice spectrum and the ending voice spectrum as the voice spectrum data.
[0043] To achieve the above object, a second aspect of the embodiments of the present application proposes a voice data processing device, and the device includes:
[0044] A voice acquisition module, configured to acquire original mixed voice; wherein, the original mixed voice is the voice collected from a host object and at least two target speaker objects;
[0045] A text conversion module, configured to perform text conversion processing on the original mixed voice to obtain a transcription text;
[0046] A keyword extraction module, configured to extract a host keyword from the transcribed text according to a preset candidate host keyword, so as to obtain a target host keyword; wherein, the target host keyword is used to represent a keyword for the host object to indicate the start or end of the speech of the target speaker.
[0047] A voice segmentation module, configured to segment the original mixed voice according to the time point of the target host keyword in the original mixed voice, so as to obtain at least two mixed voice segments.
[0048] A voice separation module, configured to separate the voice of each mixed voice segment according to the speaker, so as to obtain a single-speaker voice sub-segment.
[0049] A voice classification module, configured to classify the single-speaker voice sub-segments of at least two target speakers according to the voice features of at least two target speakers, so as to obtain at least two target single-speaker voice classification results; wherein, each target single-speaker voice classification result includes the single-speaker voice sub-segments of the same target speaker.
[0050] A voice splicing module, configured to splice the single-speaker voice sub-segments of each target single-speaker voice classification result, so as to obtain the target single-speaker voice of each target speaker.
[0051] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0052] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0053] The voice data processing method, device, electronic device, and storage medium proposed in this application perform voice data processing by combining voice segmentation and voice separation. First, considering the feature that the host object can indicate other objects to start or end speaking, the keywords of the host object (i.e., the target host keywords) are extracted from the transcribed text of the original mixed voice. According to the time points of the target host keywords in the original mixed voice, the original mixed voice is segmented into at least two first mixed voice segments. In this way, the original mixed voice with multiple speakers can be segmented into mixed voice segments that basically only contain the target speaking object, thereby preliminarily removing irrelevant audio (such as the speeches of others, etc.). Then, voice separation is performed on the mixed voice segments to obtain single-speaker voice sub-segments, so that the single-speaker voice sub-segments only contain a single speaking object, thereby further improving the quality of the voice. According to the voice features of at least two target speaking objects, at least two single-speaker voice sub-segments are classified according to the speaking object to obtain single-speaker voice sub-segments of the same target speaking object. Then, the single-speaker voice sub-segments of each same target speaking object are spliced to obtain the target single-speaker voice of each target speaking object, so that the original mixed voice of multiple speaking objects can be accurately separated into the voices of single speaking objects, improving the quality of voice data and being applicable to voice processing in complex scenarios of multi-speaker interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flowchart of the voice data processing method provided by an embodiment of this application;
[0055] Figure 2 is Figure 1 a flowchart of step 104 in
[0056] Figure 3 is Figure 2 a flowchart of step 201 in
[0057] Figure 4 is a flowchart of the voice data processing method provided by another embodiment of this application;
[0058] Figure 5 is Figure 1 a flowchart of step 106 in
[0059] Figure 6 is a flowchart of the voice data processing method provided by another embodiment of this application;
[0060] Figure 7 is Figure 6 a flowchart of step 601 in
[0061] Figure 8It is a schematic structural diagram of the voice data processing device provided by the embodiments of the present application;
[0062] Figure 9 It is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present application. Specific embodiments
[0063] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0064] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0066] First, several nouns involved in the present application are analyzed:
[0067] Artificial Intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can be a simulation of the information process of human consciousness and thinking. Artificial intelligence can also be a theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The present application can acquire and process relevant data based on artificial intelligence technology.
[0068] Speech signal processing: It is a technology used to study the speech production process, the statistical characteristics of speech signals, automatic speech recognition, machine synthesis, and speech perception.
[0069] Natural Language Processing (NLP): Natural language processing is a branch of artificial intelligence. It aims to enable computers to understand and process human language, achieve effective communication between humans and machines, and enable computers to perform tasks such as language translation, sentiment analysis, and text summarization.
[0070] The speech data processing method, device, electronic device, and storage medium provided by the embodiments of the present application will be specifically described through the following embodiments. First, the speech data processing method in the embodiments of the present application will be described.
[0071] The speech data processing method provided by the embodiments of the present application can be applied to terminals, can also be applied to the server side, or can be software running on the terminal or the server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms; the software can be an application that implements the speech data processing method, etc., but is not limited to the above forms.
[0072] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0073] It should be noted that in each specific implementation manner of this application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when this application embodiment needs to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of this application embodiment will be obtained.
[0074] Figure 1 is an optional flowchart of the voice data processing method provided by the embodiments of this application. Figure 1 The method in may include but is not limited to steps 101 to 107.
[0075] Step 101, obtain the original mixed voice; wherein, the original mixed voice is the voice collected from the host object and at least two target speaker objects.
[0076] Step 102, perform text conversion processing on the original mixed voice to obtain a transcribed text.
[0077] Step 103, extract the host keyword from the transcribed text according to the preset candidate host keywords to obtain the target host keyword; wherein, the target host keyword is used to represent the keyword for the host object to indicate the start or end of the speech of the target speaker object.
[0078] Step 104, perform voice segmentation on the original mixed voice according to the time point of the target host keyword in the original mixed voice to obtain at least two mixed voice segments.
[0079] Step 105, separate the voice of each mixed voice segment according to the speaker object to obtain the single-speaker object voice sub-segment.
[0080] Step 106, classify the at least two single-speaker object voice sub-segments according to the voice characteristics of the at least two target speaker objects to obtain at least two target single-speaker object voice classification results; wherein, each target single-speaker object voice classification result includes the single-speaker object voice sub-segments of the same target speaker object.
[0081] Step 107, splice the single-speaker object voice sub-segments of each target single-speaker object voice classification result to obtain the target single-speaker object voice of each target speaker object.
[0082] The beneficial effects of the embodiments of the present application include, but are not limited to: combining speech segmentation and speech separation to process speech data. First, considering the feature that the host object can indicate other objects to start or end speaking, the keywords of the host object (i.e., the target host keywords) are extracted from the transcribed text of the original mixed speech. According to the time points of the target host keywords in the original mixed speech, the original mixed speech is segmented into at least two first mixed speech segments, so that the original mixed speech with multiple speakers can be segmented into mixed speech segments that basically only contain the target speaking object, thereby preliminarily removing irrelevant audio (such as the speech of others, etc.). Then, speech separation is performed on the mixed speech segments to obtain single-speaker speech sub-segments, so that the single-speaker speech sub-segments only contain a single speaking object, thereby further improving the quality of the speech. According to the speech features of at least two target speaking objects, at least two single-speaker speech sub-segments are classified according to the speaking object to obtain single-speaker speech sub-segments of the same target speaking object. Then, the single-speaker speech sub-segments of each same target speaking object are spliced to obtain the target single-speaker speech of each target speaking object, so that the original mixed speech of multiple speakers can be accurately separated into the speech of a single speaker, improving the quality of the speech data and being applicable to speech processing in complex scenarios of multi-speaker interaction.
[0083] In step 101 of some embodiments, the original mixed speech is an audio collected in a multi-speaker scenario. Among them, the multi-speaker scenario refers to a scenario where multiple people are speaking, specifically a scenario where there is a host object and at least two target speaking objects (i.e., speakers) speaking.
[0084] For example, in a debate scenario, the multiple speakers may include a debate host, a positive debater, and a negative debater. Under the hosting of the debate host, the positive debater and the negative debater conduct a debate. By recording during the debate, the original mixed speech is obtained, and the original mixed speech contains the speech of the debate host, the positive debater, and the negative debater. It can be understood that since multiple speakers may speak simultaneously, the original mixed speech collected in a multi-speaker scenario is very likely to contain the mixed speech of different speakers, that is, the original mixed speech is prone to speech overlap. After obtaining the original mixed speech, the speech of different speakers in the original mixed speech can be separated from each other through subsequent speech data processing, thereby reducing the speech overlap phenomenon and improving the quality of the speech data. Further, a speech dataset can be constructed using the processed speech data for training a speech generation model, thereby further improving the diversity and accuracy of speech generation of the speech generation model.
[0085] For another example, in the fintech scenario, such as in the insurance conversation service, multiple speakers may include a conversation host, an insurance service provider, and a user. Under the management of the host, the insurance service provider and the user conduct exchanges related to insurance products. During this exchange process, the original mixed speech can be recorded. For another example, in the healthcare scenario, such as in the medical academic conference scenario, multiple speakers may include a conference host and multiple medical experts. The medical experts can conduct medical academic discussions under the coordination of the host. During the medical academic discussion process, the original mixed speech can be collected.
[0086] In step 102 of some embodiments, the transcribed text refers to the text obtained by converting the original mixed speech. The content of the transcribed text is consistent with the content of the original mixed speech.
[0087] In some embodiments, before step 102, the speech data processing method may further include: performing speech preprocessing on the original mixed speech. Specifically, the process of speech preprocessing includes: performing noise reduction processing on the original mixed speech to obtain noise-reduced speech; performing speech compression processing on the noise-reduced speech according to a preset volume threshold to obtain compressed speech, and determining the compressed speech as the original mixed speech. The embodiments of the present application can remove obvious noises in the speech for subsequent processing. And, for the problem that there may be uneven volume in different segments of the original mixed speech, the volume of the original mixed speech is adjusted through speech compression technology, so as to make the loudness levels of all segments of the original mixed speech as consistent as possible.
[0088] In step 103 of some embodiments, the candidate host keywords are the keywords that the host object may say, and the candidate host keywords are used to represent the keywords for the host object to indicate the start or end of the speech of the target speaker. For example, the candidate host keywords may include keywords such as "The first debater of the affirmative side starts to speak", "The first debater of the negative side starts to speak", "Invite the second debater of the affirmative side", "Invite the second debater of the negative side", "Please the first debater of the affirmative side to end the speech", "Please the third debater of the negative side to end the speech", etc. It should be noted that these keywords have high regularity in a debate, and can provide reliable time points to locate the start and end of the conversation in a debate, so as to accurately segment the speech into audio segments of different target speakers.
[0089] It should be noted that the target host keyword refers to the candidate host keyword that appears in the original mixed speech.
[0090] In step 104 of some embodiments, the time point of the target host keyword in the original mixed speech refers to the time position where the target host keyword appears in the original mixed speech. It should be noted that each word in the transcribed text has a corresponding time position in the original mixed speech.
[0091] Specifically, the time point can be a timestamp. A timestamp is a time format. In audio data, timestamps generally increase steadily according to the number of audio frames. For example, in the original mixed speech, the timestamp of the first frame is 0, the timestamp of the second frame is 1, and so on. The time point can also adopt other time formats, not limited to this.
[0092] In some embodiments, for example, in a debate scenario, the transcribed text of the original mixed speech includes the following conversation:
[0093] Sentence 1: "First, please invite the first debater of the affirmative side to start speaking."
[0094] Sentence 2: "Hello everyone, I am today's debater. Thank you all for giving me this opportunity to express my views."
[0095] Suppose the candidate host keywords include "the first debater of the affirmative side starts speaking" and "the second debater of the affirmative side starts speaking". Then, the keyword "the first debater of the affirmative side starts speaking" that appears in Sentence 1 can be determined as the target host keyword.
[0096] In some embodiments, the time point at which the last character of the target host keyword appears in the original mixed speech can be used as the time point of the target host keyword in the original mixed speech. Using the same example above, assume that the time point at which the last character of "the first debater of the affirmative side starts speaking" appears in the original mixed speech is 1 minute and 05 seconds. Then, the time point of the target host keyword in the original mixed speech is 1 minute and 05 seconds (01:05).
[0097] It should be noted that the mixed speech segment refers to the speech segment obtained by segmenting the original mixed speech according to the time point of the target host keyword in the original mixed speech. Using the same example above, assume that the time range of the original mixed speech is [00:00 - 05:00] and the time point of the target host keyword in the original mixed speech is (01:05). Then, the original mixed speech can be segmented into a speech segment of [00:00 - 01:05] and a speech segment of [01:05 - 05:00] according to (01:05).
[0098] In step 105 of some embodiments, the single-speaker voice sub-segment is the voice segment obtained by separating the voices of the mixed speech segment, and each single-speaker voice sub-segment includes the separate voice of the same speaker. This can reduce the voice overlap phenomenon in the voice data, improve the clarity of the voice, thereby improving the quality of the voice data, and enabling the voice of each speaker to be more accurately recognized.
[0099] It should be noted that voice separation refers to separating the mixed voices into the individual voices of each speaker. Specifically, voice separation can be performed through an automated script or a voice separation algorithm, and the embodiments of the present application do not limit this.
[0100] For example, in a debate scenario, the mixed voice segment may contain a voice formed by the voices of the host and the opposing debater mixed together. After performing voice separation on the mixed voice segment, voice segments A1 and B1 are obtained, where voice segment A2 only contains the voice of the host, and voice segment B2 only contains the voice of the opposing debater.
[0101] Another example is in the fintech scenario, such as in the conversation service of the insurance scenario. The mixed voice segment may contain a voice formed by the voices of the insurance service provider and the user mixed together. After performing voice separation on the mixed voice segment, voice segments A2 and B2 are obtained, where voice segment A2 only contains the voice of the insurance service provider, and voice segment B2 only contains the voice of the user. Another example is in the medical and health scenario, such as in the medical academic conference scenario. The mixed voice segment may contain a voice formed by the voices of medical expert DA and medical expert DB mixed together. After performing voice separation on the mixed voice segment, voice segments A3 and B3 are obtained, where voice segment A3 only contains the voice of medical expert DA, and voice segment B3 only contains the voice of medical expert DB.
[0102] In step 106 of some embodiments, a clustering algorithm based on voice features can be used to classify the single-speaker object voice sub-segments, so as to obtain the clustering result of the same speaker (i.e., the target speaker object), that is, the target single-speaker object voice classification result. By using a clustering algorithm for speaker object classification, the possibility of incorrect separation caused by voice overlap or background noise can be reduced.
[0103] In step 107 of some embodiments, the single-speaker object voice sub-segments with the same target single-speaker object voice classification result belong to the same speaker object. During the process of voice segment splicing, the single-speaker object voice sub-segments in each target single-speaker object voice classification result can be spliced in chronological order to obtain the target single-speaker object voice of each target speaker object. In this way, the segments of the same speaker can be merged in chronological order to generate an uninterrupted audio stream. It should be noted that the target single-speaker object voice contains the voice of the same target speaker object in the original mixed voice.
[0104] In some embodiments, after step 107, the speech data processing method may further include: performing speech enhancement on the target single-speaker speech. Specifically, the process of speech enhancement may include: using an advanced noise reduction algorithm (such as the rnnoise algorithm) to denoise the target single-speaker speech, thereby removing background noise in the target single-speaker speech and reducing environmental interference (such as audience interference sounds). Reverberation cancellation technology may also be used to cancel the reverberation of the target single-speaker speech, thereby optimizing the audio quality.
[0105] In some embodiments, after step 107, the speech data processing method may further include: performing text conversion on each target single-speaker speech to obtain a target transcription text; extracting features from each target single-speaker speech to obtain target speech features (including speaker features, style features, and prosody features); generating a file in a preset file format according to the target transcription text and the target speech features to obtain a speech annotation file for each target single-speaker speech, so as to construct a speech data set based on each target single-speaker speech and each speech annotation file. The speech data set constructed in this way can be used to evaluate the speech quality or perform model training, which is not limited in the embodiments of the present application.
[0106] In some embodiments, it should be noted that the current speech datasets have deficiencies in supporting specific scenarios (such as debate scenarios). For example, most of the speeches in speech datasets are mainly collected in ordinary conversation scenarios or reading scenarios, lacking speeches collected in complex debate scenarios. The speech data in debate scenarios has significant prosody, intonation, and emotional characteristics, which are not possessed by the data in the current speech datasets. Therefore, it is difficult to meet the training requirements of the debate speech generation model. Herein, the debate speech generation model refers to an artificial intelligence model used to generate debate speeches. In addition, the speeches collected in debate scenarios usually have problems such as background noise and speech overlap, resulting in poor data quality and being difficult to use directly. Specifically, in multi-speaker scenarios such as debates and conversations, multiple speakers may speak simultaneously, making it difficult to separate and annotate the speech data. Moreover, the difficulty in annotating speech data may lead to the lack of annotation of meta-information in the debate scenario for the speech data, such as style feature information like pitch, duration, and prosody of the speech data collected in the debate. These information are crucial for model training. To address the above problems, the speech data processing method of the embodiments of the present application can reduce the speech overlap phenomenon of speech data, improve the quality of speech data, thereby reducing the difficulty of annotating speech data, so as to annotate the style feature information in the debate scenario for the speech data, and can better support the debate scenario. The speech data processing method can also automatically process speech data without manual annotation, significantly reducing the labor cost and time cost, improving the processing efficiency of speech data, and meeting the data processing requirements of industrial scenarios for massive and diverse speech data.
[0107] Please refer to Figure 2 , in some embodiments, step 104 may include but is not limited to steps 201 to 203:
[0108] Step 201, perform time division on the original mixed speech according to the time points of the target host keywords in the original mixed speech to obtain at least two initial time ranges;
[0109] Step 202, update the start time or end time of the initial time range according to the start time point or end time point of the speech of any target speaker to obtain the target time range;
[0110] Step 203, perform segment interception on the original mixed speech according to the target time range to obtain the mixed speech segment.
[0111] The advantage of this embodiment is that the original mixed speech is divided in time according to the time points of the target host keywords in the original mixed speech, obtaining at least two initial time ranges, thereby initially dividing the time ranges corresponding to the speeches of different speakers. Then, the start time or end time of the initial time range is updated according to the start time point or end time point of the speech of any target speaker. In this way, the actual speech time range of the target speaker in the speech can be determined more accurately, improving the accuracy of the target time range, so that the speech segments divided according to the target time range subsequently contain as few speakers as possible, thereby reducing the speech overlap phenomenon in the speech data and improving the quality of the speech data.
[0112] In step 201 of some embodiments, specifically, the target host keywords may include a host start keyword and / or a host end keyword. As for the meanings and functions of the host start keyword and the host end keyword, reference may be made to the specific description of step 301 below.
[0113] In step 202 of some embodiments, the start time point and / or end time point of the speech of the target speaker may be obtained from the annotation data of the original mixed speech. In another embodiment, the start time point and / or end time point of the speech may also be determined according to a preset specific sound. For example, in a debate scenario, assuming that the ringing sound is used to indicate the end of a debater's speech, the time point of the ringing sound of the clock in the original mixed speech may be used as the end time point of the speech.
[0114] For example, in a debate scenario, the start time of the initial time range is (01:00) and the end time is (04:30). Assuming that the end time point of the first debater of the affirmative side is (04:32), the end time of the initial time range may be updated to (04:32) to obtain the target time range [01:00, 04:32].
[0115] In some embodiments, if the start time point of the speech of the target speaker is obtained, the time difference between the start time point of the speech of the target speaker and the start time of the initial time range may be calculated to obtain the start time difference. If the start time difference is less than or equal to a preset time difference threshold, the start time of the initial time range is updated according to the start time point of the speech of the target speaker to obtain the target time range. It should be noted that if the start time point of the speech of the target speaker is too different in time from the start time of the initial time range, it indicates that the start time point is very likely not the actual start time of the target speaker in this initial time range. Therefore, the start time is updated only when the start time difference is less than or equal to the time difference threshold, improving the reliability of updating the time range.
[0116] In another embodiment, if the speech end time point of the target speaker is obtained, the time difference between the speech end time point of the target speaker and the end time of the initial time range can be calculated to obtain the end time difference. If the end time difference is less than or equal to a preset time difference threshold, the end time of the initial time range is updated according to the speech end time point of the target speaker to obtain the target time range. It should be noted that the end time is updated only when the end time difference is less than or equal to the time difference threshold, which improves the reliability of updating the time range.
[0117] In step 203 of some embodiments, the time range of the mixed speech segment is consistent with the target time range.
[0118] Please refer to Figure 3 , in some embodiments, at least two target host keywords include a host start keyword and a host end keyword;
[0119] Step 201 may include but is not limited to steps 301 to 302:
[0120] Step 301, determine the time range between the time point of the host end keyword in the original mixed speech and the time point of the host start keyword in the original mixed speech as the host time range;
[0121] Step 302, intercept the time range outside the host time range from the original mixed speech to obtain at least two initial time ranges.
[0122] The advantage of this embodiment is that according to the time range between the host end keyword and the host start keyword, the time range of the host's speech in the voice, that is, the host time range, is determined. Since the original mixed speech is the speech collected from the host object and at least two target speakers, the time range outside the host time range can be intercepted from the original mixed speech, so that the time range of the target speakers' speech except the host's speech can be determined, and thus the time range mainly containing the conversations between the target speakers, that is, the initial time range, can be obtained more accurately, so as to accurately divide the original mixed speech of multiple speakers into the speech of a single speaker, improving the quality of the voice data.
[0123] In step 301 of some embodiments, the host start keyword refers to the keyword for the host object to indicate the start of the target speaker's speech. The host end keyword refers to the keyword for the host object to indicate the end of the target speaker's speech.
[0124] For example, in a debate scenario, the host start keywords may include "The first debater of the affirmative side starts to speak", "The first debater of the negative side starts to speak", "Invite the second debater of the affirmative side", "Invite the second debater of the negative side", etc. The host end keywords may include "Please end your speech, the first debater of the affirmative side", "Please end your speech, the third debater of the negative side", etc.
[0125] In some embodiments, step 301 includes: determining the time point of the host end keyword in the original mixed speech as the first host time point; determining the time point of the host start keyword in the original mixed speech as the second host time point; if the first host time point and the second host time point are adjacent time points and the first host time point is less than the second host time point, then determining the time range from the first host time point to the second host time point as the host time range.
[0126] It should be noted that the host time range refers to the time range during which the host object (i.e., the host) speaks in the original mixed speech. For example, assume that the time point of the host end keyword "Please end your speech, the first debater of the affirmative side" in the original mixed speech is time point T1 (01:00), and the time point of the host start keyword "Invite the first debater of the negative side" in the original mixed speech is time point T2 (02:20). Then the time period from time point T1 to time point T2 is the time during which the host speaks, that is, the host time range is [01:00, 02:20].
[0127] In step 302 of some embodiments, the initial time range refers to the time range corresponding to the voice part of the target speaker in the original mixed speech. For example, in a debate scenario, assume that the original mixed speech is a 5-minute voice, that is, the time range of the original mixed speech is [00:00, 05:00], and the host time range is [01:00, 02:20]. Then two initial time ranges can be obtained. One initial time range is [00:00, 01:00], and the other initial time range is [02:20, 05:00]. Each initial time range is equivalent to the time range of the conversation between each target speaker (such as the debaters of the affirmative side and the negative side).
[0128] Please refer to Figure 4 , in some embodiments, before step 202, the voice data processing method may include but is not limited to steps 401 to 402:
[0129] Step 401, extracting the end speech keyword from the original mixed speech according to the preset candidate end speech keywords to obtain the target end speech keyword;
[0130] Step 402, determining the time point of the target end speech keyword in the original mixed speech as the speech end time point.
[0131] The advantage of this embodiment is that the target end-of-speech keyword is extracted from the original mixed speech, and then, based on the time point of the target end-of-speech keyword in the original mixed speech, the actual end time point of the target speaker's speech is obtained, so as to obtain a more reliable speech end time point for subsequent use in segmenting the original mixed speech, thereby reducing the speech overlap phenomenon in the speech data and improving the quality of speech processing.
[0132] In step 401 of some embodiments, the candidate end-of-speech keyword refers to the keyword for the target speaker to end their speech. For example, in a debate scenario, the target speakers may include the positive debaters (such as the first positive debater, the second positive debater, etc.) and the negative debaters (such as the first negative debater, the second negative debater, etc.). Among them, the candidate end-of-speech keywords for the positive debaters may include "the positive speech ends", and the candidate end-of-speech keywords for the negative debaters may include "the negative speech ends". In addition, the candidate end-of-speech keywords may also include other keywords, not limited to this.
[0133] It should be noted that the target end-of-speech keyword is the candidate end-of-speech keyword that appears in the original mixed speech.
[0134] In step 402 of some embodiments, the speech end time point is the time point of the target end-of-speech keyword in the original mixed speech.
[0135] Please refer to Figure 5 , in some embodiments, at least two target speakers include a first speaker and a second speaker;
[0136] Step 106 may include but is not limited to steps 501 to 504:
[0137] Step 501, obtaining the speech feature of the first speaker to obtain the first speech feature;
[0138] Step 502, obtaining the speech feature of the second speaker to obtain the second speech feature;
[0139] Step 503, training the pre-constructed initial clustering model according to the first speech feature and the second speech feature to obtain the target clustering model;
[0140] Step 504, clustering at least two single-speaker speech sub-fragments through the target clustering model to obtain the first single-speaker speech classification result and the second single-speaker speech classification result; wherein, the first single-speaker speech classification result includes the single-speaker speech sub-fragments of the first speaker; the second single-speaker speech classification result includes the single-speaker speech sub-fragments of the second speaker.
[0141] The advantage of this embodiment is that, in order to classify multiple single-speaker voice sub-fragments according to the speaker, an initial clustering model is trained based on the first voice feature of the first speaker and the second voice feature of the second speaker, so that the trained target clustering model can distinguish the voice segments of different speakers (i.e., single-speaker voice sub-fragments). Then, the target clustering model performs clustering processing on at least two single-speaker voice sub-fragments, and divides the multiple single-speaker voice sub-fragments into single-speaker voice sub-fragments of the first speaker and single-speaker voice sub-fragments of the second speaker. This can accurately classify the voice segments of the same speaker, thereby improving the accuracy of voice processing.
[0142] In step 501 of some embodiments, the first voice feature is the voice feature of the first speaker, which may specifically be the voiceprint feature of the first speaker.
[0143] In step 502 of some embodiments, the second voice feature is the voice feature of the second speaker, which may specifically be the voiceprint feature of the second speaker. It should be noted that the first speaker and the second speaker are different speakers. For example, the first speaker is a positive debater (such as the first debater of the positive side), and the second speaker is a negative debater (such as the first debater of the negative side). Another example is that the first speaker is a negative debater and the second speaker is a positive debater.
[0144] In some embodiments, a first sample voice containing only the voice of the first speaker can be obtained, and the first sample voice is encoded into a first voice feature by a voice encoder. A second sample voice containing only the voice of the second speaker can be obtained, and the second sample voice is encoded into a second voice feature by a voice encoder. In another embodiment, the first voice feature and the second voice feature can also be obtained from a voice feature database. The first voice feature and the second voice feature can also be obtained by other means, and are not limited thereto.
[0145] In step 503 of some embodiments, the initial clustering model can be a clustering model based on voice features, such as a Gaussian mixture model (GMM), a hidden Markov model (HMM), etc.
[0146] In some embodiments, during the process of training the pre-constructed initial clustering model, the loss value can be calculated for the acoustic features of the speaker (i.e., the first voice feature of the first speaker and the second voice feature of the second speaker) according to the loss function of the initial clustering model, and the parameters of the initial clustering model can be adjusted according to the loss value, such as adjusting the number of clusters, so as to minimize the loss value, thereby improving the classification accuracy of the target clustering model.
[0147] In step 504 of some embodiments, for example, assume that the first speaking object is the first debater of the affirmative side and the second speaking object is the first debater of the negative side. Then the speech classification result of the first single speaking object can be the set of all speech sub - segments of the first debater of the affirmative side, and the speech classification result of the second single speaking object can be the set of all speech sub - segments of the first debater of the negative side.
[0148] Please refer to Figure 6 , in some embodiments, before step 107, the speech data processing method may include but is not limited to steps 601 to 605:
[0149] Step 601, perform time - frequency conversion on each single - speaking - object speech sub - segment to obtain speech spectrum data;
[0150] Step 602, extract low - frequency components from the speech spectrum data to obtain fundamental frequency data;
[0151] Step 603, perform energy peak analysis on the fundamental frequency data to obtain the fundamental frequency energy peak;
[0152] Step 604, if the fundamental frequency energy peak is greater than or equal to a preset energy peak threshold, determine the speech segment corresponding to the fundamental frequency energy peak in the single - speaking - object speech sub - segment as the overlapping speech segment;
[0153] Step 605, perform speech deletion processing on the overlapping speech segments of each single - speaking - object speech sub - segment to obtain the remaining speech segments, and determine the remaining speech segments as the single - speaking - object speech sub - segments.
[0154] The advantage of this embodiment is that considering that the speech separation process may not be able to completely separate the overlapping speech in the mixed speech segment, in order to detect the possible overlapping speech segments in the single - speaking - object speech sub - segment, by converting the single - speaking - object speech sub - segment into speech spectrum data and extracting the fundamental frequency data from the speech spectrum data, the human voice in the single - speaking - object speech sub - segment can be accurately analyzed. If the fundamental frequency energy peak of the fundamental frequency data is greater than or equal to the preset energy peak threshold, it means that there is more than one human voice in this part of the frequency, that is, there is overlapping speech. Therefore, the speech segment corresponding to the fundamental frequency energy peak in the single - speaking - object speech sub - segment is determined as the overlapping speech segment, and the overlapping speech segment is deleted, thereby reducing the possibility of overlapping speech in the speech data (i.e., the single - speaking - object speech sub - segment), improving the quality of the speech data, and further reducing the interference of the speech data on the subsequent model training.
[0155] In step 601 of some embodiments, time-frequency conversion can be performed on each single-speaker voice sub-fragment through an algorithm based on time-frequency analysis, such as the short-time Fourier transform (STFT) algorithm.
[0156] In step 602 of some embodiments, it should be noted that the fundamental frequency is the lowest frequency component in a sound signal, and the fundamental frequency usually corresponds to the basic pitch of the sound. In the voice spectrum data, the fundamental frequency is the lowest frequency and also the most prominent frequency. Therefore, the fundamental frequency of the voice spectrum data is equivalent to the voice of the main speaker in the voice spectrum data.
[0157] In step 603 of some embodiments, the fundamental frequency energy peak refers to the frequency part with the highest energy at the fundamental frequency.
[0158] In step 604 of some embodiments, when the fundamental frequency energy peak is greater than or equal to a preset energy peak threshold, it indicates that the energy of this part of the frequency is relatively high, that is, there is more than one human voice in the voice corresponding to this part of the frequency. In another embodiment, if the number of fundamental frequencies is relatively large, it indicates that there may be more than one human voice in this part of the voice. Considering the above situation, the number of fundamental frequencies of the voice spectrum data can also be compared with a preset fundamental frequency number threshold. If the number of fundamental frequencies is greater than or equal to the fundamental frequency number threshold and the fundamental frequency energy peak is greater than or equal to the preset energy peak threshold, the voice segment corresponding to the fundamental frequency energy peak in the single-speaker voice sub-fragment is determined as the voice overlap segment.
[0159] In step 605 of some embodiments, the remaining voice segment refers to the voice segment in the single-speaker voice sub-fragment except for the voice overlap segment.
[0160] Please refer to Figure 7 , in some embodiments, step 601 may include but is not limited to steps 701 to 705:
[0161] Step 701, perform start voice segment selection on each single-speaker voice sub-fragment according to the start time of each single-speaker voice sub-fragment and a preset overlap segment duration threshold to obtain a start voice segment;
[0162] Step 702, perform end voice segment selection on each single-speaker voice sub-fragment according to the end time of each single-speaker voice sub-fragment and the overlap segment duration threshold to obtain an end voice segment;
[0163] Step 703: Perform a first time-frequency conversion process on the starting speech segment to obtain a starting speech spectrum;
[0164] Step 704: Perform a second time-frequency conversion process on the ending speech segment to obtain an ending speech spectrum;
[0165] Step 705: Determine the starting speech spectrum and the ending speech spectrum as speech spectrum data.
[0166] The advantage of this embodiment is that when detecting the speech overlapping region, instead of detecting all single-speaker object speech sub-segments, it specifically selects some regions for detection. Specifically, considering that the start and end parts of the speech are most likely to have multiple people speaking, it specifically selects the starting speech segment and the ending speech segment of the single-speaker object speech sub-segment, and then converts the starting speech segment into a starting speech spectrum and the ending speech segment into an ending speech spectrum, so as to detect the overlapping segments of human voices in the starting speech spectrum and the ending speech spectrum, thereby being able to focus on detecting whether there are overlapping segments of human voices in the start and end parts of the single-speaker object speech sub-segment, reducing the time cost and computing resources required for detection, and improving the speech processing efficiency.
[0167] In step 701 of some embodiments, specifically, the overlapping segment duration threshold can be a relatively short duration value. For example, the overlapping segment duration threshold can be 5 seconds. The overlapping segment duration threshold can also be other values, and the embodiments of the present application do not limit this.
[0168] For example, assuming that the starting time of the single-speaker object speech sub-segment is (01:00) and the overlapping segment duration threshold is 5 seconds, the time range of the starting speech segment can be [01:00, 01:05].
[0169] In step 702 of some embodiments, for example, assuming that the ending time of the single-speaker object speech sub-segment is (04:00) and the overlapping segment duration threshold is 5 seconds, the time range of the starting speech segment can be [03:55, 04:00].
[0170] In step 703 of some embodiments, the starting speech segment can be subjected to time-frequency conversion through an algorithm based on time-frequency analysis, such as the short-time Fourier transform (STFT) algorithm, to obtain a starting speech spectrum.
[0171] In step 704 of some embodiments, the ending speech segment can be subjected to time-frequency conversion through an algorithm based on time-frequency analysis to obtain an ending speech spectrum.
[0172] In step 705 of some embodiments, the speech spectrum data includes the starting speech spectrum and the ending speech spectrum.
[0173] Please refer to Figure 8 , an embodiment of the present application further provides a voice data processing device, which can implement the above voice data processing method. The device includes:
[0174] A voice acquisition module 801, configured to acquire the original mixed voice; wherein, the original mixed voice is the voice collected from the host object and at least two target speaker objects;
[0175] A text conversion module 802, configured to perform text conversion processing on the original mixed voice to obtain a transcribed text;
[0176] A keyword extraction module 803, configured to extract a target host keyword from the transcribed text according to a preset candidate host keyword, where the target host keyword is used to represent a keyword for the host object to indicate the start or end of a target speaker's speech;
[0177] A voice segmentation module 804, configured to segment the original mixed voice according to the time point of the target host keyword in the original mixed voice to obtain at least two mixed voice segments;
[0178] A voice separation module 805, configured to separate the voice of each mixed voice segment according to the speaker to obtain a single speaker voice sub-segment;
[0179] A voice classification module 806, configured to classify the at least two single speaker voice sub-segments according to the voice features of at least two target speaker objects to obtain at least two target single speaker voice classification results; wherein, each target single speaker voice classification result includes the single speaker voice sub-segments of the same target speaker object;
[0180] A voice splicing module 807, configured to splice the single speaker voice sub-segments of each target single speaker voice classification result to obtain the target single speaker voice of each target speaker object.
[0181] In one embodiment, the voice data processing device further includes an end speech time point determination module, configured to: extract an end speech keyword from the original mixed voice according to a preset candidate end speech keyword to obtain a target end speech keyword; determine the time point of the target end speech keyword in the original mixed voice as the end speech time point.
[0182] In one embodiment, the voice data processing device further includes a voice overlapping segment deletion module, which is configured to: perform time-frequency conversion on each single-speaker voice sub-segment to obtain voice spectrum data; extract low-frequency components from the voice spectrum data to obtain fundamental frequency data; perform energy peak analysis on the fundamental frequency data to obtain the fundamental frequency energy peak; if the fundamental frequency energy peak is greater than or equal to a preset energy peak threshold, determine the voice segment corresponding to the fundamental frequency energy peak in the single-speaker voice sub-segment as a voice overlapping segment; perform voice deletion processing on the voice overlapping segment of each single-speaker voice sub-segment to obtain a remaining voice segment, and determine the remaining voice segment as the single-speaker voice sub-segment.
[0183] The specific implementation manner of this voice data processing device is basically the same as the specific embodiments of the above voice data processing method, and will not be elaborated here.
[0184] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above voice data processing method. This electronic device can include any intelligent terminal such as a tablet computer or an in-vehicle computer.
[0185] Please refer to Figure 9 , Figure 9 which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0186] A processor 901, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0187] A memory 902, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the voice data processing method of the embodiments of the present application;
[0188] An input / output interface 903, which is used to implement information input and output;
[0189] A communication interface 904, which is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0190] A bus 905, which transmits information between various components of the device (such as a processor 901, a memory 902, an input / output interface 903, and a communication interface 904);
[0191] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.
[0192] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned voice data processing method is implemented.
[0193] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0194] It should be noted that the non-company software tools or components appearing in the embodiments of the present application are only introduced by way of example and do not represent actual use.
[0195] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0196] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0197] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0198] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0199] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0200] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (item) of the following" or a similar expression means any combination of these items, including any combination of single item (s) or plural item (s). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0201] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0202] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0204] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.
[0205] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A method for processing speech data, characterized in that: The method comprises: Acquire original mixed speech; wherein the original mixed speech is speech collected from the host object and at least two target speaking objects; Performing text conversion processing on the original mixed speech to obtain a transcribed text; According to the preset candidate host keywords, the host keyword is extracted from the transcribed text to obtain the target host keyword; wherein the target host keyword is used to represent the keyword used by the host object to instruct the target speaker to start or end the speech; Segmenting the original mixed speech according to the time point of the target host keyword in the original mixed speech to obtain at least two mixed speech segments; Performing speech separation on each of the mixed speech segments according to the speaker to obtain single-speaker speech sub-segments; According to the speech features of at least two of the target speakers, at least two of the single speaker speech sub-segments are classified as speakers, to obtain at least two target single speaker speech classification results; wherein each of the target single speaker speech classification results includes the same single speaker speech sub-segment of the target speaker; The single-speaker-object speech sub-segments of each target single-speaker-object speech classification result are concatenated to obtain the target single-speaker-object speech of each target single-speaker-object.
2. The method according to claim 1, characterized in that The method of performing speech segmentation on the original mixed speech according to the time point of the target host keyword in the original mixed speech to obtain at least two mixed speech segments includes: Dividing the original mixed speech into time ranges according to the time points of the target host keyword in the original mixed speech to obtain at least two initial time ranges; The starting time or the ending time of the initial time range is updated according to the speech starting time or the speech ending time of any one of the target speakers to obtain the target time range; The original mixed speech is segmented according to the target time range to obtain mixed speech segments.
3. The method according to claim 2, characterized in that Before updating the start time or end time of the initial time range according to the speech start time or speech end time of any one of the target speech objects to obtain the target time range, the method further includes: Extracting the end-of-speech keywords from the original mixed speech according to the preset candidate end-of-speech keywords to obtain the target end-of-speech keywords; The time point at which the target end-of-speech keyword appears in the original mixed speech is determined as the speech end time point.
4. The method according to claim 2, characterized in that: The at least two target hosting keywords include a hosting start keyword and a hosting end keyword; The step of dividing the original mixed speech into time ranges according to the time points of the target host keyword in the original mixed speech to obtain at least two initial time ranges includes: Determine the time range between the time point of the host end keyword in the original mixed voice and the time point of the host start keyword in the original mixed voice as the host time range; The time range outside the hosting time range is intercepted from the original mixed speech to obtain at least two initial time ranges.
5. The method according to any one of claims 1 to 4, characterized in that: The at least two target speech objects include a first speech object and a second speech object; The step of classifying the at least two speech sub-segments of the single-speaking object according to the speech features of the at least two target speaking objects to obtain at least two speech classification results of the target single-speaking object includes: Acquire the speech feature of the first speaker to obtain a first speech feature; Acquire the voice feature of the second speaker to obtain a second voice feature; According to the first speech feature and the second speech feature, performing model training on a pre-constructed initial clustering model to obtain a target clustering model; At least two of the single-speaker object speech sub-segments are clustered by the target clustering model to obtain a first single-speaker object speech classification result and a second single-speaker object speech classification result; wherein the first single-speaker object speech classification result includes the single-speaker object speech sub-segment of the first speaker object; and the second single-speaker object speech classification result includes the single-speaker object speech sub-segment of the second speaker object.
6. The method according to any one of claims 1 to 4, characterized in that: Before performing speech segment splicing on the single-speaker speech sub-segments of each target single-speaker speech classification result to obtain the target single-speaker speech of each target speaker, the method further includes: Performing time-frequency conversion on each of the single-speaker speech sub-segments to obtain speech spectrum data; Extracting low-frequency components from the speech spectrum data to obtain fundamental frequency data; Performing energy peak analysis on the fundamental frequency data to obtain fundamental frequency energy peak value; If the fundamental frequency energy peak is greater than or equal to a preset energy peak threshold, the speech segment corresponding to the fundamental frequency energy peak in the single-speaker speech sub-segment is determined as a human voice overlapping segment; The overlapping human voice segments of each of the single-speaker speech sub-segments are subjected to speech deletion processing to obtain a remaining speech segment, and the remaining speech segment is determined as the single-speaker speech sub-segment.
7. The method according to claim 6, characterized in that The step of performing time-frequency conversion on each of the single-speaker speech sub-segments to obtain speech spectrum data includes: According to the starting time of each of the single-speaker speech sub-segments and a preset overlapping segment duration threshold, selecting a starting speech segment for each of the single-speaker speech sub-segments to obtain a starting speech segment; According to the end time of each of the single-speaker voice sub-segments and the overlapping segment duration threshold, selecting an end voice segment for each of the single-speaker voice sub-segments to obtain an end voice segment; Performing a first time-frequency conversion process on the initial speech segment to obtain an initial speech spectrum; Performing a second time-frequency conversion process on the ending speech segment to obtain an ending speech spectrum; The starting speech spectrum and the ending speech spectrum are determined as the speech spectrum data.
8. A voice data processing device, characterized in that: The device comprises: A speech acquisition module, used to acquire original mixed speech; wherein the original mixed speech is speech collected from a host object and at least two target speaking objects; A text conversion module, used for performing text conversion processing on the original mixed speech to obtain a transcribed text; A keyword extraction module is used to extract the host keyword from the transcribed text according to the preset candidate host keywords to obtain the target host keyword; wherein the target host keyword is used to represent the keyword used by the host object to instruct the target speaker to start or end the speech; A speech segmentation module, used for performing speech segmentation on the original mixed speech according to the time point of the target host keyword in the original mixed speech to obtain at least two mixed speech segments; A speech separation module, used for performing speech separation on each of the mixed speech segments according to the speaker, to obtain a single speaker speech sub-segment; A speech classification module, used to classify at least two of the single-speaker speech sub-segments according to speech features of at least two of the target speakers, to obtain at least two target single-speaker speech classification results; wherein each of the target single-speaker speech classification results includes the same single-speaker speech sub-segment of the target speaker; The speech splicing module is used to splice the single-speaker speech sub-segments of each target single-speaker speech classification result to obtain the target single-speaker speech of each target single-speaker.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the voice data processing method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the speech data processing method according to any one of claims 1 to 7 is implemented.