A voice role labeling method and device, electronic equipment and storage medium

By clustering and speaker segmentation of audio files, multiple audio segmentation datasets are generated. Voice role labeling is performed using identity identifiers and comparison results, which solves the problem of poor labeling effect in existing technologies and achieves accurate voice role labeling.

CN116756097BActive Publication Date: 2025-12-05BEIJING XUEZHITU NETWORK TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310525345.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-10
Publication Date
2025-12-05
Estimated Expiration
2043-05-10

AI Technical Summary

Technical Problem

Existing technologies have poor voice role labeling performance and cannot accurately label voice roles in audio files.

Method used

By acquiring multiple voice files of the target voice character, clustering and speaker segmentation are performed to generate first and second voice segmentation data. These data are then labeled, and accurate labeling is achieved using identity identifiers and comparison results.

Benefits of technology

It enables accurate annotation of voice roles in audio files, improving the accuracy and effectiveness of annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116756097B_ABST
    Figure CN116756097B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a voice role labeling method and device, electronic equipment and storage medium, wherein the voice role labeling method clusters a plurality of voice files of a target voice role, realizes voice aggregation of the target voice role, obtains mixed voice of the target voice role, performs speaker segmentation on the same speaker voice aggregation, obtains a first voice segmentation file corresponding to the mixed voice, obtains a plurality of independent time segments based on the mixed voice file, obtains second voice segmentation data based on the plurality of independent time segments, and performs target voice role labeling according to the first voice segmentation data and the second voice segmentation data, thereby realizing the effect of obtaining different voice segmentation data from the same clustering result, and labeling the target voice role based on mutual reference and comparison of different voice segmentation data, thereby realizing the effect of accurately labeling the voice role in the voice file.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a voice role labeling method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid increase of the acquisition channels and quantity of voice files, the management of voice files becomes more and more complex. By labeling the voice roles in the voice files, different voice roles can be analyzed, for example, the voice roles in the voice files in the sales business scenario are labeled to determine which is the customer and which is the salesperson. For the customer, the customer portrait can be further analyzed by combining voice recognition, and for the salesperson, the excellent sales skills can be extracted. Voice role labeling is the basis and premise of these works. In related technologies, the voice role labeling effect is poor, and the voice roles in the voice files cannot be accurately labeled. SUMMARY

[0003] Embodiments of the present application provide a voice role labeling method, device, electronic equipment and storage medium, wherein the method can accurately label the voice roles in the voice files, solving the problem that the voice role labeling effect is poor in related technologies and the voice roles in the voice files cannot be accurately labeled.

[0004] Other characteristics and advantages of the present application will become apparent from the following detailed description, or will be learned by practice of the present application.

[0005] According to an aspect of an embodiment of the present application, a voice role labeling method is provided, the method comprising: obtaining a plurality of voice files of a target voice role, and clustering the plurality of voice files to obtain a clustering result corresponding to the target voice role; performing speaker segmentation on the clustering result to obtain first voice segmentation data; performing splicing processing on the voice files contained in the clustering result to obtain spliced voice files, and performing speaker segmentation on the spliced voice files to obtain second voice segmentation data; and labeling the target voice role based on the first voice segmentation data and the second voice segmentation data.

[0006] In some examples, labeling the target voice role based on the first voice segmentation data and the second voice segmentation data comprises: determining an identity of the target voice role through the second voice segmentation data; comparing the first voice segmentation data and the second voice segmentation data based on the identity, and labeling the target voice role in the clustering result based on the comparison result.

[0007] In some examples, the speaker segmentation is performed on the clustering result to obtain first speech segmentation data, including: determining a constraint result based on a scene corresponding to a speech file of the target speech role, the constraint result being used to constrain a number of speech roles in the speech file; and performing speaker segmentation on the clustering result according to the constraint result to obtain the first speech segmentation data.

[0008] In some examples, the speech files contained in the clustering result are spliced to obtain spliced speech files, and speaker segmentation is performed on the spliced speech files to obtain second speech segmentation data, including: extracting a preset time length of a time segment from each speech file; splicing the obtained time segments to obtain the spliced speech files; and performing speaker segmentation on the spliced speech files to obtain the second speech segmentation data.

[0009] In some examples, after the target speech role is annotated based on the first speech segmentation data and the second speech segmentation data, the method further includes: checking an annotation result, and if the annotation result is lower than a target expectation, modifying the preset time length, and extracting new time segments from each speech file based on the modified preset time length; splicing the obtained new time segments to update the spliced speech files; and reacquiring the second speech segmentation data based on the updated spliced speech files.

[0010] In some examples, after the target speech role is annotated based on the first speech segmentation data and the second speech segmentation data, the method further includes: checking an annotation result, and if the annotation result is lower than a target expectation, extracting new time segments from each speech file based on the preset time length; splicing the obtained new time segments to update the spliced speech files; and reacquiring the second speech segmentation data based on the updated spliced speech files.

[0011] According to an aspect of an embodiment of the present application, a speech role annotation device is provided, the device comprising: a clustering module configured to acquire a plurality of speech files of a target speech role, and to cluster the plurality of speech files to obtain a clustering result corresponding to the target speech role; a segmentation module configured to perform speaker segmentation based on the clustering result to obtain first speech segmentation data; the segmentation module is further configured to splice speech files contained in the clustering result to obtain spliced speech files, and to perform speaker segmentation on the spliced speech files to obtain second speech segmentation data; and an annotation module configured to annotate the target speech role based on the first speech segmentation data and the second speech segmentation data.

[0012] In some examples, the labeling module is further configured to determine an identity of the target speech role based on the second speech segmentation data; compare the first speech segmentation data and the second speech segmentation data based on the identity; and label the target speech role in the clustering result based on a comparison result.

[0013] According to an aspect of some embodiments of the present application, an electronic device is provided. The electronic device includes one or more processors; and a storage device configured to store one or more computer programs that, when executed by the one or more processors, cause the electronic device to perform the method as described above.

[0014] According to an aspect of some embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor of an electronic device, the electronic device performs the method as described above.

[0015] According to an aspect of some embodiments of the present application, a computer program product is provided. The computer program product includes a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer program from the computer-readable storage medium. The processor executes the computer program, and causes the electronic device to perform the method as described above in various embodiments.

[0016] The scheme can be applied to advertisement analysis and processing in the field of data processing. In the technical scheme provided in the embodiments of the present application, a voice role labeling method comprises the following steps: obtaining a plurality of voice files of a target voice role, and clustering the plurality of voice files to obtain a clustering result corresponding to the target voice role; performing speaker segmentation based on the clustering result to obtain first voice segmentation data; performing splicing processing on the voice files contained in the clustering result to obtain spliced voice files, and performing speaker segmentation on the spliced voice files to obtain second voice segmentation data; and labeling the target voice role based on the first voice segmentation data and the second voice segmentation data. In the voice role labeling method, the plurality of voice files of the target voice role are clustered to realize voice aggregation of the target voice role, obtain mixed voice of the target voice role, and perform speaker segmentation on the same speaker voice aggregation to obtain first voice segmentation files corresponding to the mixed voice. Based on the mixed voice files, a plurality of independent time segments are obtained, and second voice segmentation data is obtained based on the plurality of independent time segments. The target voice role is labeled according to the first voice segmentation data and the second voice segmentation data, which realizes the effect of obtaining different voice segmentation data from the same clustering result, and the target voice role is labeled based on the mutual reference and comparison of different voice segmentation data, which realizes the effect of accurately labeling the voice role in the voice file, and avoids the problem that in related technologies, the voice role labeling effect is poor under single voice segmentation data, and the voice role in the voice file cannot be accurately labeled.

[0017] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present application. BRIEF DESCRIPTION OF DRAWINGS

[0018] The accompanying drawings, which are incorporated into and form part of the specification, illustrate one embodiment consistent with the present application and, together with the specification, serve to explain the principles of the application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. In the drawings:

[0019] Figure 1 is a basic flowchart of a voice role labeling method according to an embodiment of the present application;

[0020] Figure 2 is a basic flowchart of a voice role labeling method according to an embodiment of the present application;

[0021] Figure 3 is a schematic diagram of a voice role labeling device according to an embodiment of the present application;

[0022] Figure 4A structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. DETAILED DESCRIPTION

[0023] The exemplary embodiments will be described in detail herein below with reference to the drawings. In the following description, the same numbers in different drawings represent the same or similar elements unless otherwise represented. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.

[0024] The block diagrams shown in the drawings are merely functional entities, and do not necessarily have to correspond to physically independent entities. That is, the functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0025] The flowcharts shown in the drawings are merely exemplary illustrations, and do not necessarily include all the contents and operations, nor do they have to be executed in the order described. For example, some operations can be further divided, and some operations can be combined or partially combined, so that the actual execution order can be changed according to the actual situation.

[0026] It should also be noted that "multiple" as mentioned in the present application means two or more. The association relationship of "and / or" describes the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.

[0027] Embodiment One

[0028] To solve the above technical problems, the embodiments of the present application provide a voice role labeling method, as shown in Figure 1 The method comprises the following steps:

[0029] S101, a plurality of voice files of a target voice role are obtained, and the plurality of voice files are clustered to obtain a clustering result corresponding to the target voice role;

[0030] S102, speaker segmentation is performed based on the clustering result to obtain first voice segmentation data;

[0031] S103, the voice files contained in the clustering result are spliced to obtain spliced voice files, and speaker segmentation is performed on the spliced voice files to obtain second voice segmentation data;

[0032] S104, labeling the target voice role based on the first voice segmentation data and the second voice segmentation data.

[0033] It can be understood that the voice role labeling method provided in the embodiment can be applied to a terminal and / or a server, and the terminal can be implemented in various forms. For example, the terminal described in the embodiment can include a mobile terminal such as a mobile phone, a tablet computer, a notebook computer, a palm computer, a personal digital assistant (PDA), a portable media player (PMP), a navigation device, a wearable device, a smart bracelet, a pedometer, and the like, and a fixed terminal such as a digital TV, a desktop computer, a workstation, and the like. In the following embodiments, the voice role labeling method is applied to a mobile terminal as an example.

[0034] In the embodiment, a plurality of voice files of a target voice role are obtained, and each voice file contains the target voice role. For example, the plurality of voice files of the target voice role can be generated by the terminal itself, or obtained by the terminal from the outside. For example, the terminal obtains a plurality of voice files of each target voice role generated by the target voice role through voice communication. For another example, the reception telephone of a call center has a binding relationship with a seat and a work number, and the terminal takes each seat as a target voice role, and then obtains a plurality of voice files of each target voice role.

[0035] In the embodiment, the terminal clusters the plurality of voice files of the target voice role to obtain a clustering result corresponding to the target voice role. The clustering result contains a plurality of voice files of the same target voice role, and at least part of the voice files contain other voice roles in addition to the target voice role.

[0036] In some examples of the embodiment, speaker segmentation is performed based on the clustering result to obtain first voice segmentation data, including:

[0037] A constraint result is determined based on a scene corresponding to the voice file of the target voice role. The constraint result is used to constrain the number of voice roles in the voice file.

[0038] Speaker segmentation is performed on the clustering result based on the constraint result to obtain the first voice segmentation data.

[0039] In the embodiment, the speaker segmentation is used to automatically identify the number of voice roles in the voice file, detect the start and end time stamps of each voice file in the audio, and divide the voice file into a plurality of sub-voice files based on the start and end time stamps. Each sub-voice file contains only one voice role.

[0040] In order to make the result of speaker segmentation more accurate, the constraint result corresponding to the scene of the speech file of the target voice role can be determined, and the constraint result is used to constrain the number of voice roles in the speech file. For example, the scene corresponding to the speech file of the target voice role is a call center, a telephone recording, or the like. Since in the scenes of telephone recording and call center, most of the time is voice conversation between two people, if the scene corresponding to the speech file of the target voice role is a call, a telephone recording, or the like, the constraint result can be set to 2, and the clustering result is segmented based on the constraint result to obtain the first voice segmentation data.

[0041] As can be understood from the above example, the constraint result corresponding to each scene can be different. Therefore, before determining the constraint result based on the scene corresponding to the speech file of the target voice role, the method further includes setting a corresponding constraint result for each scene, so that more accurate first voice segmentation data is obtained when the clustering result is segmented based on the constraint result.

[0042] In some examples of the present embodiment, if the scene corresponding to the speech file cannot be confirmed or the corresponding constraint result for the scene is not set, the clustering result can be directly segmented for speaker segmentation to obtain the first voice segmentation data.

[0043] In some examples, the clustering result can also be directly segmented for speaker segmentation to obtain the first voice segmentation data without obtaining the scene corresponding to the speech file.

[0044] In some examples of the present embodiment, the speech files contained in the clustering result are spliced to obtain spliced speech files, and the spliced speech files are segmented for speaker segmentation to obtain second voice segmentation data, including:

[0045] A time segment of a preset time length is cut from each speech file;

[0046] The obtained time segments are spliced to obtain the spliced speech files;

[0047] The spliced speech files are segmented for speaker segmentation to obtain the second voice segmentation data.

[0048] In the present example, the specific length of the preset time length is not limited, and can be flexibly set by relevant personnel. For example, a time segment of 28 seconds is cut from each speech file, and then a plurality of time segments of 28 seconds are spliced to obtain the spliced speech files.

[0049] In the above example, in order to prevent the occurrence of voice endpoint detection segmentation errors during subsequent speaker segmentation, a silent segment can be added to each time file during the splicing process, and the time segments after adding the silent segments are spliced to obtain a spliced voice file. For example, taking a preset time length of 28 seconds as an example, a 28-second time segment is cut from each voice file, a 2-second silent segment is added to each time segment, so that the time length of each time segment reaches 30 seconds, and the 30-second time segments are spliced to obtain a spliced voice file.

[0050] It can be understood that the above-mentioned silent segment can be added at the beginning of each time segment or at the end of each time segment, and the present embodiment does not limit this.

[0051] In some examples of the present embodiment, labeling the target voice role based on the first voice segmentation data and the second voice segmentation data includes:

[0052] Determining the identity of the target voice role through the second voice segmentation data;

[0053] Comparing the first voice segmentation data and the second voice segmentation data based on the identity, and labeling the target voice role in the clustering result based on the comparison result.

[0054] In the above example, the second voice segmentation data is obtained by performing speaker segmentation on the spliced voice file. Since the spliced voice file is obtained based on the voice file of the target voice role, the spliced voice file contains more voice data of the target voice role. Therefore, in the second voice segmentation data obtained by performing speaker segmentation on the spliced voice file, the proportion of the target voice role will be higher than that of other voice roles. Therefore, the identity of the voice role with the highest proportion in the second voice segmentation data is determined as the identity of the target voice role.

[0055] In the above example, for example, the second voice segmentation data includes N segmentation results, each segmentation result representing voice data of a voice role. The identity of the voice role with the highest number of occurrences in the N segmentation results is obtained, and the identity of the voice role with the highest number of occurrences is taken as the identity of the target voice role.

[0056] In some examples, the first speech segmentation data and the second speech segmentation data are compared based on the identity identifier, and the target speech role is labeled in the clustering result based on a comparison result. For example, a preset time length is 28 seconds, a 28s time segment is cut from each speech file, the time segments are spliced to obtain a spliced speech file, the first speech segmentation data and the second speech segmentation data are compared based on the spliced speech file, the segmentation results of the first speech segmentation data and the second speech segmentation data are aligned, the identity identifier with the highest time coincidence degree with the target speech role is found, and the identity identifier with the highest time coincidence degree is labeled as the identity identifier of the target speech role, so that the target speech role in the multiple speech files in the clustering result is labeled.

[0057] In some examples of the embodiment, after the target speech role is labeled based on the first speech segmentation data and the second speech segmentation data, the method further includes: checking the labeling result, if the labeling result is lower than a target expectation, modifying the preset time length, and cutting new time segments from each speech file based on the modified preset time length; splicing the obtained new time segments to update the spliced speech file; and reacquiring the second speech segmentation data based on the updated spliced speech file. If the labeling result is lower than the target expectation, it indicates that the target speech role in the speech file of the clustering data has not been successfully labeled, and the preset time length needs to be modified, and new time segments are cut from each speech file based on the modified preset time length, so as to modify the second speech segmentation data. The terminal can compare after receiving a comparison instruction from a related person, or the terminal can automatically compare after labeling the target speech role. The target expectation is a value set by a related person according to actual needs, for example, the target expectation is set to 80%, if the target speech role is labeled based on the first speech segmentation data and the second speech segmentation data, if less than 80% of the speech files in the clustering result are successfully labeled, it is determined that the labeling result is lower than the target expectation.

[0058] In the case of the above example, if the annotation result is lower than the target expectation, the second speech segmentation data is updated by modifying the preset time length, and the target speech role is re-annotated based on the new second speech segmentation data and the first speech segmentation data. It can be understood that the modified preset time length is not limited in this embodiment, and the modified preset time length can be lower than or higher than the original preset time length. Preferably, the modified preset time length is higher than the original preset time length. Specifically, for example, the original preset time length is 28s, and the annotation result of the target speech role annotated based on the second speech segmentation data and the first speech segmentation data obtained based on the 28s preset time length is lower than the target expectation. At this time, the preset time length is modified to 30s, the second speech segmentation data is updated based on the modified preset time length, and the target speech role is re-annotated based on the new second speech segmentation data and the first speech segmentation data.

[0059] In some examples of this embodiment, after the target speech role is annotated based on the first speech segmentation data and the second speech segmentation data, the method further includes: checking the annotation result, and if the annotation result is lower than the target expectation, a new time segment is cut from each speech file based on the preset time length; the new time segments obtained are spliced to update the spliced speech file; and the second speech segmentation data is re-obtained based on the updated spliced speech file.

[0060] In the case of the above example, if the annotation result is lower than the target expectation, a new time segment is cut from each speech file after the annotation is completed. Specifically, for example, the time length of the speech file is 60s, the preset time length is 28s, and the time segment of 10s to 38s of the speech file is taken to obtain the second speech segmentation data. If the annotation result of the target speech role annotated based on the second speech segmentation data of the time segment of 10s to 38s and the first speech segmentation data is lower than the target expectation, a new time segment of 28s is selected from the speech file at this time, the second speech segmentation data is updated based on the new time segment of 28s, and the target speech role is re-annotated based on the new second speech segmentation data and the first speech segmentation data.

[0061] The voice role labeling method provided in the example comprises the following steps: obtaining a plurality of voice files of a target voice role, and clustering the plurality of voice files to obtain a clustering result corresponding to the target voice role; performing speaker segmentation based on the clustering result to obtain first voice segmentation data; performing splicing processing on the voice files contained in the clustering result to obtain spliced voice files, and performing speaker segmentation on the spliced voice files to obtain second voice segmentation data; and labeling the target voice role based on the first voice segmentation data and the second voice segmentation data. The plurality of voice files of the target voice role are clustered to realize voice aggregation of the target voice role, obtain mixed voice of the target voice role, and perform speaker segmentation on the same speaker voice aggregation to obtain first voice segmentation files corresponding to the mixed voice. A plurality of independent time segments are obtained based on the mixed voice files, and second voice segmentation data are obtained based on the plurality of independent time segments. The target voice role is labeled according to the first voice segmentation data and the second voice segmentation data, the effect of obtaining different voice segmentation data from the same clustering result is realized, the target voice role is labeled based on the mutual reference and comparison of different voice segmentation data, the effect of accurately labeling the voice role in the voice file is realized, and the problem that the voice role labeling effect is poor and the voice role in the voice file cannot be accurately labeled in the related art under single voice segmentation data is avoided.

[0062] Embodiment two

[0063] In order to better understand the present application, a more specific example is provided in the example for description:

[0064] The example provides a voice role labeling method, as shown in Figure 2 The example provides a voice role labeling method, as shown in

[0065] Step 1: Same speaker data aggregation

[0066] Depending on the information on the data source, the data is clustered according to the speaker, such as the reception telephone of the call center, the agent and the work number have a binding relationship, and the records of the same agent can be aggregated together in the database depending on these information.

[0067] Step 2: Speaker segmentation

[0068] There is a subdivided technical branch of voice technology: speaker diarization, which runs the speaker segmentation result of each voice in step 1. For call centers, telephone recording scenes, the result can be further constrained to be at most several people, so that the result is more accurate.

[0069] Step 3: Voice segment combination and speaker segmentation

[0070] The plurality of voice files aggregated in step 1 are spliced together, each voice file intercepting a time segment, such as 28s, a plurality of voice segments are spliced together, VAD (voice endpoint detection) segmentation errors are prevented in splicing, 2s of silence segments are added in splicing, and each voice file accounts for 30s (28+2) on average.

[0071] The spliced mixed voice is run for speaker log, because the same speaker appears frequently in the identified speaker segmentation result in each 30s short voice after splicing, and the total time length accounts for a high proportion.

[0072] Step 4: Aligning speaker segmentation results

[0073] The results of steps 2 and 3 are compared, 28s of voice in each voice file appears in the large mixed voice file, time positioning information can be tracked out, the result of step 3 can quickly find the ID of the common speaker, and with this anchor point, the 28s speaker segmentation result in each voice file and the 28s speaker segmentation result in the mixed voice are aligned, for example: the ID with the highest time coincidence degree in the two 28s segmentation results is aligned as the common speaker.

[0074] Step 5: Annotation result confirmation (optional)

[0075] The automatic checking result or the control instruction issued by the relevant personnel is checked, if the automatic checking result or the control instruction indicates that there is doubt about the annotation result, step 3 is returned to, the voice period is increased or the voice segment is selected, and the result is refreshed again to correct the unreliable result.

[0076] Embodiment three

[0077] Based on the same technical concept, the embodiment also provides a voice role annotation device, as shown in Figure 3 The device comprises:

[0078] A clustering module 1 is configured to acquire a plurality of voice files of a target voice role, and cluster the plurality of voice files to obtain a clustering result corresponding to the target voice role.

[0079] A segmentation module 2 is configured to perform speaker segmentation based on the clustering result to obtain first voice segmentation data.

[0080] The segmentation module 2 is further configured to perform splicing processing on the voice files contained in the clustering result to obtain spliced voice files, and perform speaker segmentation on the spliced voice files to obtain second voice segmentation data.

[0081] An annotation module 3 is configured to annotate the target voice role based on the first voice segmentation data and the second voice segmentation data.

[0082] The labeling module 3 is further configured to determine an identity of the target voice role based on the second voice segmentation data, compare the first voice segmentation data and the second voice segmentation data based on the identity, and label the target voice role in the clustering result based on a comparison result.

[0083] The voice role labeling method further includes: performing speaker segmentation on the clustering result to obtain first voice segmentation data, including: determining a constraint result based on a scene corresponding to the voice file of the target voice role, the constraint result being used to constrain a number of voice roles in the voice file; and performing speaker segmentation on the clustering result based on the constraint result to obtain the first voice segmentation data.

[0084] The voice role labeling method further includes: performing speaker segmentation on the clustering result to obtain first voice segmentation data, including: determining a constraint result based on a scene corresponding to the voice file of the target voice role, the constraint result being used to constrain a number of voice roles in the voice file; and performing speaker segmentation on the clustering result based on the constraint result to obtain the first voice segmentation data.

[0085] The voice role labeling method further includes: after labeling the target voice role based on the first voice segmentation data and the second voice segmentation data, checking a labeling result, and if the labeling result is lower than a target expectation, modifying the preset time length, and obtaining new time segments from each voice file based on the modified preset time length; splicing the obtained new time segments to update the spliced voice file; and reobtaining the second voice segmentation data based on the updated spliced voice file.

[0086] The voice role labeling method further includes: after labeling the target voice role based on the first voice segmentation data and the second voice segmentation data, checking a labeling result, and if the labeling result is lower than a target expectation, modifying the preset time length, and obtaining new time segments from each voice file based on the modified preset time length; splicing the obtained new time segments to update the spliced voice file; and reobtaining the second voice segmentation data based on the updated spliced voice file.

[0087] It can be understood that each module of the voice role labeling apparatus provided in the embodiment can be combined to implement each step of the voice role labeling method, and achieve the same technical effects as each step of the voice role labeling method, which will not be described herein.

[0088] Embodiment Four

[0089] Embodiments of the present application also provide an electronic device, comprising one or more processors, and a storage device, wherein the storage device is configured to store one or more computer programs, which, when executed by the one or more processors, cause the electronic device to implement the voice role labeling method as above.

[0090] Figure 4 A structural diagram of a computer system of an electronic device suitable for implementing embodiments of the present application is shown.

[0091] It should be noted that, Figure 4 The computer system 1800 of the electronic device shown is only an example and should not impose any limitation on the functions and use range of embodiments of the present application.

[0092] As Figure 4 shown, the computer system 1800 includes a processor (Central Processing Unit, CPU) 1801, which can perform various appropriate actions and processes according to programs stored in a Read-Only Memory (ROM) 1802 or programs loaded from a storage portion 1808 into a Random Access Memory (RAM) 1803, such as performing the methods in the above embodiments. In the RAM 1803, various programs and data required for system operation are also stored. The CPU 1801, the ROM 1802, and the RAM 1803 are connected to each other through a bus 1804. An Input / Output (I / O) interface 1805 is also connected to the bus 1804.

[0093] In some embodiments, the following components are connected to the I / O interface 1805: an input portion 1806 including a keyboard, a mouse, etc.; an output portion 1807 including a display such as a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc., and a speaker, etc.; a storage portion 1808 including a hard disk, etc.; and a communication portion 1809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication portion 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to the I / O interface 1805 as needed. A removable recording medium 1811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 1810 as needed, so that a computer program read therefrom is installed in the storage portion 1808 as needed.

[0094] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as a computer program. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing instructions for carrying out the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication section 1809, and / or installed from the detachable medium 1811. When the computer program is executed by the processor (CPU) 1801, various functions defined in the system of the present application are executed.

[0095] It should be noted that the computer readable medium shown in the embodiments of the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable signal medium can include a data signal propagating in a baseband or as part of a carrier wave in a propagated data signal, in which the computer readable computer program is carried. Such a propagated data signal can take on many forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium that can be used to carry or store the program for use by or in connection with an instruction execution system, apparatus or device. The computer program contained on the computer readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, or the like, or any suitable combination of the above.

[0096] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may be executed in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams or flowcharts, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer program products.

[0097] The units or modules involved in the embodiments of the present application can be implemented by software, or can be implemented by hardware, and the described units or modules can also be arranged in a processor. In some cases, the names of the units or modules do not constitute a limitation on the units or modules themselves.

[0098] Another aspect of the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the voice role labeling method as described above. The computer readable storage medium can be included in the electronic device described in the above embodiments, or can exist separately and not be assembled into the electronic device.

[0099] Another aspect of the present application also provides a computer program product, which includes a computer program stored in a computer readable storage medium. The processor of the electronic device reads the computer program from the computer readable storage medium, and the processor executes the computer program to make the electronic device execute the voice role labeling method as described above.

[0100] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units.

[0101] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope of the application being indicated by the following claims.

[0102] The above-described embodiments are merely illustrative for the present application and not intended to limit the implementation of the present application. Any modification, variation, or adaptation based on the main idea and spirit of the present application can be easily made by those skilled in the art without departing from the scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope as claimed in the following claims.

Claims

1. A voice role labeling method, characterized by, The method comprises: obtaining a plurality of voice files of a target voice role, and clustering the plurality of voice files to obtain a clustering result corresponding to the target voice role; performing speaker segmentation on the clustering result to obtain first voice segmentation data; performing splicing processing on voice files contained in the clustering result to obtain spliced voice files, and performing speaker segmentation on the spliced voice files to obtain second voice segmentation data, comprising: extracting a time segment of a preset time length from each voice file; splicing the obtained time segments to obtain the spliced voice files; performing speaker segmentation on the spliced voice files to obtain the second voice segmentation data; annotating the target voice role based on the first voice segmentation data and the second voice segmentation data, comprising: determining the identity of the target voice role through the second voice segmentation data; comparing the first voice segmentation data and the second voice segmentation data based on the identity, and annotating the target voice role in the clustering result based on the comparison result.

2. The method of claim 1, wherein, performing speaker segmentation on the clustering result to obtain first voice segmentation data, comprising: determining a constraint result based on the scene corresponding to the voice file of the target voice role, the constraint result being used to constrain the number of voice roles in the voice file; performing speaker segmentation on the clustering result based on the constraint result to obtain the first voice segmentation data.

3. The method of claim 1, wherein, After annotating the target voice role based on the first voice segmentation data and the second voice segmentation data, the method further comprises: viewing the annotation result, if the annotation result is lower than the target expectation, modifying the preset time length, and extracting new time segments from each voice file based on the modified preset time length; splicing the obtained new time segments to update the spliced voice files; re-obtaining the second voice segmentation data based on the updated spliced voice files.

4. The method of claim 1, wherein, After annotating the target voice role based on the first voice segmentation data and the second voice segmentation data, the method further comprises: viewing the annotation result, if the annotation result is lower than the target expectation, extracting new time segments from each voice file based on the preset time length; splicing the obtained new time segments to update the spliced voice files; re-obtaining the second voice segmentation data based on the updated spliced voice files.

5. A speech role labeling apparatus characterized by comprising: The device comprises: a clustering module configured to obtain a plurality of voice files of a target voice role, and cluster the plurality of voice files to obtain a clustering result corresponding to the target voice role; a segmentation module configured to perform speaker segmentation based on the clustering result to obtain first voice segmentation data; The segmentation module is further configured to perform splicing processing on the speech files contained in the clustering result to obtain spliced speech files, and perform speaker segmentation on the spliced speech files to obtain second speech segmentation data: a time segment of a preset time length is intercepted from each speech file; the obtained time segments are spliced to obtain the spliced speech files; the spliced speech files are subjected to speaker segmentation to obtain the second speech segmentation data; The labeling module is configured to label the target voice role based on the first speech segmentation data and the second speech segmentation data: determining an identity of the target voice role through the second speech segmentation data; comparing the first speech segmentation data and the second speech segmentation data based on the identity, and labeling the target voice role in the clustering result based on a comparison result.

6. An electronic device, comprising: comprising: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to perform the method of any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, a computer program is stored thereon, which, when executed by a processor of an electronic device, causes the electronic device to perform the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Speaker voice segmentation method and device, electronic equipment and storage medium

    CN114464192A

  • Speaker labeling method and device, electronic equipment and storage medium

    CN115985315A