Voice segmentation method, device, server and storage medium
Through the automated speech frequency separation and segmentation point determination method, the problem of low efficiency of manual segmentation is solved, and efficient and accurate speech segmentation is achieved.
Patent Information
- Application Number
- CN202210717004.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-06-23
AI Technical Summary
In the prior art, manual speech segmentation is performed manually, resulting in low speech segmentation efficiency.
By acquiring speech data, determining the speech frequency and separating it, using speech pause information and speech text to determine the speech segmentation points, and combining the preset frequency set and segmentation probability model for automatic speech segmentation.
Automated speech segmentation is achieved, which improves the efficiency and accuracy of speech segmentation.
Smart Images

Figure CN115101056B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a speech segmentation method, device, server, and storage medium. Background Art
[0002] Speech segmentation is a critical issue in speech processing, as longer speech segments consume significant system resources during speech recognition conversion and can result in low recognition accuracy. Speech segmentation can reduce the computational complexity of speech recognition and improve its accuracy.
[0003] In the related art, speech segmentation is usually performed manually, which results in low speech segmentation efficiency. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a speech segmentation method, device, server and storage medium to solve the problem in the related art of manually segmenting speech, resulting in low speech segmentation efficiency.
[0005] A first aspect of the embodiments of the present application provides a speech segmentation method, comprising:
[0006] Acquiring voice data, and determining whether the voice data includes multiple voice frequencies;
[0007] If the voice data includes multiple voice frequencies, the voice data is subjected to voice separation according to the preset frequency set and the voice frequencies of each part of the voice data to obtain multiple target voice data, wherein each preset frequency in the preset frequency set corresponds to a target user, and one target voice data comes from one target user;
[0008] According to the speech pause information in the target speech data and the target speech text corresponding to the target speech data, the target speech segmentation point of the target speech data is determined, and the target speech data is speech segmented according to the target speech segmentation point.
[0009] Furthermore, the method further comprises:
[0010] If the voice data includes a voice frequency, and the voice frequency included in the voice data belongs to a preset frequency set, the voice data is determined as target voice data.
[0011] Further, determining a target speech segmentation point of the target speech data according to speech pause information in the target speech data and a target speech text corresponding to the target speech data includes:
[0012] Determining, from the target speech data, a speech position at which a speech pause duration is greater than a preset duration threshold as a first speech segmentation point, and determining a segmentation probability corresponding to the first speech segmentation point based on a predetermined correspondence between speech pause duration and segmentation probability for a target user to which the target speech data belongs, wherein the speech pause information includes the speech pause duration;
[0013] Segmenting the target speech text to obtain a plurality of text sentences and a segmentation probability of each text sentence, determining an end position of a corresponding text sentence in the target speech data as a second speech segmentation point, and determining the segmentation probability of the corresponding text sentence as the segmentation probability of the second speech segmentation point;
[0014] The target speech segmentation point of the target speech data is determined according to the first speech segmentation point and the segmentation probability of the first speech segmentation point, the second speech segmentation point and the segmentation probability of the second speech segmentation point.
[0015] Furthermore, the target speech text is segmented to obtain multiple text sentences and the segmentation probability of each text sentence, including any of the following:
[0016] Input the target speech text into a pre-trained sentence segmentation model to obtain multiple text sentences and the segmentation probability of each text sentence;
[0017] If the target words exist in the target speech text, the target speech text is segmented based on the target words to obtain multiple text sentences, and the segmentation probability of each text sentence is determined based on a pre-set correspondence between the target words and the segmentation probability.
[0018] Further, determining a target speech segmentation point of the target speech data according to the first speech segmentation point and the segmentation probability of the first speech segmentation point, the second speech segmentation point and the segmentation probability of the second speech segmentation point includes:
[0019] When the first speech segmentation point and the second speech segmentation point are the same segmentation point, if the weighted sum of the segmentation probability of the first speech segmentation point and the segmentation probability of the second speech segmentation point is greater than a preset probability threshold, then one of the first speech segmentation point and the second speech segmentation point is determined as the target speech segmentation point;
[0020] When the first speech segmentation point and the second speech segmentation point are not the same segmentation point, if the product of the segmentation probability of the first speech segmentation point and the preset speech weight coefficient is greater than the preset speech segmentation probability, the corresponding first speech segmentation point is determined as the target speech segmentation point; if the product of the segmentation probability of the second speech segmentation point and the preset text weight coefficient is greater than the preset text segmentation probability, the corresponding second speech segmentation point is determined as the target speech segmentation point.
[0021] Furthermore, obtaining voice data includes:
[0022] Collecting original voice data, and performing a voice preprocessing step on the collected original voice data to obtain voice data;
[0023] The speech preprocessing step includes at least one of the following:
[0024] Deleting the speech portion of the original speech data whose corresponding amplitude does not fall within the preset amplitude range;
[0025] Deleting the speech portion of the original speech data whose corresponding frequency does not belong to the preset frequency range;
[0026] Perform content desensitization processing on the original voice data.
[0027] Furthermore, the method further comprises:
[0028] The target speech data after speech segmentation is translated into a target language to obtain a target translated text and user information of a target user corresponding to the target translated text and the target speech data.
[0029] A second aspect of the embodiments of the present application provides a speech segmentation device, including:
[0030] a data acquisition unit, configured to acquire voice data and determine whether the voice data includes multiple voice frequencies;
[0031] a data separation unit configured to, if the voice data includes multiple voice frequencies, perform voice separation on the voice data based on a preset frequency set and the voice frequencies of each portion of the voice data to obtain a plurality of target voice data, wherein each preset frequency in the preset frequency set corresponds to a target user, and one target voice data comes from one target user;
[0032] The segmentation execution unit is used to determine the target speech segmentation point of the target speech data according to the speech pause information in the target speech data and the target speech text corresponding to the target speech data, and to perform speech segmentation on the target speech data according to the target speech segmentation point.
[0033] Furthermore, the device further includes a target determination unit, configured to determine the voice data as target voice data if the voice data includes a voice frequency and the voice frequency included in the voice data belongs to a preset frequency set.
[0034] Furthermore, the split execution unit is specifically used for:
[0035] Determining, from the target speech data, a speech position at which a speech pause duration is greater than a preset duration threshold as a first speech segmentation point, and determining a segmentation probability corresponding to the first speech segmentation point based on a predetermined correspondence between speech pause duration and segmentation probability for a target user to which the target speech data belongs, wherein the speech pause information includes the speech pause duration;
[0036] Segmenting the target speech text to obtain a plurality of text sentences and a segmentation probability of each text sentence, determining an end position of a corresponding text sentence in the target speech data as a second speech segmentation point, and determining the segmentation probability of the corresponding text sentence as the segmentation probability of the second speech segmentation point;
[0037] The target speech segmentation point of the target speech data is determined according to the first speech segmentation point and the segmentation probability of the first speech segmentation point, the second speech segmentation point and the segmentation probability of the second speech segmentation point.
[0038] Furthermore, in the segmentation execution unit, the target speech text is segmented into sentences to obtain multiple text sentences and the segmentation probability of each text sentence, including any of the following:
[0039] Input the target speech text into a pre-trained sentence segmentation model to obtain multiple text sentences and the segmentation probability of each text sentence;
[0040] If the target words exist in the target speech text, the target speech text is segmented based on the target words to obtain multiple text sentences, and the segmentation probability of each text sentence is determined based on a pre-set correspondence between the target words and the segmentation probability.
[0041] Furthermore, in the segmentation execution unit, determining the target speech segmentation point of the target speech data according to the first speech segmentation point and the segmentation probability of the first speech segmentation point, the second speech segmentation point and the segmentation probability of the second speech segmentation point includes:
[0042] When the first speech segmentation point and the second speech segmentation point are the same segmentation point, if the weighted sum of the segmentation probability of the first speech segmentation point and the segmentation probability of the second speech segmentation point is greater than a preset probability threshold, then one of the first speech segmentation point and the second speech segmentation point is determined as the target speech segmentation point;
[0043] When the first speech segmentation point and the second speech segmentation point are not the same segmentation point, if the product of the segmentation probability of the first speech segmentation point and the preset speech weight coefficient is greater than the preset speech segmentation probability, the corresponding first speech segmentation point is determined as the target speech segmentation point; if the product of the segmentation probability of the second speech segmentation point and the preset text weight coefficient is greater than the preset text segmentation probability, the corresponding second speech segmentation point is determined as the target speech segmentation point.
[0044] Furthermore, in the data acquisition unit, acquiring the voice data includes:
[0045] Collecting original voice data, and performing a voice preprocessing step on the collected original voice data to obtain voice data;
[0046] The speech preprocessing step includes at least one of the following:
[0047] Deleting the speech portion of the original speech data whose corresponding amplitude does not fall within the preset amplitude range;
[0048] Deleting the speech portion of the original speech data whose corresponding frequency does not belong to the preset frequency range;
[0049] Perform content desensitization processing on the original voice data.
[0050] Furthermore, the device also includes a data translation unit for translating the target speech data after speech segmentation into a target language to obtain a target translation text and output the target translation text and user information of a target user corresponding to the target speech data.
[0051] A third aspect of an embodiment of the present application provides a server, comprising a memory, a processor, and a computer program stored in the memory and executable on the server. When the processor executes the computer program, the steps of the speech segmentation method provided in the first aspect are implemented.
[0052] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the speech segmentation method provided in the first aspect are implemented.
[0053] The implementation of a speech segmentation method, device, server, and storage medium provided in the embodiments of the present application has the following beneficial effects: it is possible to separate the speech data portion of the target user from the speech data to obtain the target speech data, and then, based on the speech pause information of the target speech data and the target speech text corresponding to the target speech data, determine the target speech segmentation point of the target speech data, and thereby segment the target speech data using the target speech segmentation point. This can automatically segment the speech data of the target user, and help improve the efficiency and accuracy of speech segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0055] Figure 1 This is a flow chart of an implementation method of a speech segmentation method provided in an embodiment of the present application;
[0056] Figure 2 This is a flowchart of another speech segmentation method provided in an embodiment of the present application;
[0057] Figure 3 This is a structural block diagram of a speech segmentation device provided in an embodiment of the present application;
[0058] Figure 4 This is a structural block diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0060] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0061] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0062] In the embodiment of the present application, artificial intelligence technology is used to automatically segment the voice data of the target user.
[0063] The speech segmentation method involved in the embodiment of the present application can be executed by a server. When the speech segmentation method is executed by a server, the execution subject is the server.
[0064] It should be noted that the aforementioned servers may include, but are not limited to, servers, mobile phones, tablets, or wearable smart devices. Furthermore, the aforementioned servers may be standalone servers or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0065] See also Figure 1 , Figure 1 The following is a flowchart of a speech segmentation method according to an embodiment of the present application, including:
[0066] Step 101: Acquire voice data and determine whether the voice data includes multiple voice frequencies.
[0067] Here, the execution subject may obtain the voice data locally or from another device connected to the communication link. Then, the execution subject may perform frequency analysis on the obtained voice data to determine whether the voice data includes multiple voice frequencies.
[0068] As an example, the execution entity may randomly obtain several speech segments from the speech data and analyze the speech frequencies of each segment. If there are speech segments with different speech frequencies, it can be determined that the speech data includes multiple speech frequencies. Conversely, if the speech frequencies of the speech segments are the same, it can be determined that the speech data includes only one speech frequency.
[0069] Step 102: If the voice data includes multiple voice frequencies, voice separation is performed on the voice data according to a preset frequency set and the voice frequencies of each part of the voice data to obtain multiple target voice data.
[0070] Each preset frequency in the preset frequency set corresponds to a target user, and each target voice data item comes from a target user. The target user is typically a pre-defined user. There can be one or more target users. By using the preset frequency set, only the target user's voice portion can be segmented.
[0071] Here, if the voice data includes multiple voice frequencies, the execution entity may determine a voice frequency belonging to a preset frequency set from the multiple voice frequencies included in the voice data. For ease of description, this may be referred to as a target voice frequency. Then, for each target voice frequency, the execution entity may separate the voice portion of the target voice frequency from the voice data to obtain the target voice data.
[0072] Here, for each target voice frequency, one target voice data can be obtained. For example, if there are three target voice frequencies, three target voice data can be obtained.
[0073] For example, the preset frequency set may include five preset frequencies, namely A, B, C, D, and E, each corresponding to a user. If the voice data includes three voice frequencies, namely A, B, and G, the execution entity can determine voice frequencies A and B belonging to the preset frequency set from these three voice frequencies. In this case, A and B are the target voice frequencies. The execution entity can then separate the voice portion containing target voice frequency A from the voice data to obtain target voice data for A, and separate the voice portion containing target voice frequency B from the voice data to obtain target voice data for B.
[0074] Step 103 : determining target speech segmentation points of the target speech data according to speech pause information in the target speech data and the target speech text corresponding to the target speech data, and performing speech segmentation on the target speech data according to the target speech segmentation points.
[0075] The target speech segmentation points are usually the locations for segmenting the target speech data. The target speech text is usually the target speech data in text form.
[0076] Here, for each target speech data, the execution entity may use the speech pause information in the target speech data and the target speech text corresponding to the target speech data to determine the target speech segmentation point of the target speech data. As an example, the execution entity may determine, from the target speech data, a speech position where the speech pause duration exceeds a preset duration threshold, as the target speech segmentation point. The target speech data is then segmented using the target speech segmentation point.
[0077] The method provided in this embodiment can separate the voice data portion of the target user from the voice data to obtain the target voice data. Then, based on the voice pause information of the target voice data and the target voice text corresponding to the target voice data, the target voice segmentation point of the target voice data is determined, and the target voice data is segmented using the target voice segmentation point. This can automatically segment the voice data of the target user, helping to improve the efficiency and accuracy of voice segmentation.
[0078] In an optional implementation of each embodiment of the present application, the above-mentioned acquisition of voice data may include: collecting original voice data, and performing a voice preprocessing step on the collected original voice data to obtain voice data.
[0079] The speech preprocessing step is usually a pre-set step for preprocessing speech data. The speech preprocessing step may include but is not limited to at least one of the following three items:
[0080] The first item is to delete the parts of the original speech data whose corresponding amplitude does not fall within the preset amplitude range. The preset amplitude range is usually a pre-set amplitude range. Here, deleting the parts with too low or too high amplitude can remove interfering noise.
[0081] The second option is to delete the speech portion of the original speech data whose corresponding frequency does not fall within a preset frequency range. The preset frequency range is usually a pre-set frequency range. Here, deleting the portion with too low or too high frequency can remove interfering noise.
[0082] The third item is to perform content desensitization on the original voice data. Here, content desensitization on the original voice data can protect user privacy.
[0083] In an optional implementation of each embodiment of the present application, determining the target speech segmentation point of the target speech data based on the speech pause information in the target speech data and the target speech text corresponding to the target speech data may include the following steps 1 to 3.
[0084] Step 1: Determine, from the target speech data, a speech position where the speech pause duration is greater than a preset duration threshold, as a first speech segmentation point, and determine the segmentation probability corresponding to the first speech segmentation point based on a predetermined correspondence between the speech pause duration and the segmentation probability for the target user to which the target speech data belongs.
[0085] The voice pause information includes the duration of the voice pause. The preset duration threshold is usually a pre-set value for indicating the duration. For example, it can be 2 seconds.
[0086] The above correspondence between speech pause duration and segmentation probability is generally used to characterize the correspondence between speech pause duration and segmentation probability. For example, if the speech pause duration is 2 seconds, the segmentation probability may be 0.2, and if the speech pause duration is 5 seconds, the segmentation probability may be 0.6.
[0087] In practice, since each target user speaks at a different speed, we can store the corresponding relationship between speech pause duration and segmentation probability for each target user. For example, the segmentation probability of a 3-second pause in Zhang San's speech data is 0.3, while the segmentation probability of a 3-second pause in Li Si's speech data may be 0.4.
[0088] Here, for each target voice data, the above-mentioned execution entity can determine, from the target voice data, a voice position where the voice pause duration is greater than a preset duration threshold, and the determined voice position can be denoted as the first voice segmentation point. Then, the above-mentioned execution entity can find the correspondence between the voice pause duration and the segmentation probability for the target user of the target voice data, and then find the segmentation probability corresponding to the first voice segmentation point from this correspondence. In practice, there are usually multiple first voice segmentation points, and for each first voice segmentation point, a segmentation probability can be found.
[0089] Step 2: Segment the target voice text into multiple text sentences and the segmentation probability of each text sentence, determine the end position of the corresponding text sentence in the target voice data as the second voice segmentation point, and determine the segmentation probability of the corresponding text sentence as the segmentation probability of the second voice segmentation point.
[0090] Here, the above-mentioned execution entity can segment the target voice text corresponding to the target voice data into multiple text sentences and obtain the segmentation probability corresponding to each text sentence. Then, for each text sentence, the above-mentioned execution entity can divide the end position of the text sentence into the second voice segmentation point, and use the segmentation probability of the text sentence as the segmentation probability of the second voice segmentation point.
[0091] For example, if the target voice text is "Could you help me carry this? I'm about to drop it.", segmenting this target voice text can obtain multiple text sentences, namely text sentence A: "Could you help me carry this?" and text sentence B: "I'm about to drop it.". If the segmentation probability of text sentence A is 0.8 and the segmentation probability of text sentence B is 0.9. At this time, the end position of text sentence A, that is, the position between "?" and "I", can be divided into a second voice segmentation point P1, and the position after "it" can be divided into another second voice segmentation point P2. At this time, the segmentation probability of the second voice segmentation point P1 is 0.8, and the segmentation probability of the second voice segmentation point P2 is 0.9.
[0092] Optionally, the above-mentioned segmentation of the target voice text into multiple text sentences and the segmentation probability of each text sentence can include any one of the following method 1 and method 2.
[0093] Method 1: Input the target voice text into a pre-trained sentence segmentation model to obtain multiple text sentences and the segmentation probability of each text sentence.
[0094] Here, the sentence segmentation model can be used to represent the corresponding relationship between the speech text, the text sentences, and the segmentation probabilities of the text sentences. Specifically, the sentence segmentation model can be a correspondence table generated based on the statistics of a large number of speech texts and storing the corresponding relationships between multiple speech texts, text sentences, and the segmentation probabilities of the text sentences, or it can be a model obtained by training an initial model (such as a Convolutional Neural Network (CNN), a ResNet, etc.) using machine learning methods based on training samples.
[0095] In the second method, if there are target words in the target speech text, the target speech text is segmented with the target words as separators to obtain multiple text sentences, and the segmentation probabilities of each text sentence are determined according to the pre-set corresponding relationship between the target words and the segmentation probabilities.
[0096] Among them, the above-mentioned target words are usually pre-set words. In practice, the target words are usually onomatopoeic words, such as "ne", "ma", etc. In addition, the target words can also be other words, such as "over", "end", "start", etc. The corresponding relationship between the above-mentioned target words and the segmentation probabilities is usually the pre-stored corresponding relationship between the target words and the segmentation probabilities.
[0097] Step 3: Determine the target speech segmentation point of the target speech data according to the first speech segmentation point and the segmentation probability of the first speech segmentation point, the second speech segmentation point and the segmentation probability of the second speech segmentation point.
[0098] As an example, the above-mentioned execution entity can use all the first speech segmentation points greater than a certain probability value as the target speech segmentation points, and can use all the second speech segmentation points greater than a certain probability value as the target speech segmentation points.
[0099] In this implementation, for each target speech data, the target speech data and the target speech text corresponding to the target speech data are analyzed simultaneously to determine the target speech segmentation point for segmenting the target speech data, which can achieve analysis from multiple perspectives, help extract accurate speech segmentation points, and thus help achieve more accurate speech segmentation.
[0100] In some optional implementation manners, in the above Step 3, the determining the target speech segmentation point of the target speech data according to the first speech segmentation point and the segmentation probability of the first speech segmentation point, the second speech segmentation point and the segmentation probability of the second speech segmentation point may include:
[0101] First, when the first speech segmentation point and the second speech segmentation point are the same segmentation point, if the weighted sum of the segmentation probability of the first speech segmentation point and the segmentation probability of the second speech segmentation point is greater than a preset probability threshold, then one of the first speech segmentation point and the second speech segmentation point is determined as the target speech segmentation point.
[0102] The above-mentioned preset probability threshold is usually a preset probability lower limit value, for example, it may be 0.6.
[0103] Here, if the first speech segmentation point and the second speech segmentation point simultaneously point to the same segmentation location, the execution entity may calculate a weighted sum based on the segmentation probability of the first speech segmentation point, the segmentation probability of the second speech segmentation point, a preset speech weight coefficient, and a preset text weight coefficient. If the weighted sum is greater than a preset probability threshold, the location is considered to be the target speech segmentation point.
[0104] For example, if for a certain target speech data, there are three first speech segmentation points, namely P1, P2 and P3, and the segmentation probabilities of the three first speech segmentation points are X1, X2 and X3 respectively, and there are two second speech segmentation points, namely P4 and P5, and the segmentation probabilities of the two second speech segmentation points are X4 and X5 respectively. The preset speech weight coefficient is W1, and the preset text weight coefficient is W2. If P1 and P4 point to the same segmentation position at the same time, the weighted sum value for P1 and P4 can be calculated, and the weighted sum value = X1×W1+X4×W2. Afterwards, the weighted sum value can be compared with the preset probability threshold. If it is greater than the preset probability threshold, P1 or P4 is determined as the target speech segmentation point.
[0105] Then, when the first speech segmentation point and the second speech segmentation point are not the same segmentation point, if the product of the segmentation probability of the first speech segmentation point and the preset speech weight coefficient is greater than the preset speech segmentation probability, the corresponding first speech segmentation point is determined as the target speech segmentation point; if the product of the segmentation probability of the second speech segmentation point and the preset text weight coefficient is greater than the preset text segmentation probability, the corresponding second speech segmentation point is determined as the target speech segmentation point.
[0106] The above-mentioned preset speech segmentation probability is usually a preset probability value. The above-mentioned preset text segmentation probability is usually a preset probability value.
[0107] Here, if the first speech segmentation point and the second speech segmentation point point to different positions, the execution entity may calculate the product of the segmentation probability of the first speech segmentation point and the preset speech weight coefficient, recorded as a first product value, and determine whether the first speech segmentation point is the target speech segmentation point by comparing the first product value with the preset speech segmentation probability. Furthermore, the execution entity may calculate the product of the segmentation probability of the second speech segmentation point and the preset text weight coefficient, recorded as a second product value, and determine whether the second speech segmentation point is the target speech segmentation point by comparing the second product value with the preset text segmentation probability.
[0108] For example, if there are three first speech segmentation points for a certain target speech data, namely P1, P2, and P3, the segmentation probabilities of the three first speech segmentation points are X1, X2, and X3, respectively. There are two second speech segmentation points, namely P4 and P5, and the segmentation probabilities of the two second speech segmentation points are X4 and X5, respectively. The preset speech weight coefficient is W1, and the preset text weight coefficient is W2. If P1 and P4 point to the same segmentation position at the same time, and the other speech segmentation points point to different positions, then for either P2 or P3, a first product value can be calculated, such as the first product value for P2 = X2 × W1, and the first product value for P3 = X3 × W1. Then, the first product value can be compared with the preset speech segmentation probabilities to determine whether P2 or P3 is the target speech segmentation point. In addition, a second product value can be calculated for P5, such as the second product value for P5 = X5 × W2. Afterwards, the second product value may be compared with a preset text segmentation probability to determine whether P5 is a target speech segmentation point.
[0109] This implementation method can accurately determine the target speech segmentation points in each target speech data, thereby achieving accurate segmentation of the target speech data.
[0110] See also Figure 2 , Figure 2 This is a flow chart of an implementation of a speech segmentation method provided in an embodiment of the present application. The speech segmentation method provided in this embodiment may include the following steps:
[0111] Step 201: Acquire voice data and determine whether the voice data includes multiple voice frequencies.
[0112] Step 202: If the voice data includes multiple voice frequencies, voice separation is performed on the voice data according to a preset frequency set and the voice frequencies of each part of the voice data to obtain multiple target voice data.
[0113] Each preset frequency in the preset frequency set corresponds to a target user, and one target voice data comes from one target user.
[0114] In this embodiment, the specific operations of steps 201-202 are the same as those in Figure 1 The operations of steps 101-102 in the illustrated embodiment are substantially the same and will not be described in detail here.
[0115] Step 203: If the voice data includes a voice frequency, and the voice frequency included in the voice data belongs to a preset frequency set, the voice data is determined as target voice data.
[0116] Here, when the voice data includes only one voice frequency, if the voice frequency belongs to a preset frequency set, the execution subject may directly determine the voice data as the target voice data.
[0117] It should be pointed out that when there is only one voice frequency, the voice data can be directly determined as the target voice data, eliminating the voice separation step, which helps to further improve data processing efficiency.
[0118] Step 204 : determining target speech segmentation points of the target speech data according to speech pause information in the target speech data and the target speech text corresponding to the target speech data, and performing speech segmentation on the target speech data according to the target speech segmentation points.
[0119] In this embodiment, the specific operation of step 204 is the same as Figure 1 The operation of step 103 in the illustrated embodiment is substantially the same and will not be described again here.
[0120] The method provided in this embodiment can directly determine the voice data as the target voice data when there is only one voice frequency, eliminating the voice separation step. The flexibility of voice data processing is good, which helps to further improve data processing efficiency.
[0121] In the optional implementation of each embodiment of the present application, the above-mentioned speech segmentation method may further include the following steps: translating the target speech data after speech segmentation into a target language to obtain a target translation text, and outputting the target translation text and user information of the target user corresponding to the target speech data.
[0122] The target language may be any of a variety of pre-set languages, such as English, French, Japanese, etc. The target user's user information may include, but is not limited to, a user name, nickname, etc.
[0123] Here, for each target voice data, after accurately segmenting the target voice data, the segmented target voice data is used for translation, which can make the target translation text more accurate. In addition, the target translation text is output in correspondence with the user information of the target user corresponding to the target voice data, which can achieve an intuitive presentation of the target translation text for each target user, which is more practical and helps to improve the user experience.
[0124] See also Figure 3 , Figure 3 This is a structural block diagram of a speech segmentation device 300 provided in an embodiment of the present application. In this embodiment, the speech segmentation device includes various units for performing Figure 1-Figure 2 For details, please refer to the steps in the corresponding embodiment. Figure 1-Figure 2 as well as Figure 1-Figure 2 For the convenience of explanation, only the parts related to this embodiment are shown. Figure 3 , the speech segmentation device 300 includes:
[0125] The data acquisition unit 301 is used to acquire voice data and determine whether the voice data includes multiple voice frequencies;
[0126] The data separation unit 302 is configured to, if the voice data includes multiple voice frequencies, perform voice separation on the voice data based on a preset frequency set and the voice frequencies of each portion of the voice data to obtain multiple target voice data, wherein each preset frequency in the preset frequency set corresponds to a target user, and each target voice data comes from one target user;
[0127] The segmentation execution unit 303 is used to determine the target speech segmentation points of the target speech data according to the speech pause information in the target speech data and the target speech text corresponding to the target speech data, and to perform speech segmentation on the target speech data according to the target speech segmentation points.
[0128] As an embodiment of the present application, the device further includes a target determination unit (not shown in the figure). The target determination unit is configured to determine the voice data as target voice data if the voice data includes a voice frequency and the voice frequency included in the voice data belongs to a preset frequency set.
[0129] As an embodiment of the present application, the split execution unit 303 is specifically configured to:
[0130] Determining, from the target speech data, a speech position at which a speech pause duration is greater than a preset duration threshold as a first speech segmentation point, and determining a segmentation probability corresponding to the first speech segmentation point based on a predetermined correspondence between speech pause duration and segmentation probability for a target user to which the target speech data belongs, wherein the speech pause information includes the speech pause duration;
[0131] Segmenting the target speech text to obtain a plurality of text sentences and a segmentation probability of each text sentence, determining an end position of a corresponding text sentence in the target speech data as a second speech segmentation point, and determining the segmentation probability of the corresponding text sentence as the segmentation probability of the second speech segmentation point;
[0132] The target speech segmentation point of the target speech data is determined according to the first speech segmentation point and the segmentation probability of the first speech segmentation point, the second speech segmentation point and the segmentation probability of the second speech segmentation point.
[0133] As an embodiment of the present application, in the segmentation execution unit 303, the target speech text is segmented into sentences to obtain multiple text sentences and the segmentation probability of each text sentence, including any of the following:
[0134] Input the target speech text into a pre-trained sentence segmentation model to obtain multiple text sentences and the segmentation probability of each text sentence;
[0135] If the target words exist in the target speech text, the target speech text is segmented based on the target words to obtain multiple text sentences, and the segmentation probability of each text sentence is determined based on a pre-set correspondence between the target words and the segmentation probability.
[0136] As an embodiment of the present application, in the segmentation execution unit 303, determining the target speech segmentation point of the target speech data according to the first speech segmentation point and the segmentation probability of the first speech segmentation point, the second speech segmentation point and the segmentation probability of the second speech segmentation point includes:
[0137] When the first speech segmentation point and the second speech segmentation point are the same segmentation point, if the weighted sum of the segmentation probability of the first speech segmentation point and the segmentation probability of the second speech segmentation point is greater than a preset probability threshold, then one of the first speech segmentation point and the second speech segmentation point is determined as the target speech segmentation point;
[0138] When the first speech segmentation point and the second speech segmentation point are not the same segmentation point, if the product of the segmentation probability of the first speech segmentation point and the preset speech weight coefficient is greater than the preset speech segmentation probability, the corresponding first speech segmentation point is determined as the target speech segmentation point; if the product of the segmentation probability of the second speech segmentation point and the preset text weight coefficient is greater than the preset text segmentation probability, the corresponding second speech segmentation point is determined as the target speech segmentation point.
[0139] As an embodiment of the present application, in the data acquisition unit 301, acquiring voice data includes:
[0140] Collecting original voice data, and performing a voice preprocessing step on the collected original voice data to obtain voice data;
[0141] The speech preprocessing step includes at least one of the following:
[0142] Deleting the speech portion of the original speech data whose corresponding amplitude does not fall within the preset amplitude range;
[0143] Deleting the speech portion of the original speech data whose corresponding frequency does not belong to the preset frequency range;
[0144] Perform content desensitization processing on the original voice data.
[0145] As an embodiment of the present application, the device further includes a data translation unit (not shown in the figure). The data translation unit is used to translate the target speech data after speech segmentation into a target language to obtain a target translated text and user information of a target user corresponding to the target translated text and the target speech data.
[0146] The device provided in this embodiment can separate the voice data portion of the target user from the voice data to obtain the target voice data. Then, based on the voice pause information of the target voice data and the target voice text corresponding to the target voice data, the target voice segmentation point of the target voice data is determined, and the target voice data is segmented using the target voice segmentation point. This can automatically segment the voice data of the target user, helping to improve the efficiency and accuracy of voice segmentation.
[0147] It should be understood that Figure 3 In the structural block diagram of the speech segmentation device shown, each unit is used to perform Figure 1-Figure 2 The steps in the corresponding embodiments, and Figure 1-Figure 2 Each step in the corresponding embodiment has been explained in detail in the above embodiment. Figure 1-Figure 2 as well as Figure 1-Figure 2 The relevant descriptions in the corresponding embodiments will not be repeated here.
[0148] Figure 4 This is a structural block diagram of a server provided by another embodiment of the present application. Figure 4As shown, the server 400 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401, such as a program for a speech segmentation method. When the processor 401 executes the computer program 403, the steps in each embodiment of the above-mentioned speech segmentation method are implemented, such as Figure 1 Alternatively, the processor 401 executes the computer program 403 to implement the above steps 101 to 103. Figure 3 The functions of each unit in the corresponding embodiment are, for example, Figure 3 For details on the functions of units 301 to 303, please refer to Figure 3 The relevant descriptions in the corresponding embodiments are not repeated here.
[0149] Exemplarily, computer program 403 may be divided into one or more units, one or more of which are stored in memory 402 and executed by processor 401 to implement the present application. One or more units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of computer program 403 in server 400. For example, computer program 403 may be divided into a data acquisition unit, a data separation unit, and a segmentation execution unit, with the specific functions of each unit being as described above.
[0150] The server may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4 This is only an example of the server 400 and does not constitute a limitation on the server 400. The server 400 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the turntable device may also include input and output devices, network access devices, buses, etc.
[0151] The processor 401 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0152] Memory 402 can be an internal storage unit of server 400, such as a hard drive or memory of server 400. Memory 402 can also be an external storage device of server 400, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on server 400. Furthermore, memory 402 can include both an internal storage unit of server 400 and an external storage device. Memory 402 is used to store computer programs and other programs and data required by the turntable device. Memory 402 can also be used to temporarily store data that has been output or is about to be output.
[0153] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0154] If the integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable storage medium may include: any entity or device that can carry computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in computer-readable storage media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunications signals.
[0155] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A speech segmentation method, characterized in that: The method comprises: Acquiring voice data, and determining whether the voice data includes multiple voice frequencies; If the voice data includes multiple voice frequencies, performing voice separation on the voice data according to a preset frequency set and the voice frequencies of each part of the voice data to obtain multiple target voice data, wherein each preset frequency in the preset frequency set corresponds to a target user, and one target voice data comes from one target user; Determining a target speech segmentation point of the target speech data according to speech pause information in the target speech data and a target speech text corresponding to the target speech data, and performing speech segmentation on the target speech data according to the target speech segmentation point; The step of determining the target speech segmentation point of the target speech data according to the speech pause information in the target speech data and the target speech text corresponding to the target speech data includes: Determining, from the target speech data, a speech position at which a speech pause duration is greater than a preset duration threshold as a first speech segmentation point, and determining a segmentation probability corresponding to the first speech segmentation point based on a predetermined correspondence between speech pause duration and segmentation probability for a target user to which the target speech data belongs, wherein the speech pause information includes speech pause duration; Segmenting the target speech text to obtain a plurality of text sentences and a segmentation probability of each text sentence, determining an end position of a corresponding text sentence in the target speech data as a second speech segmentation point, and determining the segmentation probability of the corresponding text sentence as the segmentation probability of the second speech segmentation point; When the first speech segmentation point and the second speech segmentation point are the same segmentation point, if a weighted sum of the segmentation probability of the first speech segmentation point and the segmentation probability of the second speech segmentation point is greater than a preset probability threshold, determining one of the first speech segmentation point and the second speech segmentation point as the target speech segmentation point; When the first speech segmentation point and the second speech segmentation point are not the same segmentation point, if the product of the segmentation probability of the first speech segmentation point and the preset speech weight coefficient is greater than the preset speech segmentation probability, the corresponding first speech segmentation point is determined as the target speech segmentation point; if the product of the segmentation probability of the second speech segmentation point and the preset text weight coefficient is greater than the preset text segmentation probability, the corresponding second speech segmentation point is determined as the target speech segmentation point.
2. The speech segmentation method according to claim 1, wherein: The method further comprises: If the voice data includes a voice frequency, and the voice frequency included in the voice data belongs to the preset frequency set, the voice data is determined as target voice data.
3. The speech segmentation method according to claim 1, wherein: The sentence segmentation of the target speech text to obtain a plurality of text sentences and a segmentation probability of each text sentence includes any of the following: Inputting the target speech text into a pre-trained sentence segmentation model to obtain multiple text sentences and the segmentation probability of each text sentence; If the target speech text contains a target word, the target speech text is segmented using the target word as a separator to obtain a plurality of text sentences, and the segmentation probability of each text sentence is determined based on a pre-set correspondence between the target word and the segmentation probability.
4. The speech segmentation method according to claim 1, wherein: The acquiring of voice data comprises: Collecting original voice data, and performing a voice preprocessing step on the collected original voice data to obtain the voice data; The speech preprocessing step includes at least one of the following: Deleting a speech portion of the original speech data whose corresponding amplitude does not fall within a preset amplitude range; Deleting a speech portion of the original speech data whose corresponding frequency does not belong to a preset frequency interval; Perform content desensitization processing on the original voice data.
5. The speech segmentation method according to any one of claims 1 to 4, characterized in that: The method further comprises: The target speech data after speech segmentation is translated into a target language to obtain a target translation text, and the target translation text and user information of a target user corresponding to the target speech data are output accordingly.
6. A speech segmentation device, characterized in that: The device comprises: a data acquisition unit, configured to acquire voice data and determine whether the voice data includes multiple voice frequencies; a data separation unit configured to, if the voice data includes multiple voice frequencies, perform voice separation on the voice data based on a preset frequency set and the voice frequencies of each portion of the voice data to obtain a plurality of target voice data, wherein each preset frequency in the preset frequency set corresponds to a target user, and one target voice data comes from one target user; a segmentation execution unit, configured to determine a target speech segmentation point of the target speech data according to speech pause information in the target speech data and a target speech text corresponding to the target speech data, and perform speech segmentation on the target speech data according to the target speech segmentation point; The split execution unit is specifically used to: Determining, from the target speech data, a speech position at which a speech pause duration is greater than a preset duration threshold as a first speech segmentation point, and determining a segmentation probability corresponding to the first speech segmentation point based on a predetermined correspondence between speech pause duration and segmentation probability for a target user to which the target speech data belongs, wherein the speech pause information includes speech pause duration; Segmenting the target speech text to obtain a plurality of text sentences and a segmentation probability of each text sentence, determining an end position of a corresponding text sentence in the target speech data as a second speech segmentation point, and determining the segmentation probability of the corresponding text sentence as the segmentation probability of the second speech segmentation point; When the first speech segmentation point and the second speech segmentation point are the same segmentation point, if a weighted sum of the segmentation probability of the first speech segmentation point and the segmentation probability of the second speech segmentation point is greater than a preset probability threshold, determining one of the first speech segmentation point and the second speech segmentation point as the target speech segmentation point; When the first speech segmentation point and the second speech segmentation point are not the same segmentation point, if the product of the segmentation probability of the first speech segmentation point and the preset speech weight coefficient is greater than the preset speech segmentation probability, the corresponding first speech segmentation point is determined as the target speech segmentation point; if the product of the segmentation probability of the second speech segmentation point and the preset text weight coefficient is greater than the preset text segmentation probability, the corresponding second speech segmentation point is determined as the target speech segmentation point.
7. A server comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Voice separation method and device, electronic equipment, and computer readable storage medium
CN110718228A
VAD tail point detection method and device, server and computer readable medium
CN111627423A