A data processing method, apparatus and electronic device

By identifying the first and second voice segments in the speech data and determining the punctuation based on the duration threshold and the punctuation recognition rules, the problem of poor punctuation recognition accuracy in the text conversion of the speech data is solved, and the efficiency and accuracy of punctuation recognition are improved.

CN114038466BActive Publication Date: 2025-07-04ZHEJIANG DASOUCHE SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111114057.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-23
Publication Date
2025-07-04
Estimated Expiration
2041-09-23

AI Technical Summary

Technical Problem

In the prior art, the accuracy of punctuation recognition during the text conversion of speech data is poor, mainly because the punctuation itself does not pronounce and the acoustic characteristics of different people are large, resulting in insufficient recognition accuracy based on acoustic feature modeling.

Method used

By identifying the first voice segment and the second voice segment in the voice data, punctuation points are determined using the preset duration threshold and punctuation recognition rules. The first voice segment contains vocal data and the second voice segment does not contain vocal data. The first punctuation point is determined when the duration of the second voice segment is less than the threshold. When the duration is not less than the threshold, the second punctuation point is determined based on the adjacent voice segment and the preset rules to realize text conversion.

Benefits of technology

The accuracy and efficiency of punctuation recognition are improved, the recognition accuracy problems caused by differences in acoustic characteristics are avoided, and the punctuation recognition effect in speech data text conversion is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114038466B_ABST
    Figure CN114038466B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a data processing method, apparatus and electronic device. The method includes: obtaining target speech data to be recognized, and determining a first speech segment and a second speech segment in the target speech data; when the duration of the second speech segment is less than a first preset duration threshold, determining that the punctuation corresponding to the second speech segment is a first punctuation; when the duration of the second speech segment is not less than the first preset duration threshold, determining the punctuation corresponding to the second speech segment in a second punctuation based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule; converting the first speech segment into target text data, and performing text conversion on the target speech data based on the target text data and the punctuation corresponding to the second speech segment to obtain text conversion data including punctuation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a data processing method, apparatus, and electronic device. Background Art

[0002] With the rapid development of communication technology and information processing technology, and the increasing computing power of devices, the application of speech recognition technology is becoming more and more extensive, such as simultaneous interpretation, speech transcription, human-computer interaction, voice control, etc.

[0003] During the process of recognizing speech data, in order to recognize the punctuation in the speech data, acoustic features such as the voice and intonation of the speaker can be modeled to recognize the punctuation in the speech data. However, since punctuation itself has no pronunciation and only appears as silence in acoustic features, and the acoustic features of different people are quite different, this leads to the problem of poor recognition accuracy in the method of modeling based on acoustic features to recognize the punctuation in speech data. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a data processing method, apparatus, and electronic device to solve the problem of poor punctuation recognition accuracy in the process of text conversion of speech data in the prior art.

[0005] To solve the above technical problems, the embodiments of the present invention are implemented as follows:

[0006] In a first aspect, a data processing method provided by an embodiment of the present invention includes:

[0007] Obtain target speech data to be recognized, and determine a first speech segment and a second speech segment in the target speech data, where the first speech segment is a segment of speech data containing human voice data, and the second speech segment is a segment of speech data not containing human voice data;

[0008] When the duration of the second speech segment is less than a first preset duration threshold, determine that the punctuation corresponding to the second speech segment is a first punctuation, and the first punctuation is used to represent a sentence pause;

[0009] When the duration of the second speech segment is not less than the first preset duration threshold, based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule, determine the punctuation corresponding to the second speech segment in the second punctuation, and the second punctuation is used to represent the end of a sentence;

[0010] Convert the first speech segment into target text data, and perform text conversion on the target speech data based on the target text data and the punctuation corresponding to the second speech segment to obtain text conversion data including punctuation.

[0011] Optionally, determining the punctuation in the second punctuation corresponding to the second speech segment based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule includes:

[0012] Determine the punctuation in the second punctuation corresponding to the second speech segment based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model, where the punctuation recognition model is trained based on a model constructed by a preset machine learning algorithm using historical first speech segment pairs.

[0013] Optionally, before determining the punctuation in the second punctuation corresponding to the second speech segment based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model, it further includes:

[0014] Obtain the historical first speech segments within a preset training period;

[0015] Perform supervised training on a preset neural network model based on the historical first speech segments and corresponding labels to obtain the pre-trained punctuation recognition model, where the label corresponding to the historical first speech segment is the punctuation in the second punctuation corresponding to the historical first speech segment.

[0016] Optionally, before determining the punctuation in the second punctuation corresponding to the second speech segment based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model, it further includes:

[0017] Obtain the historical first speech segments within a preset training period;

[0018] Convert the historical first speech segments into historical text data based on a preset speech recognition algorithm;

[0019] Perform supervised training on a preset neural network model based on the historical text data and corresponding labels to obtain the pre-trained punctuation recognition model, where the label corresponding to the historical text data is the punctuation in the second punctuation corresponding to the historical text data;

[0020] Determining the punctuation in the second punctuation corresponding to the second speech segment based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model includes:

[0021] Based on the preset speech recognition algorithm, convert the target speech segment adjacent to the second speech segment in the first speech segment into first target text data;

[0022] Based on the first target text data and the pre-trained punctuation recognition model, determine the punctuation corresponding to the second speech segment in the second punctuation.

[0023] Optionally, before converting the first speech segment into target text data, it further includes:

[0024] If there is a target second speech segment with a duration less than the second preset duration threshold in the second speech segment, and there are two first speech segments adjacent to the target second speech segment in the first speech segment, then merge the two first speech segments and the target second speech segment into the first speech segment, where the second preset duration threshold is less than the first preset duration threshold.

[0025] In a second aspect, an embodiment of the present invention provides a data processing device, and the device includes:

[0026] A first acquisition module, configured to acquire target speech data to be recognized, and determine a first speech segment and a second speech segment in the target speech data, where the first speech segment is a segment of speech data containing human voice data, and the second speech segment is a segment of speech data not containing human voice data;

[0027] A first determination module, configured to determine that the punctuation corresponding to the second speech segment is a first punctuation when the duration of the second speech segment is less than the first preset duration threshold, and the first punctuation is used to represent a sentence pause;

[0028] A second determination module, configured to, when the duration of the second speech segment is not less than the first preset duration threshold, based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule, determine the punctuation corresponding to the second speech segment in the second punctuation, and the second punctuation is used to represent the end of a sentence;

[0029] A recognition module, configured to convert the first speech segment into target text data, and perform text conversion on the target speech data based on the target text data and the punctuation corresponding to the second speech segment to obtain text conversion data containing punctuation.

[0030] Optionally, the second determination module is configured to:

[0031] Based on the target speech segment adjacent to the second speech segment in the first speech segment, and a pre-trained punctuation recognition model, determine the punctuation corresponding to the second speech segment in the second punctuation. The punctuation recognition model is obtained by training a model constructed by a preset machine learning algorithm based on historical first speech segment pairs.

[0032] Optionally, the apparatus further includes:

[0033] A second acquisition module, configured to acquire the historical first speech segments within a preset training period;

[0034] A first training module, configured to perform supervised training on a preset neural network model based on the historical first speech segments and corresponding labels, to obtain the pre-trained punctuation recognition model. The label corresponding to the historical first speech segment is the punctuation corresponding to the historical first speech segment in the second punctuation.

[0035] Optionally, the apparatus further includes:

[0036] A third acquisition module, configured to acquire the historical first speech segments within a preset training period;

[0037] A first conversion module, configured to convert the historical first speech segments into historical text data based on a preset speech recognition algorithm;

[0038] A second training module, configured to perform supervised training on a preset neural network model based on the historical text data and corresponding labels, to obtain the pre-trained punctuation recognition model. The label corresponding to the historical text data is the punctuation corresponding to the historical text data in the second punctuation;

[0039] The second determination module is configured to:

[0040] Based on the preset speech recognition algorithm, convert the target speech segment adjacent to the second speech segment in the first speech segment into first target text data;

[0041] Based on the first target text data and the pre-trained punctuation recognition model, determine the punctuation corresponding to the second speech segment in the second punctuation.

[0042] Optionally, the apparatus further includes:

[0043] A merging module, configured to, if there is a target second speech segment in the second speech segment whose duration is less than a second preset duration threshold, and there are two first speech segments adjacent to the target second speech segment in the first speech segment, merge the two first speech segments and the target second speech segment into the first speech segment, where the second preset duration threshold is less than the first preset duration threshold.

[0044] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the data processing method provided in the above embodiment are implemented.

[0045] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data processing method provided in the above embodiment are implemented.

[0046] As can be seen from the technical solutions provided in the embodiments of the present invention above, in the embodiments of the present invention, by obtaining target speech data to be recognized, and determining a first speech segment and a second speech segment in the target speech data, where the first speech segment is a segment of speech data containing human voice data, and the second speech segment is a segment of speech data not containing human voice data. When the duration of the second speech segment is less than a first preset duration threshold, determining that the punctuation corresponding to the second speech segment is a first punctuation, where the first punctuation is used to represent a sentence pause. When the duration of the second speech segment is not less than the first preset duration threshold, based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule, determining the punctuation corresponding to the second speech segment in a second punctuation, where the second punctuation is used to represent the end of a sentence. Converting the first speech segment into target text data, and performing text conversion on the target speech data based on the target text data and the punctuation corresponding to the second speech segment to obtain text conversion data including punctuation. In this way, the punctuation corresponding to the second speech segment not containing human voice data can be determined according to its duration, which can improve the efficiency of punctuation recognition. In addition, when the duration of the second speech segment is not less than the first preset time threshold, the punctuation corresponding to the second speech segment can be determined based on the preset punctuation recognition rule and the target speech segment adjacent to the second speech segment, avoiding the problem of poor punctuation recognition accuracy caused by the large difference in acoustic features corresponding to different human voice data, and improving the accuracy of punctuation recognition in the process of text conversion for the target speech data. Description of the Drawings

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0048] Figure 1 Flow schematic diagram of a data processing method of the present invention;

[0049] Figure 2 Schematic diagram of a target voice data of the present invention;

[0050] Figure 3 Flow schematic diagram of another data processing method of the present invention;

[0051] Figure 4 Schematic diagram of a processed target voice data of the present invention;

[0052] Figure 5 Structural schematic diagram of a data processing device of the present invention;

[0053] Figure 6 Structural schematic diagram of an electronic device of the present invention. Detailed implementation manners

[0054] The embodiments of the present invention provide a data processing method, device and electronic device.

[0055] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] Embodiment 1

[0057] As Figure 1 shown, the embodiments of the present invention provide a data processing method. The execution subject of this method can be a terminal device or a server. The server can be an independent server or a server cluster composed of multiple servers. This method can specifically include the following steps:

[0058] In S102, obtain the target voice data to be recognized, and determine the first voice segment and the second voice segment in the target voice data.

[0059] Among them, the first voice segment can be a segment of voice data containing human voice data, the second voice segment can be a segment of voice data without human voice data, and the target voice data can be any voice data to be recognized. For example, the target voice data can be the user voice data collected by a smart home device, or the target voice data can also be the voice data that has been sent or received in an instant messaging application, or the target voice data can also be any one or more segments of voice data contained in a certain video data. The embodiments of the present invention do not make specific limitations on the acquisition scenario of the target voice data, the duration of the target voice data, etc.

[0060] In practice, with the rapid development of communication technology and information processing technology, and the increasing computing power of devices, the application of speech recognition technology is becoming more and more extensive, such as: simultaneous interpretation, speech transcription, human-computer interaction, voice control, etc. In the process of recognizing voice data, in order to recognize the punctuation in the voice data, the acoustic features such as the voice and intonation of the speaker can be modeled to recognize the punctuation in the voice data. However, since the punctuation itself has no pronunciation and only appears as silence in the acoustic features, and the acoustic features of different people are quite different, this leads to the problem of poor recognition accuracy in the method of modeling based on acoustic features to recognize the punctuation in voice data. For this reason, the embodiments of the present invention provide a technical solution that can solve the above problems, which can specifically include the following content:

[0061] After obtaining the target voice data to be recognized, the first voice segment and the second voice segment in the target voice data can be determined through a preset voice data preprocessing method, where the preset voice preprocessing method can be Voice Activity Detection (VAD).

[0062] For example, the target voice data can be a segment of voice data received in an instant messaging application, and the VAD method can be used to process the target voice data. The processed target voice data can be as Figure 2 shown. The target voice data can include three first voice segments and three second voice segments.

[0063] It is possible to determine whether to continue executing S104 or continue executing S106 after S102 according to the relationship between the duration of the second voice segment and the first preset duration threshold.

[0064] Among them, the first preset duration threshold can be any threshold. For example, the first preset duration threshold can be 350 ms, etc. The first preset duration threshold can vary according to different actual application scenarios. For example, assuming that the target voice data to be recognized is the voice data to be recognized in an instant messaging application program, then the first preset duration threshold can be different thresholds set by the target user of the instant messaging application program for different contacts. For example, for a contact with a faster speaking speed, the target user can set the first preset duration threshold to a smaller threshold; for a contact with a slower speaking speed, the target user can set the first preset duration threshold to a larger threshold. In this way, the first preset duration threshold can be set flexibly according to different actual application scenarios to more accurately recognize the punctuation marks contained in the target voice data. The embodiments of the present invention do not make specific limitations on the first preset duration threshold.

[0065] The duration of each voice segment (i.e., the first voice segment and the second voice segment) can be calculated according to the number of frames contained therein. For example, it can be preset that the duration of one frame of voice is 50 ms. If a voice segment contains 10 frames, then it can be determined that the duration of this voice segment is 500 ms.

[0066] In S104, when the duration of the second voice segment is less than the first preset duration threshold, it is determined that the punctuation mark corresponding to the second voice segment is the first punctuation mark.

[0067] Among them, the first punctuation mark is used to indicate a pause in a sentence. For example, the first punctuation mark can be a comma, a semicolon, or a colon, etc., which are punctuation marks used to indicate a general pause in a sentence.

[0068] In implementation, different first punctuation marks can be preset according to different application scenarios. For example, if the target voice data is the voice data received in an instant messaging application program, then the first punctuation mark in this scenario can be preset as a comma. That is, when the duration of the second voice segment is less than the first preset duration threshold, it can be determined that the punctuation mark corresponding to this second voice segment is a comma.

[0069] Or, the punctuation mark corresponding to the first voice segment can also be determined according to the duration, quantity, etc. of the first voice segment adjacent to the second voice segment. For example, as Figure 2 shown, assuming that the durations of the first voice segment 1 and the first voice segment 2 are equal, and the duration of the second voice segment 1 is less than the first preset duration threshold, then it can be determined that the punctuation mark corresponding to the second voice segment 1 is a semicolon or a colon.

[0070] Or, the punctuation mark corresponding to the first voice segment can also be determined by applying the application scenario, as well as the duration, quantity, etc. of the first voice segment adjacent to the second voice segment.

[0071] In addition, the method for determining the punctuation corresponding to the first speech segment described above is an optional and implementable method. In actual application scenarios, there can be multiple different determination methods, which can vary according to different actual application scenarios. The embodiments of the present invention do not make specific limitations in this regard.

[0072] In S106, when the duration of the second speech segment is not less than the first preset duration threshold, based on the target speech segment adjacent to the second speech segment in the first speech segment and the preset punctuation recognition rule, determine the punctuation corresponding to the second speech segment in the second punctuation.

[0073] Among them, the second punctuation can be used to indicate the end of a sentence. For example, the second punctuation can be a full stop, an exclamation mark, a question mark, or other punctuation marks used to indicate the end.

[0074] In implementation, when the duration of the second speech segment is not less than the first preset duration threshold, the target speech segment adjacent to the second speech segment can be obtained, and then based on the target speech segment and the preset punctuation recognition rule, determine the punctuation corresponding to the second speech segment.

[0075] For example, the punctuation corresponding to the second speech segment can be determined according to the duration of the target speech segment. For example, if the duration of the target speech segment is less than the preset duration (such as 600 ms), the punctuation corresponding to the second speech segment can be determined as an exclamation mark; if the duration of the target speech segment is not less than the preset duration, the punctuation corresponding to the second speech segment can be determined as a full stop or a question mark.

[0076] Alternatively, the human voice data included in the target speech segment can also be converted into text data. Based on the preset keyword extraction algorithm, extract the keywords in the text data, match the extracted keywords with the preset keywords, and determine the punctuation corresponding to the second speech segment according to the matching result. For example, assume that the text data converted from the target speech segment is "What's the weather like tomorrow", and the second punctuation that matches the keyword "like what" can be a question mark, that is, the punctuation corresponding to the second speech segment can be a question mark.

[0077] The method for determining the punctuation corresponding to the second speech segment in the second punctuation described above is an optional and implementable method. In actual application scenarios, there can be multiple different determination methods, which can vary according to different actual application scenarios. The embodiments of the present invention do not make specific limitations in this regard.

[0078] In S108, convert the first speech segment into target text data, and perform text conversion on the target speech data based on the target text data and the punctuation corresponding to the second speech segment to obtain text conversion data including punctuation.

[0079] In implementation, based on a preset speech recognition algorithm (such as an Automatic Speech Recognition (ASR) algorithm), the first speech segment can be converted into target text data.

[0080] Based on the target text data and the punctuation corresponding to the second speech segment, text conversion is performed on the target speech data to obtain text conversion data including punctuation. For example, as Figure 2 shown, based on the ASR algorithm, the first speech segment 1, the first speech segment 2, and the first speech segment 3 can be respectively converted into target text data, the punctuation corresponding to the second speech segment 1, the second speech segment 2, and the second speech segment 3 is respectively obtained, and based on the order of the speech segments, the target text data and the punctuation are combined to obtain text conversion data including punctuation corresponding to the target speech data.

[0081] An embodiment of the present invention provides a data processing method. By obtaining target speech data to be recognized, and determining a first speech segment and a second speech segment in the target speech data, the first speech segment is a segment of speech data including human voice data, and the second speech segment is a segment of speech data not including human voice data. When the duration of the second speech segment is less than a first preset duration threshold, it is determined that the punctuation corresponding to the second speech segment is a first punctuation, and the first punctuation is used to represent a sentence pause. When the duration of the second speech segment is not less than the first preset duration threshold, based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule, the punctuation corresponding to the second speech segment in the second punctuation is determined, and the second punctuation is used to represent the end of a sentence. The first speech segment is converted into target text data, and based on the target text data and the punctuation corresponding to the second speech segment, text conversion is performed on the target speech data to obtain text conversion data including punctuation. In this way, the punctuation corresponding to the second speech segment not including human voice data can be determined according to its duration, which can improve the efficiency of punctuation recognition. In addition, when the duration of the second speech segment is not less than the first preset time threshold, based on the preset punctuation recognition rule and the target speech segment adjacent to the second speech segment, the punctuation corresponding to the second speech segment can be determined, avoiding the problem of poor punctuation recognition accuracy caused by the large difference in acoustic features corresponding to different human voice data, and improving the accuracy of punctuation recognition in the process of text conversion for the target speech data.

[0082] Embodiment 2

[0083] As Figure 3As shown in the figure, an embodiment of the present invention provides a data processing method. The execution subject of this method can be a terminal device or a server. The server can be an independent server or a server cluster composed of multiple servers. The method may specifically include the following steps:

[0084] In S302, obtain the target voice data to be recognized, and determine the first voice segment and the second voice segment in the target voice data.

[0085] For the specific processing process of the above S302, reference can be made to the relevant content of S102 in the first embodiment above, which will not be elaborated here.

[0086] In S304, if there is a target second voice segment in the second voice segment whose duration is less than the second preset duration threshold, and there are two first voice segments in the first voice segment adjacent to the target second voice segment, then merge the two first voice segments and the target second voice segment into the first voice segment.

[0087] Among them, the second preset duration threshold can be less than the first preset duration threshold.

[0088] In implementation, for example, as Figure 2 shown, assume that the duration of the second voice segment 1 is 200 ms, and the second preset duration threshold is 250 ms, and the first voice segments adjacent to the second voice segment 1 include the first voice segment 1 and the first voice segment 2. Then, the first voice segment 1, the second voice segment 1, and the first voice segment 2 can be merged into the first voice segment, and the merged result can be as Figure 4 shown.

[0089] In S306, when the duration of the second voice segment is less than the first preset duration threshold, determine that the punctuation corresponding to the second voice segment is the first punctuation.

[0090] For the specific processing process of the above S306, reference can be made to the relevant content of S104 in the first embodiment above, which will not be elaborated here.

[0091] As Figure 3 shown, when the duration of the second voice segment is not less than the first preset duration threshold, after S308, S310 - S312 can be continued, or after S308, S314 - S320 can be continued.

[0092] In S308, obtain the historical first voice segments within the preset training period.

[0093] Among them, the preset training period can be any period, such as the recent one month, the recent half year, etc. The historical first voice segments are voice segments containing human voice data.

[0094] In S310, based on the historical first speech segment and the corresponding label, a supervised training is performed on a preset neural network model to obtain a pre-trained punctuation recognition model.

[0095] Among them, the label corresponding to the historical first speech segment is the punctuation in the second punctuation corresponding to the historical first speech segment.

[0096] In S312, when the duration of the second speech segment is not less than the first preset duration threshold, based on the target speech segment adjacent to the second speech segment in the first speech segment and the pre-trained punctuation recognition model, the punctuation in the second punctuation corresponding to the second speech segment is determined.

[0097] Among them, the punctuation recognition model can be obtained by training a model constructed by a preset machine learning algorithm based on the historical first speech segment.

[0098] In implementation, the target speech segment can be input into the pre-trained punctuation recognition model to determine the punctuation corresponding to the second speech segment.

[0099] In S314, the historical first speech segment is converted into historical text data based on a preset speech recognition algorithm.

[0100] Among them, the preset speech recognition algorithm can be an Automatic Speech Recognition (ASR) algorithm, and the historical first speech segment can be converted into historical text data based on the ASR algorithm.

[0101] In S316, based on the historical text data and the corresponding label, a supervised training is performed on a preset neural network model to obtain a pre-trained punctuation recognition model.

[0102] Among them, the label corresponding to the historical text data can be the punctuation in the second punctuation corresponding to the historical text data.

[0103] In implementation, training the model based on the historical text data can avoid the problem of inaccurate recognition caused by training the model with speech data, and there is no need to limit the ASR algorithm. It has good adaptability, strong portability, and the model is trained based on text features, so the robustness of the model is strong.

[0104] In S318, when the duration of the second speech segment is not less than the first preset duration threshold, based on the preset speech recognition algorithm, the target speech segment adjacent to the second speech segment in the first speech segment is converted into the first target text data.

[0105] In S320, based on the first target text data and the pre-trained punctuation recognition model, the punctuation in the second punctuation corresponding to the second speech segment is determined.

[0106] In S322, the first voice segment is converted into target text data, and text conversion is performed on the target voice data based on the target text data and the punctuation corresponding to the second voice segment, to obtain text conversion data including punctuation.

[0107] For the specific processing procedure of the above S322, reference may be made to the relevant content of S108 in the first embodiment above, which will not be elaborated herein.

[0108] An embodiment of the present invention provides a data processing method. By acquiring target voice data to be recognized, and determining a first voice segment and a second voice segment in the target voice data, where the first voice segment is a segment of voice data including human voice data, and the second voice segment is a segment of voice data not including human voice data. When the duration of the second voice segment is less than a first preset duration threshold, it is determined that the punctuation corresponding to the second voice segment is a first punctuation, and the first punctuation is used to represent a sentence pause. When the duration of the second voice segment is not less than the first preset duration threshold, based on the target voice segment adjacent to the second voice segment in the first voice segment and a preset punctuation recognition rule, it is determined that the punctuation corresponding to the second voice segment in the second punctuation, and the second punctuation is used to represent the end of a sentence. The first voice segment is converted into target text data, and text conversion is performed on the target voice data based on the target text data and the punctuation corresponding to the second voice segment, to obtain text conversion data including punctuation. In this way, the punctuation corresponding to the second voice segment not including human voice data can be determined according to its duration, which can improve the efficiency of punctuation recognition. In addition, when the duration of the second voice segment is not less than the first preset time threshold, based on the preset punctuation recognition rule and the target voice segment adjacent to the second voice segment, the punctuation corresponding to the second voice segment can be determined, avoiding the problem of poor punctuation recognition accuracy caused by the large difference in acoustic features corresponding to different human voice data, and improving the accuracy of punctuation recognition in the process of text conversion for the target voice data.

[0109] Embodiment III

[0110] The above is the data processing method provided by the embodiment of the present invention. Based on the same idea, the embodiment of the present invention further provides a data processing device, as Figure 5 shown.

[0111] The data processing device includes: a first acquisition module 501, a first determination module 502, a second determination module 503, and an identification module 504, where:

[0112] The first acquisition module 501 is configured to acquire target voice data to be recognized, and determine a first voice segment and a second voice segment in the target voice data, where the first voice segment is a segment of voice data containing human voice data, and the second voice segment is a segment of voice data not containing human voice data;

[0113] The first determination module 502 is configured to, when the duration of the second voice segment is less than a first preset duration threshold, determine that the punctuation corresponding to the second voice segment is a first punctuation, where the first punctuation is used to indicate a sentence pause;

[0114] The second determination module 503 is configured to, when the duration of the second voice segment is not less than the first preset duration threshold, determine the punctuation corresponding to the second voice segment in the second punctuation based on the target voice segment adjacent to the second voice segment in the first voice segment and a preset punctuation recognition rule, where the second punctuation is used to indicate the end of a sentence;

[0115] The recognition module 504 is configured to convert the first voice segment into target text data, and perform text conversion on the target voice data based on the target text data and the punctuation corresponding to the second voice segment to obtain text conversion data including punctuation.

[0116] In an embodiment of the present invention, the second determination module 503 is configured to:

[0117] Based on the target voice segment adjacent to the second voice segment in the first voice segment and a pre-trained punctuation recognition model, determine the punctuation corresponding to the second voice segment in the second punctuation, where the punctuation recognition model is obtained by training a model constructed by a preset machine learning algorithm based on historical first voice segment pairs.

[0118] In an embodiment of the present invention, the device further includes:

[0119] The second acquisition module is configured to acquire the historical first voice segments within a preset training period;

[0120] The first training module is configured to perform supervised training on a preset neural network model based on the historical first voice segments and corresponding labels to obtain the pre-trained punctuation recognition model, where the label corresponding to the historical first voice segment is the punctuation corresponding to the historical first voice segment in the second punctuation.

[0121] In an embodiment of the present invention, the device further includes:

[0122] The third acquisition module is configured to acquire the historical first voice segments within a preset training period;

[0123] A first conversion module, configured to convert the historical first speech segment into historical text data based on a preset speech recognition algorithm;

[0124] A second training module, configured to perform supervised training on a preset neural network model based on the historical text data and corresponding labels, to obtain the pre-trained punctuation recognition model, where the label corresponding to the historical text data is the punctuation in the second punctuation corresponding to the historical text data;

[0125] The second determination module 503 is configured to:

[0126] Convert a target speech segment adjacent to the second speech segment in the first speech segment into first target text data based on the preset speech recognition algorithm;

[0127] Determine the punctuation in the second punctuation corresponding to the second speech segment based on the first target text data and the pre-trained punctuation recognition model.

[0128] In an embodiment of the present invention, the apparatus further includes:

[0129] A merging module, configured to, if there is a target second speech segment in the second speech segment with a duration less than a second preset duration threshold, and there are two first speech segments in the first speech segment adjacent to the target second speech segment, merge the two first speech segments and the target second speech segment into the first speech segment, where the second preset duration threshold is less than the first preset duration threshold.

[0130] An embodiment of the present invention provides a data processing device. By obtaining target voice data to be recognized and determining a first voice segment and a second voice segment in the target voice data, where the first voice segment is a segment of voice data containing human voice data, and the second voice segment is a segment of voice data not containing human voice data. When the duration of the second voice segment is less than a first preset duration threshold, determine that the punctuation corresponding to the second voice segment is a first punctuation, and the first punctuation is used to represent a sentence pause. When the duration of the second voice segment is not less than the first preset duration threshold, based on the target voice segment adjacent to the second voice segment in the first voice segment and a preset punctuation recognition rule, determine the punctuation corresponding to the second voice segment in a second punctuation, and the second punctuation is used to represent the end of a sentence. Convert the first voice segment into target text data, and perform text conversion on the target voice data based on the target text data and the punctuation corresponding to the second voice segment to obtain text conversion data containing punctuation. In this way, the punctuation corresponding to the second voice segment not containing human voice data can be determined according to its duration, which can improve the efficiency of punctuation recognition. In addition, when the duration of the second voice segment is not less than the first preset time threshold, the punctuation corresponding to the second voice segment can be determined based on the preset punctuation recognition rule and the target voice segment adjacent to the second voice segment, avoiding the problem of poor punctuation recognition accuracy caused by the large difference in acoustic features corresponding to different human voice data, and improving the accuracy of punctuation recognition in the process of text conversion for target voice data.

[0131] Embodiment 4

[0132] Figure 6 A schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.

[0133] The electronic device 600 includes, but is not limited to: a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609, a processor 610, and a power supply 611, etc. Those skilled in the art can understand that Figure 6 the electronic device structure shown in does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, mobile phones, tablet computers, laptop computers, handheld computers, vehicle-mounted terminals, wearable devices, and pedometers, etc.

[0134] Among them, the processor 610 is used to: obtain target speech data to be recognized, and determine a first speech segment and a second speech segment in the target speech data, where the first speech segment is a segment of speech data containing human voice data, and the second speech segment is a segment of speech data not containing human voice data; when the duration of the second speech segment is less than a first preset duration threshold, determine that the punctuation corresponding to the second speech segment is a first punctuation, and the first punctuation is used to indicate a sentence pause; when the duration of the second speech segment is not less than the first preset duration threshold, based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule, determine the punctuation corresponding to the second speech segment in the second punctuation, and the second punctuation is used to indicate the end of a sentence; convert the first speech segment into target text data, and perform text conversion on the target speech data based on the target text data and the punctuation corresponding to the second speech segment to obtain text conversion data containing punctuation.

[0135] The processor 610 is further used to: based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model, determine the punctuation corresponding to the second speech segment in the second punctuation, where the punctuation recognition model is obtained by training a model constructed by a preset machine learning algorithm based on historical first speech segment pairs.

[0136] The processor 610 is further used to: obtain the historical first speech segments within a preset training period; perform supervised training on a preset neural network model based on the historical first speech segments and corresponding labels to obtain the pre-trained punctuation recognition model, where the label corresponding to the historical first speech segment is the punctuation corresponding to the historical first speech segment in the second punctuation.

[0137] The processor 610 is further used to: obtain the historical first speech segments within a preset training period; convert the historical first speech segments into historical text data based on a preset speech recognition algorithm; perform supervised training on a preset neural network model based on the historical text data and corresponding labels to obtain the pre-trained punctuation recognition model, where the label corresponding to the historical text data is the punctuation corresponding to the historical text data in the second punctuation; convert the target speech segment adjacent to the second speech segment in the first speech segment into first target text data based on the preset speech recognition algorithm; based on the first target text data and the pre-trained punctuation recognition model, determine the punctuation corresponding to the second speech segment in the second punctuation.

[0138] In addition, the processor 610 is further configured to: if there is a target second voice segment in the second voice segment whose duration is less than a second preset duration threshold, and there are two first voice segments adjacent to the target second voice segment in the first voice segment, merge the two first voice segments and the target second voice segment into the first voice segment, where the second preset duration threshold is less than the first preset duration threshold.

[0139] An embodiment of the present invention provides an electronic device. By acquiring target voice data to be recognized and determining a first voice segment and a second voice segment in the target voice data, where the first voice segment is a segment of voice data containing human voice data, and the second voice segment is a segment of voice data not containing human voice data. When the duration of the second voice segment is less than a first preset duration threshold, determine that the punctuation corresponding to the second voice segment is a first punctuation, where the first punctuation is used to indicate a sentence pause. When the duration of the second voice segment is not less than the first preset duration threshold, based on the target voice segment adjacent to the second voice segment in the first voice segment and a preset punctuation recognition rule, determine the punctuation corresponding to the second voice segment in the second punctuation, where the second punctuation is used to indicate the end of a sentence. Convert the first voice segment into target text data, and perform text conversion on the target voice data based on the target text data and the punctuation corresponding to the second voice segment to obtain text conversion data containing punctuation. In this way, the punctuation corresponding to the second voice segment not containing human voice data can be determined according to its duration, which can improve the efficiency of punctuation recognition. In addition, when the duration of the second voice segment is not less than the first preset time threshold, the punctuation corresponding to the second voice segment can be determined based on the preset punctuation recognition rule and the target voice segment adjacent to the second voice segment, avoiding the problem of poor punctuation recognition accuracy caused by the large difference in acoustic features corresponding to different human voice data, and improving the accuracy of punctuation recognition during the text conversion of the target voice data.

[0140] It should be understood that in an embodiment of the present invention, the radio frequency unit 601 can be used for receiving and sending signals during information reception and transmission or a call process. Specifically, after receiving downlink data from a base station, it is given to the processor 610 for processing; in addition, the uplink data is sent to the base station. Usually, the radio frequency unit 601 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc. In addition, the radio frequency unit 601 can also communicate with a network and other electronic devices through a wireless communication system.

[0141] The electronic device provides the user with wireless broadband Internet access through the network module 602, such as helping the user to send and receive emails, browse web pages, and access streaming media, etc.

[0142] The audio output unit 603 can convert the audio data received by the radio frequency unit 601 or the network module 602 or stored in the memory 609 into an audio signal and output it as sound. Moreover, the audio output unit 603 can also provide an audio output related to a specific function executed by the electronic device 600 (e.g., call signal reception sound, message reception sound, etc.). The audio output unit 603 includes a speaker, a buzzer, a receiver, etc.

[0143] The input unit 604 is used to receive audio or video signals. The input unit 604 may include a Graphics Processing Unit (GPU) 6041 and a microphone 6042. The graphics processor 6041 processes the image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The processed image frame can be displayed on the display unit 606. The image frame processed by the graphics processor 6041 can be stored in the memory 609 (or other storage media) or transmitted via the radio frequency unit 601 or the network module 602. The microphone 6042 can receive sound and can process such sound into audio data. The processed audio data can be output in a format that can be transmitted to a mobile communication base station via the radio frequency unit 601 in the case of a phone call mode.

[0144] The electronic device 600 further includes at least one sensor 605, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 6061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 6061 and / or the backlight when the electronic device 600 is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used to identify the posture of the electronic device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as a pedometer, tapping), etc.; the sensor 605 can also include a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be elaborated here.

[0145] The display unit 606 is used to display the information input by the user or the information provided to the user. The display unit 606 may include a display panel 6061, and the display panel 6061 can be configured in the form of a Liquid Crystal Display (LCD), an Organic Light-Emitting Diode (OLED), etc.

[0146] The user input unit 607 can be used to receive input numerical or character information and generate key signal inputs related to user settings and function control of the electronic device. Specifically, the user input unit 607 includes a touch panel 6071 and other input devices 6072. The touch panel 6071, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 6071). The touch panel 6071 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 610, and receives and executes the command sent by the processor 610. In addition, the touch panel 6071 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 6071, the user input unit 607 can also include other input devices 6072. Specifically, the other input devices 6072 can include but are not limited to a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, which will not be elaborated here.

[0147] Furthermore, the touch panel 6071 can cover the display panel 6061. After the touch panel 6071 detects a touch operation thereon or nearby, it transmits the operation to the processor 610 to determine the type of the touch event. Subsequently, the processor 610 provides a corresponding visual output on the display panel 6061 according to the type of the touch event. Although in Figure 6 the touch panel 6071 and the display panel 6061 are implemented as two independent components to realize the input and output functions of the electronic device, in some embodiments, the touch panel 6071 and the display panel 6061 can be integrated to realize the input and output functions of the electronic device, and the specific implementation here is not limited.

[0148] The interface unit 608 is an interface for connecting an external device to the electronic device 600. For example, the external device can include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headset port, etc. The interface unit 608 can be used to receive inputs from the external device (such as data information, power, etc.) and transmit the received inputs to one or more components within the electronic device 600 or can be used to transmit data between the electronic device 600 and the external device.

[0149] The memory 609 can be used to store software programs and various data. The memory 609 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 609 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0150] The processor 610 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 609, and calling data stored in the memory 609, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. The processor 610 can include one or more processing units; preferably, the processor 610 can integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 610.

[0151] The electronic device 600 can also include a power supply 611 (such as a battery) for powering each component. Preferably, the power supply 611 can be logically connected to the processor 610 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.

[0152] Preferably, an embodiment of the present invention further provides an electronic device, including a processor 610, a memory 609, and a computer program stored on the memory 609 and executable on the processor 610. When the computer program is executed by the processor 610, it implements each process of the above-mentioned data processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0153] Embodiment Five

[0154] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned data processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (Read-Only Memory, abbreviated as ROM), a random access memory (Random Access Memory, abbreviated as RAM), a magnetic disk, or an optical disc, etc.

[0155] An embodiment of the present invention provides a computer-readable storage medium. By obtaining target voice data to be recognized and determining a first voice segment and a second voice segment in the target voice data, where the first voice segment is a segment of voice data containing human voice data, and the second voice segment is a segment of voice data not containing human voice data. When the duration of the second voice segment is less than a first preset duration threshold, determine that the punctuation corresponding to the second voice segment is a first punctuation, and the first punctuation is used to indicate a sentence pause. When the duration of the second voice segment is not less than the first preset duration threshold, based on the target voice segment adjacent to the second voice segment in the first voice segment and a preset punctuation recognition rule, determine the punctuation corresponding to the second voice segment in a second punctuation, and the second punctuation is used to indicate the end of a sentence. Convert the first voice segment into target text data, and perform text conversion on the target voice data based on the target text data and the punctuation corresponding to the second voice segment to obtain text conversion data containing punctuation. In this way, the punctuation corresponding to the second voice segment not containing human voice data can be determined according to its duration, which can improve the efficiency of punctuation recognition. In addition, when the duration of the second voice segment is not less than the first preset time threshold, based on the preset punctuation recognition rule and the target voice segment adjacent to the second voice segment, the punctuation corresponding to the second voice segment can be determined, avoiding the problem of poor punctuation recognition accuracy caused by the large difference in acoustic features corresponding to different human voice data, and improving the punctuation recognition accuracy in the process of text conversion for the target voice data.

[0156] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0157] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0158] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the block or blocks.

[0160] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0161] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.

[0162] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0163] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0164] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0165] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A data processing method, characterized in that, The method includes: Obtaining target speech data to be recognized, and determining a first speech segment and a second speech segment in the target speech data, where the first speech segment is a segment of speech data containing human voice data, and the second speech segment is a segment of speech data not containing human voice data; When the duration of the second speech segment is less than a first preset duration threshold, determining that the punctuation corresponding to the second speech segment is a first punctuation, and the first punctuation is used to indicate a sentence pause; When the duration of the second speech segment is not less than the first preset duration threshold, based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule, determining the punctuation corresponding to the second speech segment in a second punctuation, and the second punctuation is used to indicate the end of a sentence; Converting the first speech segment into target text data, and performing text conversion on the target speech data based on the target text data and the punctuation corresponding to the second speech segment to obtain text conversion data including punctuation.

2. The method according to claim 1, characterized in that, The determining the punctuation corresponding to the second speech segment in the second punctuation based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule includes: Based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model, determining the punctuation corresponding to the second speech segment in the second punctuation, where the punctuation recognition model is obtained by training a model constructed by a preset machine learning algorithm based on historical first speech segment pairs.

3. The method according to claim 2, wherein Before the determining the punctuation corresponding to the second speech segment in the second punctuation based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model, it further includes: Obtaining the historical first speech segments within a preset training period; Performing supervised training on a preset neural network model based on the historical first speech segments and corresponding labels to obtain the pre-trained punctuation recognition model, where the label corresponding to the historical first speech segment is the punctuation corresponding to the historical first speech segment in the second punctuation.

4. The method according to claim 2, wherein Before the determining the punctuation corresponding to the second speech segment in the second punctuation based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model, it further includes: Obtaining the historical first speech segments within a preset training period; Converting the historical first speech segments into historical text data based on a preset speech recognition algorithm; Performing supervised training on a preset neural network model based on the historical text data and corresponding labels to obtain the pre-trained punctuation recognition model, where the label corresponding to the historical text data is the punctuation corresponding to the historical text data in the second punctuation; The determining the punctuation corresponding to the second speech segment in the second punctuation based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model includes: Based on the preset speech recognition algorithm, convert the target speech segment adjacent to the second speech segment in the first speech segment into first target text data; Based on the first target text data and the pre-trained punctuation recognition model, determine the punctuation corresponding to the second speech segment in the second punctuation.

5. The method according to claim 1, wherein Before converting the first speech segment into target text data, it further includes: If there is a target second speech segment with a duration less than the second preset duration threshold in the second speech segment, and there are two first speech segments adjacent to the target second speech segment in the first speech segment, then merge the two first speech segments and the target second speech segment into the first speech segment, where the second preset duration threshold is less than the first preset duration threshold.

6. A data processing device, characterized in that, The device includes: A first acquisition module, configured to acquire target speech data to be recognized, and determine a first speech segment and a second speech segment in the target speech data, where the first speech segment is a segment of speech data containing human voice data, and the second speech segment is a segment of speech data not containing human voice data; A first determination module, configured to determine that the punctuation corresponding to the second speech segment is a first punctuation when the duration of the second speech segment is less than the first preset duration threshold, and the first punctuation is used to indicate a sentence pause; A second determination module, configured to, when the duration of the second speech segment is not less than the first preset duration threshold, based on the target speech segment adjacent to the second speech segment in the first speech segment and a preset punctuation recognition rule, determine the punctuation corresponding to the second speech segment in the second punctuation, and the second punctuation is used to indicate the end of a sentence; A recognition module, configured to convert the first speech segment into target text data, and perform text conversion on the target speech data based on the target text data and the punctuation corresponding to the second speech segment to obtain text conversion data including punctuation.

7. The device according to claim 6, characterized in that, The second determination module is configured to: Based on the target speech segment adjacent to the second speech segment in the first speech segment and a pre-trained punctuation recognition model, determine the punctuation corresponding to the second speech segment in the second punctuation, and the punctuation recognition model is obtained by training a model constructed by a preset machine learning algorithm based on historical first speech segment pairs.

8. The device according to claim 7, characterized in that, The device further includes: A second acquisition module, configured to acquire the historical first speech segments within a preset training period; A first training module, configured to perform supervised training on a preset neural network model based on the historical first speech segments and corresponding labels to obtain the pre-trained punctuation recognition model, and the label corresponding to the historical first speech segment is the punctuation corresponding to the historical first speech segment in the second punctuation.

9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the data processing method according to any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Call emotion real-time identification method and device, computer equipment and storage medium

    CN113241095A

  • Audio signal processing method, model training method and device, equipment and medium

    CN113380238A