Intelligent voice transfer method, system and device and storage medium

By separating voice data and extracting feature, combining linear prediction analysis and Bert model correction processing, target text sequences including punctuation marks are generated, which solves the problem of low accuracy of the existing speech transcription system and achieves higher speech transcription accuracy.

CN120279919APending Publication Date: 2025-07-08NANJING PUTIAN TELEGE INTELLIGENT BUILDING
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510575255.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing speech transfer system has the problem of low accuracy in speech transfer, especially when different users change their voices, such as multiple characters, few characters, homophones, similar tones, close pronunciation and inaccurate semantic understanding.

Method used

By collecting voice data, separating voices, identifying voice fragments from different users, determining acoustic feature information using linear predictive analysis, and generating a target text sequence including punctuation mark information through feature extraction and correction processing, and typo correction is performed in combination with the Bert model to improve the accuracy of speech translation.

Benefits of technology

It effectively improves the accuracy of speech transfer and solves the problem of inaccurate speech transfer in the prior art, especially in multi-user scenarios, which significantly improves the accuracy of text recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279919A_ABST
    Figure CN120279919A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent voice transcription method, system and device and a storage medium, and the method comprises the steps: collecting voice data, carrying out the human voice separation of the voice data, and obtaining voice segments corresponding to different users; determining a to-be-processed text sequence corresponding to the voice segment, and determining acoustic feature information of the voice segment through linear prediction analysis; feature extraction is carried out on the to-be-processed text sequence, correction processing is carried out on the to-be-processed text sequence based on a feature extraction result, and a processed text sequence is obtained; and according to the processed text sequence and the acoustic feature information, determining a target text sequence including punctuation mark information and corresponding to the voice segment. Compared with the prior art, the speech data is separated, the speech segments are converted into the characters, then the to-be-processed text sequence is corrected, and finally the target text sequence which corresponds to the speech segments and comprises the punctuation mark information is determined, so that the speech transcription accuracy is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an intelligent speech transcription method, system, device, and storage medium. Background Art

[0002] Speech is the most natural way of communication, and many occasions involve communication through speech, such as: phone calls, speeches, medical consultations, court trials, meetings, and so on. The content of these communications needs to be recorded, and usually, the communication content is recorded by recording audio or manual typing. However, audio recording is not convenient for retrieval and query, and it is impossible to quickly locate the content that needs to be located; while using a court stenographer to quickly record the content of a voice communication is restricted by the typing speed. Sometimes, when the speech speed is too fast, some important content may be missed, and it is also affected by the state of the court stenographer.

[0003] A speech transcription system is a speech processing system that can transcribe speech into text. Through this system, meeting minutes can be automatically formed to improve meeting efficiency, give full play to the functions of meetings, avoid waste of human, material, and financial resources, reduce meeting costs, and achieve the efficiency of human resources. However, in the existing solutions, there are problems such as multiple characters, missing characters, homophones with different tones, near-homophones with the same tone, and inaccurate semantic understanding in speech transcription by different users, resulting in low accuracy of speech transcription.

[0004] Therefore, there is an urgent need for an intelligent speech transcription method that can effectively improve the accuracy of speech transcription. Summary of the Invention

[0005] The main objective of the present invention is to provide an intelligent speech transcription method, system, device, and storage medium, aiming to solve the technical problem of low accuracy of speech transcription in the existing technology.

[0006] To achieve the above objective, the present invention provides an intelligent speech transcription method, and the method includes the following steps:

[0007] Collect speech data, and perform voice separation on the speech data to obtain speech segments corresponding to different users;

[0008] Determine the text sequence to be processed corresponding to the speech segment, and determine the acoustic feature information of the speech segment through linear prediction analysis;

[0009] Extract features from the text sequence to be processed, and perform correction processing on the text sequence to be processed based on the feature extraction result to obtain a processed text sequence;

[0010] Determine a target text sequence including punctuation symbol information corresponding to the speech segment according to the processed text sequence and the acoustic feature information.

[0011] Optionally, the step of separating human voices from the voice data to obtain voice segments corresponding to different users includes:

[0012] Performing voiceprint recognition on the voice data through a preset voiceprint recognition model to obtain voiceprint features;

[0013] Performing clustering processing on the voiceprint features to obtain a clustering result;

[0014] If the clustering result indicates that the voice data contains at least two voice segments corresponding to different users, then perform voice separation on the voice data according to the clustering result to obtain voice segments corresponding to different users.

[0015] Optionally, after the step of performing clustering processing on the voiceprint features to obtain a clustering result, the method further includes:

[0016] If the clustering result indicates that the voice data contains only one voice segment corresponding to a user, then execute the step of determining the text sequence to be processed corresponding to the voice segment and determining the acoustic feature information of the voice segment through linear prediction analysis.

[0017] Optionally, the step of determining the text sequence to be processed corresponding to the voice segment and determining the acoustic feature information of the voice segment through linear prediction analysis includes:

[0018] Performing intent recognition on the voice segment to determine all intents corresponding to each voice segment;

[0019] Obtaining a preset set of scenario intents, and recording the voice segments whose all intents are not in the preset set of scenario intents as segments to be excluded;

[0020] Performing speech recognition on the voice segments other than the segments to be excluded to obtain the corresponding text sequence to be processed;

[0021] Determining the acoustic feature information of the voice segments other than the segments to be excluded through linear prediction analysis.

[0022] Optionally, the step of performing feature extraction on the text sequence to be processed, performing correction processing on the text sequence to be processed based on the feature extraction result, and obtaining a processed text sequence includes:

[0023] Performing encoding conversion on the text sequence to be processed to obtain encoded integer symbol information;

[0024] Performing feature extraction on the encoded integer symbol information through a preset feature extraction model to obtain semantic feature information and part-of-speech feature information, and using the semantic feature information and part-of-speech feature information as the feature extraction result;

[0025] Perform correction processing on the to-be-processed text sequence based on the feature extraction result to obtain a processed text sequence.

[0026] Optionally, the step of performing correction processing on the to-be-processed text sequence based on the feature extraction result to obtain a processed text sequence includes:

[0027] Perform merging processing on the semantic feature information and the part-of-speech feature information to obtain encoder feature information;

[0028] Perform decoding processing based on the encoder feature information to obtain corrected text information corresponding to the to-be-processed text sequence;

[0029] Compare the corrected text information with the to-be-processed text sequence, and determine error position information in the to-be-processed text sequence according to the comparison result;

[0030] Perform correction processing on the to-be-processed text sequence based on the error position information and the corrected text information to obtain a processed text sequence.

[0031] Optionally, the step of determining a target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information includes:

[0032] Determine first punctuation information related to the text semantic information of the speech segment according to the processed text sequence;

[0033] Determine second punctuation information related to the text semantic information and the acoustic feature information of the speech segment according to the first punctuation information and the acoustic feature information;

[0034] Determine a target text sequence including punctuation information corresponding to the speech segment according to the second punctuation information and the processed text sequence.

[0035] In addition, to achieve the above object, the present invention also proposes an intelligent speech transcription system, and the system includes:

[0036] A data acquisition module, configured to acquire speech data and perform voice separation on the speech data to obtain speech segments corresponding to different users;

[0037] A speech recognition module, configured to determine a to-be-processed text sequence corresponding to the speech segment and determine acoustic feature information of the speech segment through linear prediction analysis;

[0038] A text correction module, configured to extract features from the to-be-processed text sequence, and perform correction processing on the to-be-processed text sequence based on the feature extraction result to obtain a processed text sequence;

[0039] A result output module, configured to determine a target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information.

[0040] In addition, to achieve the above object, the present invention also provides an intelligent speech transcription device, including: a memory, a processor, and an intelligent speech transcription program stored on the memory and executable on the processor, where the intelligent speech transcription program is configured to implement the steps of the intelligent speech transcription method as described above.

[0041] In addition, to achieve the above object, the present invention also provides a storage medium, on which an intelligent speech transcription program is stored, and when the intelligent speech transcription program is executed by a processor, it implements the steps of the intelligent speech transcription method as described above.

[0042] The present invention discloses collecting speech data, performing voice separation on the speech data to obtain speech segments corresponding to different users; determining a to-be-processed text sequence corresponding to the speech segment, and determining acoustic feature information of the speech segment through linear prediction analysis; performing feature extraction on the to-be-processed text sequence, and performing correction processing on the to-be-processed text sequence based on the feature extraction result to obtain a processed text sequence; determining a target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information. Since the present invention separates speech data, converts speech segments into text, then performs correction processing on the to-be-processed text sequence according to the feature extraction result of the to-be-processed text sequence, and finally determines a target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information, compared with the prior art, the present invention effectively improves the accuracy of speech transcription. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic flowchart of the first embodiment of the intelligent speech transcription method of the present invention;

[0044] Figure 2 It is a schematic flowchart of the second embodiment of the intelligent speech transcription method of the present invention;

[0045] Figure 3 It is a schematic flowchart of the third embodiment of the intelligent speech transcription method of the present invention;

[0046] Figure 4 It is a structural block diagram of the first embodiment of the intelligent speech transcription device of the present invention;

[0047] Figure 5 It is a schematic structural diagram of an intelligent speech transcription device in the hardware operating environment involved in the embodiment solution of the present invention.

[0048] The realization, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments

[0049] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0050] The embodiment of the present invention provides an intelligent speech transcription method. Refer to Figure 1 , Figure 1 It is a schematic flowchart of the first embodiment of the intelligent speech transcription method of the present invention.

[0051] In this embodiment, the intelligent speech transcription method includes steps S10 to S40:

[0052] Step S10: Collect speech data and perform voice separation on the speech data to obtain speech segments corresponding to different users.

[0053] It should be noted that the execution subject of this embodiment can be a computer server device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a smart phone, a smart watch, etc., or an electronic device, intelligent speech transcription, etc. that can implement the above functions. Hereinafter, an intelligent speech transcription system including an intelligent speech transcription device (hereinafter referred to as the system) will be taken as an example to illustrate this embodiment and the following embodiments.

[0054] It should be understood that the above voice separation can be to separate the speech segments corresponding to different users in the speech data.

[0055] In a specific implementation, it can be to perform voiceprint recognition on the speech data through a preset voiceprint recognition model to obtain voiceprint features; perform clustering processing on the voiceprint features to obtain a clustering result; if the clustering result indicates that the speech data contains at least two speech segments corresponding to different users, then perform voice separation on the speech data according to the clustering result to obtain speech segments corresponding to different users.

[0056] It should be noted that if the clustering result indicates that the speech data contains only one speech segment corresponding to a user, then perform the step of determining the text sequence to be processed corresponding to the speech segment and determining the acoustic feature information of the speech segment through linear prediction analysis.

[0057] It should be noted that the above-mentioned preset voiceprint recognition model is a mathematical model established through specific algorithms and technical means based on the speaker information contained in the speech waveform for automatically identifying the speaker's identity.

[0058] Step S20: Determine the text sequence to be processed corresponding to the speech segment, and determine the acoustic feature information of the speech segment through linear prediction analysis.

[0059] It should be noted that the text sequence to be processed corresponding to the speech segment can be determined through an acoustic model and a language model. Among them, the acoustic model can convert the input speech signal into the posterior probability score of an acoustic modeling unit (also known as a phoneme, pronunciation unit); the language model can be used to predict the prior probability of the occurrence of a word sequence. For a given word sequence, then, through the decoder, by combining the acoustic model score and the language model score, a decoding network can be constructed, and the text sequence to be processed can be obtained through optimal path search.

[0060] It should be understood that the above-mentioned text sequence to be processed can be a text sequence that does not include punctuation information.

[0061] It can be understood that the acoustic feature information includes at least one of the following acoustic feature information: Bottleneck feature, fbank feature, word duration, post-word silence duration, pitch feature, etc.

[0062] In a specific implementation, the text sequence to be processed corresponding to the speech segment can be determined through an acoustic model and a language model, and then the acoustic feature information of the speech segment can be determined through linear prediction analysis.

[0063] It should be noted that linear prediction analysis is a method for predicting current or future sample values based on past known sample values of a random signal.

[0064] Step S30: Extract features from the text sequence to be processed, and perform correction processing on the text sequence to be processed based on the feature extraction result to obtain a processed text sequence.

[0065] It should be understood that in the existing solutions, there are problems such as homophones with different tones and near-homophones with the same tone caused by different user voice conversions in speech transcription, resulting in low accuracy of speech transcription. Therefore, the features of the sentences with typos in the text sequence to be processed can be extracted through the Bert model, and then the key context information can be extracted using the attention mechanism of the Bert model, combined with the decoder, to achieve sequence-to-sequence correction of typos, thereby improving the accuracy of speech transcription.

[0066] Step S40: Determine the target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information.

[0067] It should be understood that in the existing solutions, punctuation prediction is usually only considered based on text semantics, without considering the input of the ASR system, that is, acoustic feature information, resulting in inaccurate recognition of punctuation in speech texts, which in turn affects text semantics.

[0068] It should be explained that the above punctuation information includes punctuation information related to the text semantic information and acoustic feature information of the speech segment.

[0069] It should be noted that according to the processed text sequence, the first punctuation information related to the text semantic information of the speech segment is determined; according to the first punctuation information and the acoustic feature information, the second punctuation information related to the text semantic information and the acoustic feature information of the speech segment is determined; according to the second punctuation information and the processed text sequence, the target text sequence including punctuation information corresponding to the speech segment is determined.

[0070] This embodiment discloses collecting speech data, separating the human voices from the speech data to obtain speech segments corresponding to different users; determining the text sequence to be processed corresponding to the speech segment, and determining the acoustic feature information of the speech segment through linear prediction analysis; extracting features from the text sequence to be processed, and performing correction processing on the text sequence to be processed based on the feature extraction result to obtain a processed text sequence; determining the target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information. Since this embodiment separates speech data, converts speech segments into text, then corrects the text sequence to be processed according to the feature extraction result of the text sequence to be processed, and finally determines the target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information, compared with the prior art, this embodiment effectively improves the accuracy of speech transcription.

[0071] Reference Figure 2 , Figure 2 is a schematic flowchart of the second embodiment of the intelligent speech transcription method of the present invention.

[0072] Based on the above first embodiment, in this embodiment, the step S20 includes steps S201 to S204:

[0073] Step S201: Perform intent recognition on the speech segment to determine all intents corresponding to each speech segment.

[0074] Step S202: Obtain a preset scene intent set, and record the speech segments whose all intents are not in the preset scene intent set as segments to be excluded.

[0075] Step S203: Perform speech recognition on the speech segments other than the to-be-excluded segments to obtain a corresponding sequence of text to be processed.

[0076] Step S204: Determine the acoustic feature information of the speech segments other than the to-be-excluded segments through linear prediction analysis.

[0077] It should be understood that the above preset scenario intention set can be a scenario intention set obtained by integrating all possible speech intentions in the current application scenario.

[0078] This embodiment discloses performing intention recognition on the speech segments to determine all the intentions corresponding to each speech segment; obtaining a preset scenario intention set, and recording the speech segments whose all intentions are not in the preset scenario intention set as to-be-excluded segments; performing speech recognition on the speech segments other than the to-be-excluded segments to obtain a corresponding sequence of text to be processed; and determining the acoustic feature information of the speech segments other than the to-be-excluded segments through linear prediction analysis. Compared with the prior art, this embodiment can speed up the recognition speed and reduce the amount of recognition data by excluding the speech segments whose all intentions are not in the preset scenario intention set.

[0079] Reference Figure 3 , Figure 3 is a schematic flowchart of the third embodiment of the intelligent speech transcribing method of the present invention.

[0080] Based on the above embodiments, in this embodiment, the step S30 includes steps S301 to S303:

[0081] Step S301: Perform encoding conversion on the sequence of text to be processed to obtain encoded integer symbol information.

[0082] In a specific implementation, the text character information corresponding to the sequence of text to be processed can be determined, and then the encoded integer symbol information corresponding to the text character information can be determined based on a preset mapping table.

[0083] It should be noted that the preset mapping table can be a table used to represent the mapping relationship between text character information and encoded integer symbol information. Therefore, the effect of converting text character information into encoded integer symbol information can be achieved through the preset mapping table.

[0084] It should be understood that the above preset mapping table can be custom-set according to specific circumstances, and this embodiment does not limit this.

[0085] Step S302: Perform feature extraction on the encoded integer symbol information through a preset feature extraction model to obtain semantic feature information and part-of-speech feature information, and use the semantic feature information and part-of-speech feature information as the feature extraction results.

[0086] Step S303: Based on the feature extraction result, correct the to-be-processed text sequence to obtain the processed text sequence.

[0087] In a specific implementation, the semantic feature information and the part-of-speech feature information can be merged to obtain encoder feature information; based on the encoder feature information, decoding processing is performed to obtain the corrected text information corresponding to the to-be-processed text sequence; the corrected text information is compared with the to-be-processed text sequence, and according to the comparison result, the error position information in the to-be-processed text sequence is determined; based on the error position information and the corrected text information, the to-be-processed text sequence is corrected to obtain the processed text sequence.

[0088] In this embodiment, the to-be-processed text sequence is encoded and converted to obtain encoded integer symbol information; the encoded integer symbol information is subjected to feature extraction through a preset feature extraction model to obtain semantic feature information and part-of-speech feature information, and the semantic feature information and the part-of-speech feature information are used as the feature extraction result; based on the feature extraction result, the to-be-processed text sequence is corrected to obtain the processed text sequence. Compared with the prior art, this embodiment can maximize the utilization of the content contained in the to-be-processed text sequence, make the result of ASR transcription closest to the true and correct result, realize accurate error correction of the speech recognition text, and improve the accuracy rate of speech transcription.

[0089] In addition, an embodiment of the present invention also proposes a storage medium, on which an intelligent speech transcription program is stored, and when the intelligent speech transcription program is executed by a processor, the steps of the intelligent speech transcription method described above are implemented.

[0090] Refer to Figure 4 , Figure 4 which is the structural block diagram of the first embodiment of the intelligent speech transcription system of the present invention.

[0091] As Figure 4 shown, the intelligent speech transcription system proposed by the embodiment of the present invention includes: a data acquisition module 401, a speech recognition module 402, a text correction module 403, and a result output module 404.

[0092] The data acquisition module 401 is used to collect speech data and perform voice separation on the speech data to obtain speech segments corresponding to different users.

[0093] The speech recognition module 402 is used to determine the to-be-processed text sequence corresponding to the speech segment and determine the acoustic feature information of the speech segment through linear prediction analysis.

[0094] The text correction module 403 is configured to extract features from the to-be-processed text sequence, and correct the to-be-processed text sequence based on the feature extraction result to obtain a processed text sequence.

[0095] The result output module 404 is configured to determine a target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information.

[0096] The data acquisition module 401 is further configured to perform voiceprint recognition on the voice data through a preset voiceprint recognition model to obtain voiceprint features; perform clustering processing on the voiceprint features to obtain a clustering result; if the clustering result indicates that the voice data contains at least two voice segments corresponding to different users, separate the voices in the voice data according to the clustering result to obtain voice segments corresponding to different users.

[0097] The data acquisition module 401 is further configured to, if the clustering result indicates that the voice data contains only one voice segment corresponding to a user, perform the steps of determining the to-be-processed text sequence corresponding to the voice segment and determining the acoustic feature information of the voice segment through linear prediction analysis.

[0098] The result output module 404 is further configured to determine first punctuation information related to the text semantic information of the voice segment according to the processed text sequence; determine second punctuation information related to the text semantic information and the acoustic feature information of the voice segment according to the first punctuation information and the acoustic feature information; determine a target text sequence including punctuation information corresponding to the voice segment according to the second punctuation information and the processed text sequence.

[0099] This system embodiment discloses collecting voice data, separating the voices in the voice data to obtain voice segments corresponding to different users; determining the to-be-processed text sequence corresponding to the voice segment and determining the acoustic feature information of the voice segment through linear prediction analysis; extracting features from the to-be-processed text sequence, and correcting the to-be-processed text sequence based on the feature extraction result to obtain a processed text sequence; determining a target text sequence including punctuation information corresponding to the voice segment according to the processed text sequence and the acoustic feature information. Since this system embodiment separates the voice data, converts the voice segment into text, then corrects the to-be-processed text sequence according to the feature extraction result of the to-be-processed text sequence, and finally determines a target text sequence including punctuation information corresponding to the voice segment according to the processed text sequence and the acoustic feature information, compared with the prior art, this system embodiment effectively improves the accuracy of speech transcription.

[0100] Based on the first embodiment of the intelligent speech transcription system of the present invention, a second embodiment of the intelligent speech transcription system of the present invention is proposed.

[0101] In this embodiment, the speech recognition module 402 is further configured to perform intent recognition on the speech segments, determine all intents corresponding to each speech segment; obtain a preset scenario intent set, and record the speech segments whose all intents are not in the preset scenario intent set as segments to be excluded; perform speech recognition on the speech segments other than the segments to be excluded to obtain a corresponding text sequence to be processed; and determine the acoustic feature information of the speech segments other than the segments to be excluded through linear prediction analysis.

[0102] For other embodiments or specific implementation manners of the intelligent speech transcription system of the present invention, reference may be made to the above method embodiments, which will not be elaborated herein.

[0103] This application provides an intelligent speech transcription device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the intelligent speech transcription method in the first embodiment above.

[0104] Next, with reference to Figure 5 , which shows a schematic structural diagram of an intelligent speech transcription device suitable for implementing the embodiments of the present application. The intelligent speech transcription device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The shown intelligent speech transcription device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0105] As Figure 5As shown in the figure, the intelligent voice transcription device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the intelligent voice transcription device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the intelligent voice transcription device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an intelligent voice transcription device having various systems, it should be understood that it is not required to implement or include all the shown systems. More or fewer systems may be implemented or included alternatively.

[0106] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart may be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the method of the embodiments disclosed in the present application are executed.

[0107] The intelligent voice transcription device provided by the present application adopts the intelligent voice transcription method in the above embodiments, and can solve the technical problem of low accuracy of voice transcription in the prior art. Compared with the prior art, the beneficial effects of the intelligent voice transcription device provided by the present application are the same as those of the intelligent voice transcription method provided by the above embodiments, and other technical features in the intelligent voice transcription device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0108] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0109] As mentioned above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all such changes or substitutions should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0110] It should be noted that in this text, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or system including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or system including such element.

[0111] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0113] The above are only the preferred embodiments of the present invention, and thus do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. An intelligent speech transcription method, characterized in that, The method includes: Collecting voice data, separating voices from the voice data, and obtaining voice segments corresponding to different users; Determining a text sequence to be processed corresponding to the voice segment, and determining acoustic feature information of the voice segment through linear prediction analysis; Performing feature extraction on the text sequence to be processed, and performing correction processing on the text sequence to be processed based on the feature extraction result to obtain a processed text sequence; Determining a target text sequence including punctuation information corresponding to the voice segment according to the processed text sequence and the acoustic feature information.

2. The intelligent speech transcription method according to claim 1, wherein The step of separating voices from the voice data and obtaining voice segments corresponding to different users includes: Performing voiceprint recognition on the voice data through a preset voiceprint recognition model to obtain voiceprint features; Performing clustering processing on the voiceprint features to obtain a clustering result; If the clustering result indicates that the voice data contains at least two voice segments corresponding to different users, separating voices from the voice data according to the clustering result to obtain voice segments corresponding to different users.

3. The intelligent speech transcription method according to claim 2, characterized in that, After the step of performing clustering processing on the voiceprint features and obtaining a clustering result, it further includes: If the clustering result indicates that the voice data contains only one voice segment corresponding to a user, performing the step of determining the text sequence to be processed corresponding to the voice segment and determining the acoustic feature information of the voice segment through linear prediction analysis.

4. The intelligent speech transcription method according to claim 1, wherein The step of determining the text sequence to be processed corresponding to the voice segment and determining the acoustic feature information of the voice segment through linear prediction analysis includes: Performing intent recognition on the voice segment to determine all intents corresponding to each voice segment; Obtaining a preset set of scenario intents, and recording voice segments whose all intents are not in the preset set of scenario intents as segments to be excluded; Performing speech recognition on voice segments other than the segments to be excluded to obtain a corresponding text sequence to be processed; Determining the acoustic feature information of the voice segments other than the segments to be excluded through linear prediction analysis.

5. The intelligent speech transcription method according to claim 1, wherein The step of performing feature extraction on the text sequence to be processed, and performing correction processing on the text sequence to be processed based on the feature extraction result to obtain a processed text sequence includes: Performing encoding conversion on the text sequence to be processed to obtain encoded integer symbol information; Performing feature extraction on the encoded integer symbol information through a preset feature extraction model to obtain semantic feature information and part-of-speech feature information, and using the semantic feature information and the part-of-speech feature information as the feature extraction result; Performing correction processing on the text sequence to be processed based on the feature extraction result to obtain a processed text sequence.

6. The intelligent speech transcription method according to claim 5, wherein, The step of performing correction processing on the text sequence to be processed based on the feature extraction result to obtain a processed text sequence includes: Performing merging processing on the semantic feature information and the part-of-speech feature information to obtain encoder feature information; Performing decoding processing based on the encoder feature information to obtain correction text information corresponding to the text sequence to be processed. Compare the corrected text information with the text sequence to be processed, and determine the error location information in the text sequence to be processed according to the comparison result; Based on the error location information and the corrected text information, perform correction processing on the text sequence to be processed to obtain a processed text sequence.

7. The intelligent speech transcription method according to claim 1, wherein The step of determining the target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information includes: According to the processed text sequence, determine the first punctuation information related to the text semantic information of the speech segment; According to the first punctuation information and the acoustic feature information, determine the second punctuation information related to the text semantic information and the acoustic feature information of the speech segment; According to the second punctuation information and the processed text sequence, determine the target text sequence including punctuation information corresponding to the speech segment.

8. An intelligent speech transcription system, characterized in that, The system includes: A data acquisition module, configured to acquire speech data, and perform voice separation on the speech data to obtain speech segments corresponding to different users; A speech recognition module, configured to determine the text sequence to be processed corresponding to the speech segment, and determine the acoustic feature information of the speech segment through linear prediction analysis; A text correction module, configured to extract features from the text sequence to be processed, and perform correction processing on the text sequence to be processed based on the feature extraction result to obtain a processed text sequence; A result output module, configured to determine the target text sequence including punctuation information corresponding to the speech segment according to the processed text sequence and the acoustic feature information.

9. An intelligent speech transcription device, characterized in that, The device includes: a memory, a processor, and an intelligent speech transcription program stored on the memory and executable on the processor, where the intelligent speech transcription program is configured to implement the steps of the intelligent speech transcription method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, An intelligent speech transcription program is stored on the storage medium, and when the intelligent speech transcription program is executed by a processor, it implements the steps of the intelligent speech transcription method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Conference voice data processing method and device, computer equipment and storage medium

    CN110322872A

  • Speech recognition text error correction method and device, electronic equipment and storage medium

    CN116070595A

  • Speech recognition method and device, electronic equipment and storage medium

    CN116504247A

  • Voice transliteration method and apparatus, and related system and device

    WO2021098637A1