Role perception real-time voice-to-text conversion method and system based on voice model

Through the real-time voice-to-text method based on the voice model, the problem of inability to distinguish the speech content of different speakers in the prior art is solved, and a large number of text processing with high accuracy and real-time performance is achieved, which is suitable for multi-person dialogue scenarios.

CN120126512APending Publication Date: 2025-06-10XIAN TPRI THERMAL CONTROL TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510276680.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing speech recognition technology cannot effectively distinguish the speech content of different speakers, resulting in inaccurate recognition and increased workload for post-processing texts.

Method used

The real-time voice-to-text method based on the voice model is adopted, and the audio data stream is obtained for preprocessing. The voice activity detection model is used to detect the voice activity interval, identify the speaker's identity, and segment the audio when the speaker changes to realize real-time voice-to-text.

Benefits of technology

It improves the accuracy of speech recognition, reduces the workload of post-processing text, and can use voice to text in real time, suitable for multi-person conversation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126512A_ABST
    Figure CN120126512A_ABST
Patent Text Reader

Abstract

The invention discloses a role perception real-time voice-to-text method, system and device based on a voice model, a storage medium and a program, and belongs to the technical field of voice translation. The method comprises the following steps: acquiring and preprocessing an audio data stream to obtain an audio clip; splicing the audio clips, detecting a voice activity interval by using a voice activity detection model, and generating a voice activity list; performing voice similarity comparison on the voice activities in the voice activity list, identifying the identity of a speaker, and segmenting the audio when the speaker changes to obtain segmented audio clips and speaker tags; performing voice recognition on the segmented audio clips to obtain an audio text; and integrating the speaker tag and the corresponding audio text to obtain a final recognition result. According to the invention, the speech content of each speaker in the speech can be accurately distinguished, the speech recognition accuracy is improved, the workload of post text processing is reduced, and speech-to-text conversion can be carried out in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech translation, and particularly relates to a method, system, device, storage medium and program for role-aware real-time speech-to-text based on a speech model. Background Art

[0002] With the rapid development of information technology, electronic devices have penetrated into all aspects of people's lives and become an indispensable part of modern life. When users conduct business through electronic devices, such as filling out electronic forms, querying information, sending and receiving messages, etc., they mainly complete operations by manually inputting text. However, this method has problems of low efficiency and easy errors, especially when a large amount of information needs to be input quickly, this problem is more prominent.

[0003] To solve the limitations of manual input, the industry has introduced speech recognition technology, which converts the user's speech input into text through a speech model to achieve a quick response to business operations. However, most existing speech recognition methods adopt the method of receiving the entire speech input in a whole paragraph and returning the recognition result at one time. This method cannot bring a real-time feedback effect to users. When users are speaking, they cannot immediately see that what they are saying is correctly recognized, which is likely to cause a sense of uncertainty and anxiety. This not only does not conform to the user's usage habits but also limits the application of speech recognition technology in a wider range of scenarios.

[0004] In addition, in some specific usage scenarios, such as meeting records, multi-person conversations, etc., it is far from enough to simply recognize text through speech. In these scenarios, it is often necessary to distinguish the speech content of different speakers, but most existing speech recognition methods cannot effectively distinguish speakers. This results in a large amount of time and effort being required to manually sort out the speech content of different speakers during subsequent text processing, seriously affecting work efficiency. Summary of the Invention

[0005] Aiming at the problem in the prior art that speech recognition cannot effectively distinguish the speech content of different speakers, resulting in inaccurate recognition and an increased workload of post-processing text. The present invention provides a role-aware real-time speech-to-text method based on a speech model, which can accurately distinguish the speech content of each speaker in the speech, improve the accuracy of speech recognition, reduce the workload of post-processing text, and can perform speech-to-text in real time.

[0006] To achieve the above object, the present invention provides the following technical solutions.

[0007] In a first aspect, the present invention provides a role-aware real-time speech-to-text method based on a speech model, including: Obtaining an audio data stream and performing preprocessing to obtain audio segments; Splice the audio clips, use the voice activity detection model to detect the voice activity intervals, and generate a voice activity list; Perform voice similarity comparison on the voice activities in the voice activity list, identify the speaker, and segment the audio when the speaker changes, to obtain the segmented audio segments and speaker labels; Perform speech recognition on the segmented audio segments to obtain audio text; Integrate the speaker labels and the corresponding audio text to get the final recognition result.

[0008] As a further improvement of the present invention, the step of obtaining an audio data stream and preprocessing it to obtain an audio clip includes: The requester initiates a persistent connection request to the server, and the server responds to the persistent connection request sent by the requester, establishing a real-time data transmission channel between the requester and the server; The requesting end sends an audio data stream to the server through a real-time data transmission channel. The server receives the audio data stream and pre-processes it to obtain an audio clip.

[0009] As a further improvement of the present invention, the step of splicing the audio clips, detecting the voice activity intervals using a voice activity detection model, and generating a voice activity list includes: splicing the audio clip data to obtain a spliced ​​combined audio; The voice activity detection model performs voice activity analysis on the spliced ​​combined audio to determine whether there is voice activity in the current combined audio; If there is voice activity, then there is content in the current combined audio that needs to be recognized as text, and then, check whether the number of tags in the voice activity list is more than one; If so, generate the current combined sound to form a voice activity list.

[0010] As a further improvement of the present invention, the voice similarity comparison of the voice activities in the voice activity list, identifying the speaker identity, and segmenting the audio when the speaker changes to obtain the segmented audio segments and speaker labels include: For the voice activity list Latest Voice Activities And the new voice activity Perform similarity comparison; If the similarity between the two is greater than or equal to the minimum audio similarity, the latest voice activity And the new voice activity For the same speaker, there is no need to perform audio segmentation processing. The speaker identity is recorded, the speaker label is generated, and speech recognition is performed to obtain the audio text; If the similarity between the two is less than the minimum audio similarity, the latest speech activity and the penultimate speech activity are from different speakers. When the speaker changes, the audio is segmented, the speaker identity is recorded, a speaker label is generated, and the segmented audio segments and speaker labels are obtained.

[0011] As a further improvement of the present invention, the method includes performing speech recognition on the segmented audio segments to obtain audio texts, including: Preprocessing the segmented audio segments to obtain formatted audio data ; Using an automatic speech recognition model to perform speech recognition on the formatted audio data to obtain the audio text corresponding to the segmented audio segments .

[0012] As a further improvement of the present invention, the method of integrating the speaker labels and the corresponding audio texts to obtain the final recognition result includes: Sorting the text segments in the order of the audio segmentation time, and inserting the speaker and the corresponding text of the cut-off part, the currently speaking speaker and the corresponding text into the formatting result to form a one-to-one corresponding structured answer and obtain the final recognition result.

[0013] In a second aspect, the present invention provides a role-aware real-time speech-to-text system based on a speech model, including: An audio segment acquisition module: used to acquire an audio data stream and perform preprocessing to obtain audio segments; A speech list generation module: used to splice the audio segments, detect the speech activity interval using a speech activity detection model, and generate a speech activity list; An identity label recognition module: used to compare the speech similarities of the speech activities in the speech activity list, recognize the speaker identity, and segment the audio when the speaker changes to obtain the segmented audio segments and speaker labels; An audio text acquisition module: used to perform speech recognition on the segmented audio segments to obtain audio texts; An integrated recognition result module: used to integrate the speaker labels and the corresponding audio texts to obtain the final recognition result.

[0014] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for real-time speech-to-text conversion with role perception based on a speech model are implemented.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the method for real-time speech-to-text conversion with role perception based on a speech model are implemented.

[0016] In a fifth aspect, the present invention provides a computer program product including computer instructions, and when the computer instructions are executed by a processor, the steps of the method for real-time speech-to-text conversion with role perception based on a speech model are implemented.

[0017] Compared with the prior art, the present invention has the following beneficial effects: By acquiring the audio data stream and performing fine preprocessing, the present invention ensures the quality of the audio segments, laying a solid foundation for subsequent accurate recognition. Secondly, the present invention uses a voice activity detection model to detect the voice activity intervals and generate a voice activity list, which improves the processing efficiency and can accurately locate the valid parts in the voice, avoiding the interference of invalid audio on the recognition result. The introduction of the voice activity detection model enables the system to intelligently identify which parts are the speaker's speech and which parts are silence or background noise, thus greatly improving the accuracy and pertinence of recognition. By comparing the speech similarities of the voice activities in the voice activity list, the present invention realizes the identification of the speaker's identity and timely segments the audio when the speaker changes, solving the problem in the prior art that it is impossible to effectively distinguish the speech contents of different speakers, and making the recognition result clearer and more accurate. In the scenario of multi-person conversation, this function is particularly important, which can ensure that the speech of each speaker is accurately recorded, avoiding confusion and misjudgment, and greatly improving the practicality and credibility of speech recognition. In addition, the present invention performs speech recognition on the segmented audio segments to obtain high-quality audio text. Finally, the present invention integrates the speaker labels and the corresponding audio text to obtain the final recognition result, clarifying the identity of each speaker. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are only for the purpose of explanation and are not intended to limit the scope of the disclosure of the present invention in any way. In the drawings: Figure 1 is a flowchart of a method for real-time speech-to-text conversion with role perception based on a speech model according to the present invention; Figure 2Structural connection diagram of a real-time speech-to-text system with role perception based on a speech model according to the present invention; Figure 3 Schematic diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners

[0019] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0021] Aiming at the problem in the prior art that speech recognition cannot effectively distinguish the speech content of different speakers, resulting in inaccurate recognition and increased workload of post-processing text. The present invention provides a role perception real-time speech-to-text method based on a speech model, as Figure 1 shown, the method includes: S100: Obtain an audio data stream and perform preprocessing to obtain audio segments; S200: Concatenate the audio segments, use a voice activity detection model to detect the voice activity interval, and generate a voice activity list; S300: Compare the speech similarity of the speech activities in the voice activity list, identify the speaker's identity, and split the audio when the speaker changes to obtain the split audio segments and speaker labels; S400: Perform speech recognition on the split audio segments to obtain audio text; S500: Integrate the speaker labels and the corresponding audio text to obtain the final recognition result.

[0022] The present invention can accurately distinguish the speech content of each speaker in the speech, improve the accuracy of speech recognition, reduce the workload of post-processing text, and can perform speech-to-text in real time.

[0023] The following further explains the present invention.

[0024] The present invention relates to a method for real-time speech-to-text conversion based on a speech model with role perception, and the specific steps include: S1: Establish a stable long connection interface between the server and the request-side user interface UI for real-time data transmission, continuously receive the user's audio input, and obtain the real-time received audio data.

[0025] Specifically, after the long connection interface receives the connection opening request, first use the connection information to initialize the entire system and clarify the parameter information that remains unchanged during this long connection. This information includes the upper limit of the audio segmentation duration the minimum audio similarity the number of input device channels the sampling bits of the input device the sampling rate of the input device .

[0026] After the initialization is completed, a channel for continuously receiving and sending data is established between the system and the request-side user interface UI.

[0027] During the entire long connection period, as the request side, audio data will be sent to the system at a fixed frequency. The sent data will be received, processed by the main audio processing module, and the result will be returned. This process will continue until the user sends a connection termination request.

[0028] When the connection termination request is received, stop sending and receiving data and end the long connection. Then, clear the cache used during the entire long connection period and release the resources consumed.

[0029] The above is the processing flow method for establishing a stable long connection between the request side and the server side. This step can ensure that requests can be continuously sent and received between the system and the request-side user interface UI.

[0030] S2: For the received real-time audio data, use a voice activity detection model ( model) to analyze and process the speech, and segment it according to speaker change and duration limit to obtain the processed audio data.

[0031] After the stable long connection is established, an audio segment data will be received from the request side at the audio data sending frequency . By combining the analysis results of the previous several rounds to analyze and process this segment, text that meets the requirements can be obtained as the answer. The specific steps are as follows: S21: Append the audio segment data to the end of the combined audio :

[0032] S22: Send the combined audio after splicing to the model for voice activity analysis. The model receives the combined audio as input and outputs a list consisting of voice activity tags :

[0033] These activity tags indicate the intervals where all human voice activities in the combined audio occur, that is, the intervals of several consecutive sounds separated by blanks and noise. Each activity tag has two parameters, representing the start time and the end time of this activity respectively

[0034] Then, each voice activity can be obtained through this time interval information

[0035] S23: Determine whether there is voice activity in the current combined audio :

[0036] If so, it means that there is content in the current combined audio that needs to be recognized as text, and continue with S24.

[0037] If not, it means that there is no content in the current combined audio that needs to be recognized as text, and directly return a null value as the result of this round of processing.

[0038] S24: Check whether the number of tags in the voice activity list is more than one:

[0039] If so, it means that the current combined audio meets the conditions for identifying the speaker. Send the combined audio and the voice activity list to the S3 speaker recognition and segmentation model for processing.

[0040] If not, it means that the current combined audio If the conditions for identifying the speaker are not met, directly jump to S25 responsible for duration detection.

[0041] Voice activity list must have more than one mark in it. Because the combined audio the last voice activity in may be an ongoing speech and is unfinished. And such a voice activity itself has too much uncertainty, so it should be excluded from the analysis and regarded as an invalid activity. Therefore, the voice activity list should have at least two voice activities for the speaker identification and segmentation module to process:

[0042] S25: Determine whether the number of marks in the voice activity list is more than one and has not been segmented due to speaker change analysis:

[0043] If so, it means that the current combined audio meets the conditions for duration analysis. Send the combined audio and the voice activity list to the duration analysis and segmentation module for processing.

[0044] If not, it means that the processing conditions are not met, and jump to S26.

[0045] Repeat to determine whether the number of marks in the voice activity list is more than one. Because the combined audio at this step and the voice activity list , may be after being segmented by S24 or may not have been segmented, so it must be determined again. As for why the number of marks in the voice activity list must be more than one to meet the analysis conditions, it is the same as the reason described in S24.

[0046] S26: Send the second half of the combined audio formed after segmentation to the speech recognition module for recognition.

[0047] S27: Return the formatted result . If processed by the segmentation module, the result will consist of two parts; in the case of no segmentation processing, there is only one part of the result including two situations, and the three situations are respectively: Situation 1: The result without segmentation: A: (A's speech. If A has not completed the current speech, then there is only half a sentence) Case 2: Result without segmentation: B: (Speech of B, if B has not completed the current speech, then only half a sentence) Case 3: Result after segmentation: A: (Complete speech of A) B: (Speech of B, if B has not completed the current speech, then only half a sentence) The above is the process method for processing the input audio segment and returning the formatted text. This step, as the main process module of this system, can splice the audio and perform voice activity analysis on the spliced combined audio to obtain a voice activity list . According to the different situations of the voice activity list , call the appropriate module for corresponding processing, and finally perform speech recognition to return the final answer .

[0048] S3: Identify the speaker's identity by comparing the speech similarity in the audio data, update the speaker information database according to the recognition result, segment the audio and perform speech recognition to ensure accurate speaker annotation.

[0049] When receiving the combined audio that meets the speaker recognition requirements and the voice activity list , the speaker recognition analysis can be started, and when a speaker change is detected, the combined audio is segmented. The specific steps include: S31: Judge whether there are only two voice activities in the combined audio:

[0050] If so, it means that there is only one valid voice activity in the current combined audio , that is, the voice activity recorded by the first activity marker in the voice activity list from the start time to the end time : :

[0051] Then this voice activity needs to be compared with the previously recorded latest speaker voice activity , mark this voice activity as the latest voice activity , and mark the voice activity in the record as the second latest voice activity 。

[0052] If not, it indicates that the current combined audio contains multiple valid voice activities. Then it is necessary to compare the last valid voice activity in the combined audio :

[0053] and the penultimate valid voice activity :

[0054] Mark the last voice activity as the latest voice activity , and the penultimate voice activity is marked as the second latest voice activity .

[0055] It should be noted that only the latest valid voice activity and the second latest valid voice activity are compared. The reason is that if such a comparison is made in each round, it can ensure that all voice activities are compared during the entire long connection period. For example: when the combined audio contains two valid voice activities, then the first and second valid voice activities will be compared; if in case ii, there is always one user speaking, that is, the case where no segmentation is required, then when the number of valid voice activities in the combined audio increases to three, the similarity between the second and third valid voice activities will be compared, and so on; similarly, for case i, where the speaker changes, that is, the case where segmentation is required.

[0056] S32: Use the speaker recognition model ( Model) to compare the similarity between the latest voice activity and the second latest valid voice activity . The model takes two audio segments as input and outputs a percentage number as an indication of the similarity between the two:

[0057] Compare the similarity between the two and the lowest audio similarity : If the similarity between the two is greater than or equal to the lowest audio similarity :

[0058] It indicates the latest valid voice activity and the second latest valid voice activity are from the same speaker, and no splitting process is required. The processing of this module can be directly ended.

[0059] If the similarity between the two is less than the minimum audio similarity :

[0060] It indicates the latest valid voice activity and the second latest valid voice activity are from different speakers, and S34 needs to be continued to determine whether splitting is required.

[0061] S34: Determine whether the current latest voice activity is the first valid voice activity in the combined audio , or the last valid voice activity in the combined audio : If it is the first valid voice activity , it means that the combined audio in this round starts with the speech of a new speaker. No splitting process is required, but it is necessary to jump to S36 for the identity confirmation of the new speaker.

[0062] If it is the last valid voice activity , it means that the combined audio in this round starts with the previous speaker and changes the speaker in the middle. Splitting is required, and the identity of the new speaker needs to be confirmed.

[0063] S35: Use the start time of the last valid voice activity as the splitting point. Split the combined audio into two parts. One part is the speech audio of the previous speaker :

[0064] and mark the speaker of this part of the speech audio with the speaker identity in the record :

[0065] The other part is the speech audio of the latest speaker, that is, the new combined audio :

[0066] S36: For the new combined audio obtained through segmentation processing, or the original combined audio without processing , the effective speech activity of the new speaker is subject to speaker recognition.

[0067] First, compare this effective speech activity one by one with the example audios of all speakers in the speaker database in the manner described in S32: If none of the example audios of the speakers in the speaker database has a similarity degree with this effective speech activity exceeding the minimum audio similarity :

[0068] It can be determined that this new speaker has not spoken before and is a newly discovered speaker. Add the identity of this speaker , as well as the effective speech activity serving as the example audio to the speaker database :

[0069] And update the speaker speech activity in the record to the latest effective speech activity :

[0070] And update the speaker identity in the record to the latest speaker identity :

[0071] If at least one of the example audios of the speakers in the speaker database has a similarity degree with this effective speech activity exceeding the minimum audio similarity :

[0072] It can be determined that this new speaker has spoken before and is a discovered speaker. Select the speaker with the highest similarity degree:

[0073] Update the speaker speech activity in the record Sample audio for this speaker :

[0074] And update the speaker identity in the record For the identity of this speaker :

[0075] The above is the processing flow method for speaker recognition and segmentation based on the recognition result. This step can be based on the combined audio The speech activity existing in The specific situation of the combined audio Perform segmentation. At the same time, identify the identity of the new speaker and record it

[0076] S4: After receiving the combined audio that meets the duration analysis conditions, perform duration analysis on the combined audio, and use the upper limit of the audio segmentation duration set when the connection is opened to determine whether the combined audio meets the segmentation conditions

[0077] When receiving the combined audio that meets the duration analysis requirements And the speech activity list , you can start the duration analysis, and when it is detected that the duration of the combined audio Is longer than the upper limit of the audio segmentation duration , perform segmentation on the combined audio . The specific steps are as follows S41: Check the last activity marker in the speech activity list Note: The activity marker marked here is not the last valid speech activity described in S3 , but the real last speech activity , and whether the end time Exceeds the upper limit of the audio segmentation duration :

[0078] If so, it means that the combined audio Has exceeded the range that can be safely used as a processing object and should be segmented

[0079] If not, it means that the combined audio Is still within the safe range and does not need to be segmented, and directly exit the current module

[0080] S42: Use the penultimate speech activity The start time is used as the segmentation point. The combined audio is segmented into two parts. The first part contains a series of complete speech activities, which, like in S35, are marked as :

[0081] Mark the speaker of this part of the audio as the current speaker :

[0082] The second part is the ongoing speech activity, that is, the new combined audio :

[0083] Since it is necessary to wait until the complete number of rounds of this part of the speech activity to determine the true speaker identity, so for this round, it is temporarily assumed that the speaker remains unchanged and there is no need to update the latest speaker in the record:

[0084] The above is the processing flow method for performing duration analysis and segmenting according to the analysis results. This step can segment the combined audio based on whether its duration exceeds the audio segmentation duration upper limit , and segment the combined audio .

[0085] S5: By analyzing attributes such as audio operations, the number of channels, sampling bits, and sampling rate of the processed audio data, and sending it into an automatic speech recognition model (ASR model) for speech recognition, finally insert the speaker and the corresponding text information into the result.

[0086] Sort out the operations performed in this round, and according to different actual situations, recognize the audio recorded in this round and convert it into text. Finally, insert one or more groups of corresponding speakers and the corresponding text into the final answer. The specific steps include: S51: First, judge whether segmentation has been performed in this round:

[0087] If so, it means that the final result needs to include two parts. One part is the speaker of the segmented part and the corresponding complete speech . So continue with S52 to recognize this part of the speech .

[0088] If not, it means that the combined audio It is still the most primitive combined audio, and only a part of the results will be included in the final result. Jump to S53 to perform speech recognition on this part of the speech recognition.

[0089] S52: Use the number of channels of the input device and the sampling bit depth of the input device and the sampling rate of the input device to preprocess the segmented audio to form formatted audio data , and then send it for speech recognition using the model. The model receives a segment of audio data and outputs a corresponding text. By calling this model, a text representing the content of the segmented part of the audio is obtained . .

[0090] S53: Use the number of channels of the input device and the sampling bit depth of the input device and the sampling rate of the input device to preprocess the combined audio to form formatted audio data , and then send it for speech recognition using the model. By calling this model, a text representing the content of the combined audio is obtained .

[0091] S54: If the text of the segmented part exists, that is, when the combined audio has been segmented, the speaker of the segmented part and the corresponding text should first be inserted into the formatted result :

[0092] After that, the speaker who is currently speaking and the corresponding text should be inserted into the formatted result :

[0093] It should be noted that depending on whether the segmentation is due to a change in the speaker, the speaker of the segmented part may be the same as the current speaker :

[0094] The above is the process method of performing speech recognition on the audio according to the situation and integrating it into the final result. This step can perform reasonable speech recognition processing according to whether segmentation has been done and the reason for segmentation, and form a structured answer corresponding one by one with all the results in order, and use the answer to form a formatted result. .

[0095] Through the method and system for real-time speech to text with role perception based on a speech model, the present invention can achieve an accuracy rate of 95% in speech recognition to text, can solve the problem that users need to manually input to carry out various information input services based on electronic devices, and has simple invocation, high accuracy rate, and high general performance.

[0096] The common use of the real-time communication module and the two segmentation modules in the present invention can support the system to stably receive requests and return results at a frequency greater than 0.5 seconds. It can solve the problem that traditional speech models can only support whole-segment speech input and one-time text output, and cannot provide real-time effect feedback to users.

[0097] The speaker recognition and segmentation module in the present invention can accurately distinguish the speech content of each speaker. In the usage scenario with less than five people, the accuracy rate of speaker discrimination reaches 85%. It can solve the problem that in some business scenarios, the speech content of different speakers cannot be distinguished, thus increasing the workload of post-processing text.

[0098] In summary, the method and system provided by the present invention are a method and system for real-time speech to text with role perception based on a speech model, and have high accuracy in generating text and distinguishing roles. In addition, the method provided by the present invention can be applied to any scenario that requires text information interaction using an electronic device, and has good general performance.

[0099] The second object of the present invention is to propose a system for real-time speech to text with role perception based on a speech model, as Figure 2 shown, including: Audio segment acquisition module 100: used to acquire the audio data stream and perform preprocessing to obtain audio segments; Speech list generation module 200: used to splice the audio segments, detect the speech activity interval using a voice activity detection model, and generate a speech activity list; Identity label recognition module 300: used to perform speech similarity comparison on the speech activities in the speech activity list, identify the speaker identity, and segment the audio when the speaker changes to obtain the segmented audio segments and speaker labels; Audio text acquisition module 400: used to perform speech recognition on the segmented audio segments to obtain audio text; Integrated Recognition Result Module 500: It is used to integrate the speaker label and the corresponding audio text to obtain the final recognition result.

[0100] As Figure 3 shown, the third object of the present invention is to provide an electronic device, which includes: a processor 601, a memory 602, and a display screen 603. Among them, the memory 602 and the display screen 603 are both connected to the processor 601, such as being connected through a bus 604. Optionally, the electronic device may further include a transceiver 605. It should be noted that in practical applications, the transceiver 605 is not limited to one, and the structure of the electronic device does not constitute a limitation to the embodiments of the present application.

[0101] The processor 601 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 601 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0102] The bus 604 may include a path for transmitting information between the above components. The bus 604 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 604 may be divided into an address bus, a data bus, a control bus, etc.

[0103] The memory 602 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0104] The memory 602 is used to store the application program code for executing the solution of this application, and is controlled by the processor 601 for execution. The processor 601 is used to execute the application program code stored in the memory 602 to implement the content shown in the foregoing method embodiments.

[0105] Figure 3 The illustrated electronic device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0106] The fourth object of the present invention is to provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements each process of the method embodiment as described above. Figure 1 For example, a memory including instructions, and the above instructions can be executed by the processor of the electronic device to complete the above method.

[0107] A computer-readable storage medium can be a tangible device that holds and stores instructions used by an instruction execution device. A computer-readable storage medium can be, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above. Specifically, a computer-readable storage medium can be a portable computer disk, a hard disk, a USB flash drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, an optical disc, a magnetic disk, a mechanical encoding device, and any combination of the above.

[0108] The fifth objective of the present invention is to provide a computer program product, including computer instructions, which implement the various processes of the method embodiments shown above when executed by a processor, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Figure 1 The above method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0109] Upon reading the above description, many embodiments and many applications beyond the provided examples will be obvious to those skilled in the art. Therefore, the scope of this teaching should not be determined with reference to the above description, but rather should be determined with reference to the full scope of the foregoing claims and the equivalents thereof. For the sake of completeness, all articles and references, including patent applications and published disclosures, are incorporated herein by reference. The omission of any aspect of the subject matter disclosed herein in the foregoing claims is not intended to abandon such subject matter, nor should it be considered that the applicant has not considered such subject matter to be part of the disclosed inventive subject matter.

[0110] The above is a further detailed description of the present invention. It cannot be determined that the specific implementation of the present invention is limited to this. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope determined by the claims submitted for the present invention.

Claims

1. A role-aware real-time speech-to-text method based on speech model, characterized in that: include: Acquire the audio data stream and preprocess it to obtain the audio clip; Splice the audio clips, use the voice activity detection model to detect the voice activity intervals, and generate a voice activity list; Perform voice similarity comparison on the voice activities in the voice activity list, identify the speaker, and segment the audio when the speaker changes, to obtain the segmented audio segments and speaker labels; Perform speech recognition on the segmented audio segments to obtain audio text; Integrate the speaker labels and the corresponding audio text to get the final recognition result.

2. The method for character-aware real-time speech-to-text conversion based on speech model according to claim 1, characterized in that: The step of obtaining an audio data stream and preprocessing the audio data stream to obtain an audio clip includes: The requester initiates a persistent connection request to the server, and the server responds to the persistent connection request sent by the requester, establishing a real-time data transmission channel between the requester and the server; The requesting end sends an audio data stream to the server through a real-time data transmission channel. The server receives the audio data stream and pre-processes it to obtain an audio clip.

3. The method for character-aware real-time speech-to-text conversion based on speech model according to claim 1, characterized in that: The step of splicing the audio clips, detecting the voice activity intervals using a voice activity detection model, and generating a voice activity list includes: splicing the audio clip data to obtain a spliced ​​combined audio; The voice activity detection model performs voice activity analysis on the spliced ​​combined audio to determine whether there is voice activity in the current combined audio; If there is voice activity, then there is content in the current combined audio that needs to be recognized as text, and then, check whether the number of tags in the voice activity list is more than one; If so, generate the current combined sound to form a voice activity list.

4. The method for character-aware real-time speech-to-text conversion based on speech model according to claim 1, characterized in that: The voice activities in the voice activity list are compared for voice similarity, the speaker identity is identified, and the audio is segmented when the speaker changes, to obtain segmented audio segments and speaker labels, including: For the voice activity list Latest Voice Activities And the new voice activity Perform similarity comparison; If the similarity between the two is greater than or equal to the minimum audio similarity, the latest voice activity And the new voice activity For the same speaker, there is no need to perform audio segmentation processing. The speaker identity is recorded, the speaker label is generated, and speech recognition is performed to obtain the audio text; If the similarity between the two is less than the minimum audio similarity, the latest voice activity And the new voice activity The audio is from different speakers. When the speaker changes, the audio is segmented, the speaker identity is recorded, and the speaker label is generated to obtain the segmented audio clips and speaker labels.

5. The method for character-aware real-time speech-to-text conversion based on speech model according to claim 1, characterized in that: The method includes performing speech recognition on the segmented audio segments to obtain audio text, including: Preprocess the segmented audio segments to obtain formatted audio data ; Use automatic speech recognition models to format audio data Perform speech recognition to obtain the audio text corresponding to the segmented audio segment .

6. The method for character-aware real-time speech-to-text conversion based on speech model according to claim 1, characterized in that: The integration of the speaker label and the corresponding audio text to obtain the final recognition result includes: Sort the text segments according to the time sequence of the audio segmentation, and separate the speakers of the segmented parts. and the corresponding text , the current speaker And the corresponding text Insert formatted results , forming a one-to-one corresponding structured answer and forming the final recognition result.

7. A role-aware real-time speech-to-text system based on speech model, characterized in that: include: Audio segment acquisition module: used to acquire audio data stream and pre-process it to obtain audio segments; The voice list generation module is used to splice audio clips, detect voice activity intervals using a voice activity detection model, and generate a voice activity list; Identification label module: used to compare the voice similarity of the voice activities in the voice activity list, identify the speaker, and segment the audio when the speaker changes, and obtain the segmented audio segments and speaker labels; The module for obtaining audio text is used to perform speech recognition on the segmented audio segments to obtain audio text; Integration recognition result module: used to integrate speaker labels and corresponding audio texts to obtain the final recognition results.

8. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of a character-aware real-time speech-to-text conversion method based on a speech model as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the character-aware real-time speech-to-text conversion method based on a speech model as described in any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of a character-aware real-time speech-to-text conversion method based on a speech model as described in any one of claims 1 to 6.