Voice recording method, related device, equipment and medium
By acquiring and integrating real-time operational data into a speech transcription method, the problem of transcription accuracy under personalized needs in existing technologies has been solved, achieving higher transcription accuracy.
Patent Information
- Application Number
- CN202510825797.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-17
AI Technical Summary
Existing speech recognition transcription technology is difficult to improve the accuracy of transcription records while meeting personalized needs.
By performing real-time transcription based on the speech stream, the effective data of real-time operations is obtained as reference data, the transcribed content in the recording interface is updated in real time, and the effective data of real-time operations is incorporated into subsequent transcription processes to improve transcription accuracy.
While meeting individual needs, it significantly improves the accuracy of transcription records.
Smart Images

Figure CN120808786A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice processing, and particularly relates to a voice recording method and related device, equipment and medium. BACKGROUND
[0002] With the continuous development of information technology, voice data is widely used in many fields such as daily life and business office. For example, by recognizing and transcribing voice data, recognized text can be obtained, so that recording can be performed in the form of text.
[0003] At present, although the existing recognition and transcription technology has made great progress after years of development. For example, a neural network model based on deep learning can significantly improve the accuracy of speech recognition. However, it is still difficult to improve the accuracy of transcription recording as much as possible under the premise of meeting individual needs. SUMMARY
[0004] The technical problem solved by the present application is to provide a voice recording method and related device, equipment and medium, which can improve the accuracy of transcription recording as much as possible under the premise of meeting individual needs.
[0005] In order to solve the above technical problem, the first aspect of the present application provides a voice recording method, comprising: real-time transcription based on a voice stream to obtain a first recognized text; wherein the first recognized text contains a plurality of first subtexts and a speaker label to which each first subtext belongs; updating the transcription content in the recording interface in real time based on the first recognized text; in response to a real-time operation triggered on the recording interface, obtaining the effective data of the real-time operation as reference data for real-time transcription; continuing real-time transcription on the voice stream based on the reference data to obtain a second recognized text; wherein the second recognized text contains a plurality of second subtexts and a speaker label to which each second subtext belongs, and the transcription content in the recording interface is updated based on the second recognized text.
[0006] To solve the above technical problems, the second aspect of the present application provides a voice recording method, comprising: performing real-time transcription based on a voice stream to obtain a first recognized text; wherein the first recognized text comprises a plurality of first subtexts and a speaker label to which each first subtext belongs; updating the transcription content in a recording interface in real time based on the first recognized text; obtaining effective data of a real-time operation triggered on the recording interface; in response to detecting overall transcription triggered on the voice stream, obtaining voice data from the start of recording to the triggering of overall transcription as to-be-recognized voice; performing overall transcription on the to-be-recognized voice based on reference data of the overall transcription to obtain a third recognized text; wherein the third recognized text comprises a plurality of third subtexts and a speaker label to which each third subtext belongs, the reference data of the overall transcription comprises effective data of each real-time operation triggered in the process from the start of recording to the triggering of overall transcription of the voice stream, and the transcription content in the recording interface is updated based on the third recognized text.
[0007] To solve the above technical problems, the third aspect of the present application provides a voice recording device, comprising: a first transcription module, an interface updating module, a reference obtaining module, and a second transcription module. The first transcription module is configured to perform real-time transcription based on a voice stream to obtain a first recognized text; wherein the first recognized text comprises a plurality of first subtexts and a speaker label to which each first subtext belongs. The interface updating module is configured to update the transcription content in a recording interface in real time based on the first recognized text. The reference obtaining module is configured to obtain effective data of a real-time operation triggered on the recording interface as reference data for real-time transcription in response to the real-time operation. The second transcription module is configured to continue to perform real-time transcription on the voice stream based on the reference data to obtain a second recognized text; wherein the second recognized text comprises a plurality of second subtexts and a speaker label to which each second subtext belongs, and the transcription content in the recording interface is updated based on the second recognized text.
[0008] To solve the above technical problems, the fourth aspect of the present application provides a voice recording device, comprising: a real-time transcription module, an interface updating module, a reference obtaining module, a voice obtaining module, and a whole transcription module. The real-time transcription module is configured to perform real-time transcription based on a voice stream to obtain a first recognized text. The first recognized text includes a plurality of first subtexts and a speaker label to which each first subtext belongs. The interface updating module is configured to update the transcription content in the recording interface in real time based on the first recognized text. The reference obtaining module is configured to obtain effective data triggered by real-time operations of the recording interface. The voice obtaining module is configured to obtain voice data from the start of recording to the triggering of whole transcription of the voice stream as to-be-recognized voice in response to detecting the whole transcription triggered by the voice stream. The whole transcription module is configured to perform whole transcription on the to-be-recognized voice based on reference data of the whole transcription to obtain a third recognized text. The third recognized text includes a plurality of third subtexts and a speaker label to which each third subtext belongs. The reference data of the whole transcription includes effective data of each real-time operation triggered in the process from the start of recording to the triggering of whole transcription of the voice stream. The transcription content in the recording interface is updated based on the third recognized text.
[0009] To solve the above technical problems, the fifth aspect of the present application provides an electronic device, comprising at least a memory and a processor coupled with each other. The memory stores at least program instructions. The processor is configured to execute the program instructions to implement the voice recording method in the first aspect or the second aspect.
[0010] To solve the above technical problems, the sixth aspect of the present application provides a computer-readable storage medium, which stores program instructions capable of being executed by a processor. The program instructions are configured to implement the voice recording method in the first aspect or the second aspect.
[0011] The above scheme is based on the voice stream to obtain the first recognition text in real time, and the first recognition text includes a plurality of first subtexts and a speaker label to which each first subtext belongs. The first recognition text is used to update the transcription content in the recording interface in real time, so as to obtain the effective data of the real-time operation triggered in the recording interface as the reference data of real-time transcription, and then continue to transcribe the voice stream in real time based on the reference data to obtain the second recognition text, and the second recognition text includes a plurality of second subtexts and a speaker label to which each second subtext belongs. The transcription content in the recording interface is updated based on the second recognition text. Therefore, on the one hand, the real-time operation triggered in the recording interface can be accepted in the real-time transcription process, which helps to meet the personalized needs. On the other hand, after accepting the real-time operation, the real-time transcription of the subsequent voice stream is further fed back according to the effective data of the real-time operation. Compared with the conventional recognition transcription technology, the accuracy of the transcription record can be improved as much as possible. Therefore, the accuracy of the transcription record can be improved as much as possible on the premise of meeting the personalized needs. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a flowchart of an embodiment of the voice recording method of the present application; Figure 2 is a schematic diagram of an embodiment of the recording interface of the present application; Figure 3 is a flowchart of an embodiment of the homework correction method of the present application; Figure 4 is a schematic diagram of an embodiment of the voice recording device of the present application; Figure 5 is a schematic diagram of an embodiment of the voice recording device of the present application; Figure 6 is a schematic diagram of an embodiment of the electronic device of the present application; Figure 7 is a schematic diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION
[0013] The schemes of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0014] In the following description, specific details such as specific system structures, interfaces, techniques, etc. are presented in order to thoroughly understand the present application, but are not intended to limit the present application.
[0015] The terms "system" and "network" are often used interchangeably herein. The term "and / or", merely describes association between associated objects, indicates that there can be three cases: A and / or B, A alone, and B alone. In addition, the segment " / " herein generally indicates that the associated objects before and after are in an "or" relationship. In addition, "multiple" herein means two or more than two.
[0016] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the voice recording method of the present application. Specifically, it can include the following steps: Step S11: Real-time transcription based on the voice stream to obtain the first recognized text.
[0017] In the embodiments of the present disclosure, the first recognized text contains a plurality of first subtexts and a speaker label to which each first subtext belongs. Taking a conference scenario as an example, the speaker label can include but is not limited to "Speaker 1", "Speaker 2", "Speaker 3", etc. Alternatively, reference voices of different speakers and corresponding speaker labels can be pre-set, and the speaker label to which each first subtext in the first recognized text belongs can be the pre-set speaker label. For example, still taking the conference scenario as an example, the reference voices of speakers "Zhang San", "Li Si", and "Wang Wu" and the corresponding speaker labels (i.e., "Zhang San", "Li Si", and "Wang Wu") can be pre-set, and the speaker label to which the first subtext in the first recognized text belongs is one of "Zhang San", "Li Si", and "Wang Wu". Of course, the above examples are only a few possible examples in actual application, and other possible cases are not limited herein, and will not be exemplified one by one. In addition, real-time transcription can be realized by a deep learning speech recognition model, or it can also be realized by a large language model, and the implementation of real-time transcription is not limited herein.
[0018] In one implementation scenario, the voice stream can be transcribed in real time based on a deep learning neural network model. It should be noted that the deep learning neural network model can include but is not limited to a long short-term memory network, a recurrent neural network, etc., and the network structure of the neural network model is not limited herein. Alternatively, the voice stream can be transcribed in real time based on a large language model. It should be noted that the large language model can include but is not limited to open source large models such as Llama and Bloom, or the large language model can be obtained by fine-tuning the open source large model based on a specific data set (such as a voice data set), or the large language model can also be a custom large model, and the specific source of the large language model is not limited herein. Of course, the above examples are only a few possible examples of real-time transcription, and the implementation of real-time transcription is not limited herein, and will not be exemplified one by one.
[0019] In one implementation scenario, the first recognized text can take the form of a combination of "speaker label + first subtext". Still taking the aforementioned meeting scenario as an example, the first recognized text can include: "Zhang San: XXXXX (i.e. the first subtext belonging to speaker Zhang San); Li Si: XXXXX (i.e. the first subtext belonging to speaker Li Si); Wang Wu: XXXXX (i.e. the first subtext belonging to speaker Wang Wu); Li Si: XXXXX (i.e. the first subtext belonging to speaker Li Si)". Of course, the above example is only one possible example of the first recognized text in actual application, and the specific content of the first recognized text is not limited herein, nor will it be exemplified one by one.
[0020] Step S12: Based on the first recognized text, updating the transcribed content in the recording interface in real time.
[0021] Specifically, in the real-time transcription of the speech stream, the first recognized text can be updated in real time, and accordingly the first recognized text can be updated in the recording interface to follow the continuous output of the speech stream, and the transcribed content of the speech stream in the recording interface is updated accordingly. Still taking the aforementioned meeting scenario as an example, when speaker Zhang San continues to speak, the speech stream continues to output, at which time the first recognized text is updated to include the content "Zhang San: XXXXX (i.e. the new first subtext belonging to speaker Zhang San)", and then the transcribed content in the recording interface can be updated in real time based on this.
[0022] Step S13: In response to the real-time operation triggered in the recording interface, obtaining the effective data of the real-time operation as reference data for real-time transcription.
[0023] In one implementation scenario, the real-time operation triggered in the recording interface can be image insertion, and at this time the effective data of the real-time operation can include but is not limited to the recognized content of the inserted image, the insertion position, etc., which are not limited herein. Accordingly, the reference data for real-time transcription can also include the recognized content of the inserted image, the insertion position, etc., which are not limited herein. For example, the recognized content of the inserted image and the insertion position can be obtained as reference data. Still taking the aforementioned meeting scenario as an example, when speaker Zhang San speaks "please look at the PPT, at the current stage of artificial intelligence... ", the first recognized text is also updated accordingly through the aforementioned real-time transcription, at which time the real-time operation "image insertion" can be triggered in the recording interface to insert an image such as a screenshot of the meeting PPT (the screenshot can introduce related content about the current stage of artificial intelligence) in the transcribed content of the recording interface. It should be noted that in the real-time operation "image insertion", not only the inserted image can be specified, but also the insertion position can be specified, so that the recognized content of the inserted image and the insertion position can be obtained as reference data.
[0024] In an implementation scenario, the real-time operation triggered in the recording interface can also be the speaker modification, and in this case, the effective data of the real-time operation can include but is not limited to a speech segment to which the first subtext targeted by the speaker modification belongs, a speaker label in which the speaker modification takes effect, and the like, without limitation. In order to facilitate distinction, the speech segment to which the first subtext targeted by the speaker modification belongs can be taken as a first target segment, and the speaker label in which the speaker modification takes effect can be taken as a target label. That is, the speech segment to which the first subtext targeted by the speaker modification belongs can be taken as the first target segment, and the speaker label in which the speaker modification takes effect can be taken as the target label. Still taking the aforementioned conference scenario as an example, when the first recognition text continues to add the transcribed content “Zhang San: Please look at the PPT, in the current stage of artificial intelligence …”, if it is found that the speaker label to which the first subtext belongs is incorrect, the real-time operation “speaker modification” can be triggered in the recording interface (for example, the real-time operation “speaker modification” can be triggered by clicking the speaker label of the first subtext, in this case, the speaker label of the first subtext can be switched to an editable state to wait for entering a new speaker label, that is, the target label), so as to modify the speaker label “Zhang San” of the first subtext to the speaker label specified by the real-time operation “speaker modification” in the recording interface, such as “Wang Wu” and the like. It should be noted that the real-time operation “speaker modification” can not only specify the speaker label to which the first subtext belongs, but also specify the speaker label in which the speaker modification takes effect, so as to obtain the speech segment to which the first subtext targeted by the speaker modification belongs (that is, the aforementioned first target segment) and the speaker label in which the speaker modification takes effect (that is, the aforementioned target label) as the reference data.
[0025] In an implementation scenario, the real-time operation triggered in the recording interface can also be a transcription modification, and the effective data of the real-time operation can include, but is not limited to, an original word or phrase of the real-time transcription to which the transcription modification is directed, a target word or phrase to which the transcription modification is effective, and the like, without limitation. For example, a reference word or phrase pair formed by the original word or phrase of the real-time transcription to which the transcription modification is directed and the target word or phrase to which the transcription modification is effective can be obtained as reference data. Still taking the conference scenario as an example, when the first identified text continues to add the transcription content “Zhang San: Please look at the PPT, in the current stage of artificial intelligence...”, if it is found that the original word or phrase “PPT” in the first subtext does not conform to the written specification, a real-time operation “transcription modification” can be performed on the original word or phrase “PPT” in the recording interface, such as specifying a target word or phrase “presentation” to which the real-time operation “transcription modification” is effective. Of course, the above example is only one possible example in the actual application process, and the specific purpose of performing the “transcription modification” and the modification content are not limited herein (such as, the specific purpose can also be to correct the transcription error of the original word or phrase, and the like), and will not be exemplified one by one. It should be noted that the real-time operation “transcription modification” can specify not only the original word or phrase, but also the target word or phrase. In order to facilitate subsequent processing, the original word or phrase to which any real-time operation “transcription modification” is directed and the target word or phrase to which the real-time operation is effective can form a reference word or phrase pair to obtain reference data.
[0026] In an implementation scenario, the real-time operation triggered in the recording interface can also be note recording, and the effective data of the real-time operation can include, but is not limited to, recording content of the note recording, recording time of the note recording, and the like, without limitation. For example, the recording content and the recording time of the note recording can be obtained as reference data. Still taking the conference scenario as an example, when the speaker “Zhang San” speaks “Please look at the PPT, in the current stage of artificial intelligence...” at time t, if the real-time operation triggered in the recording interface is note recording, the recording content of the note recording (such as “According to Zhang San’s PPT, AI is in the current stage...”) and the recording time of the note recording (such as time t) can be obtained as the effective data of the real-time operation “note recording”, that is, as reference data. Of course, the above example is only one possible example in the actual application process, and the specific content of the effective data when performing “note recording” is not limited herein, and will not be exemplified one by one.
[0027] It should be noted that the image insertion, speaker modification, transcription modification, and note recording are only several possible examples of real-time operations triggered in the recording interface in actual application, and the specific types of real-time operations are not limited herein, nor are they exemplified one by one. For example, in actual application, the real-time operation can be any of the above, or a combination of two or more of the above, such as first triggering image insertion in the recording interface, and then triggering transcription modification in the recording interface. Alternatively, first triggering transcription modification in the recording interface, and then triggering speaker modification in the recording interface, which will not be exemplified one by one herein. In addition, please refer to Figure 2 , Figure 2 is a schematic diagram of an embodiment of the recording interface of the present application. As shown in Figure 2 , in the case where the real-time operation includes note recording, the recording content of the note recording and the transcription content can be displayed in separate zones in the recording interface. For example, Figure 2 , the transcription content (left side) and the recording content (right side) in the recording interface can be displayed in left and right zones. Of course, other partitioning methods (such as, top and bottom partitioning, etc.) can also be used, which are not limited herein, nor are they exemplified one by one. Alternatively, in some embodiments, the transcription content and the recording content can also be displayed without partitioning, i.e., the recording interface can also not be partitioned, i.e., the recording interface can be only one partition as a whole, and at this time the transcription content and the recording content can be in the same display area. Please continue to refer to Figure 2 , the time axis of the voice stream and the recording interface can be displayed on the same screen, and the transcription content in the recording interface can be associated with the timestamp information of the corresponding voice segment, and in response to the point selection operation triggered on the time axis, the voice stream and the transcription content in the recording interface can jump to the corresponding time of the point selection operation. For example, Figure 2 , the transcription content "Ladies and gentlemen, today we are gathered here" can be associated with the timestamp information "00:00:00" of the corresponding voice segment, the transcription content "Yes, here we also want to thank the organizers" can be associated with the timestamp information "00:02:10" of the corresponding voice segment, and the transcription content "So, our meeting today officially begins" can be associated with the timestamp information "00:03:40" of the corresponding voice segment. When the corresponding time of the point selection operation on the time axis is "00:02:10", the transcription content in the recording interface can jump to "So, our meeting today officially begins", and the voice stream can also jump to "So, our meeting today officially begins". Of course, the above example is only one possible example of the transcription content and the voice stream jumping in response to the point selection operation on the time axis, and other possible cases are not limited herein, nor are they exemplified one by one.
[0028] Step S14: continue real-time transcription of the voice stream based on the reference data to obtain a second recognition text.
[0029] In the embodiments of the present disclosure, the second recognized text can include a plurality of second subtexts and a speaker label to which each second subtext belongs, and the transcribed content in the recording interface can be updated based on the second recognized text. That is, after obtaining the second recognized text, the transcribed content in the recording interface can be updated again based on the second recognized text. In this way, in the process of continuous input of the speech stream, real-time transcription, recording interface updating, and further real-time transcription of subsequent speech streams based on the effective data of real-time operation when there is real-time operation can be continuously performed in a loop until the speech stream stops inputting or recording ends. It should be noted that when the real-time transcription is continued based on the reference data, the second subtext and the speaker label to which the second subtext belongs can be determined based on the existing manner of real-time transcription. For example, when the real-time transcription is originally implemented based on a large language model, the second subtext and the speaker label to which the second subtext belongs can also be determined based on the large language model. On this basis, if the operation type of the real-time operation (such as image insertion, speaker modification, note recording, etc.) affects the content of the subtext, the second subtext is further determined based on the existing manner and further integrated with the reference data. Or, if the operation type of the real-time operation (such as speaker modification) affects the speaker label to which the subtext belongs, the speaker label to which the second subtext belongs is further determined based on the existing manner and further integrated with the reference data. For ease of understanding, the specific process of continuing real-time transcription of the speech stream based on the effective data of the real-time operation will be exemplarily described below by taking the real-time operation as image insertion, speaker modification, transcription modification, and note recording, respectively.
[0030] In one implementation scenario, as described above, when the real-time operation is image insertion, the recognized content of the inserted image and the insertion position can be obtained as reference data. On this basis, in response to the current first speech segment to be transcribed being adjacent to the reference speech segment in the process of continuing real-time transcription, the first speech segment can be transcribed based on the recognized content to obtain a second subtext corresponding to the first speech segment. It should be noted that the reference speech segment can be the speech segment to which the first subtext located at the insertion position belongs. Please refer to Figure 2When the speaker "Wang Wu" speaks "So, our meeting today officially begins", the inserted image (which can contain the summary of the next speech of the speaker "Wang Wu") can be obtained by shooting the conference large screen (such as the display of the speaker "Wang Wu" presentation on the conference large screen), screen capture (such as the display of the speaker "Wang Wu" presentation on the communication software of the participants) and the like, and the inserted image is placed below the transcribed content "So, our meeting today officially begins". At this time, the recognition content of the inserted image can be obtained by recognizing the inserted image through relationship character recognition technology and the like, and the insertion position is below the speech "So, our meeting today officially begins" of the speaker "Wang Wu". That is, in this example, the first subtext of the insertion position can be "So, our meeting today officially begins", and the reference voice segment is the voice segment to which the first subtext "So, our meeting today officially begins" belongs. The current first voice segment to be transcribed in the real-time transcription process can be the continued speech of the speaker "Wang Wu", such as "This meeting contains four topics, which are …". The first voice segment can be transcribed with the assistance of the recognition content of the aforementioned inserted image to obtain the second subtext corresponding to the first voice segment. As a possible implementation, the prompt instruction can be constructed based on the first voice segment and the recognition content of the inserted image, and the prompt instruction is used to instruct the large language model to recognize the first voice segment with the assistance of the recognition content of the inserted image. The output content of the large language model in response to the prompt instruction can be obtained as the second subtext corresponding to the first voice segment. The above-mentioned method uses the recognition content of the inserted image to transcribe the first voice segment adjacent to the reference voice segment in the continued real-time transcription process, which can help improve the transcription accuracy of the first voice segment. For example, in the aforementioned example, if the first voice segment contains irregular words, professional terms and the like, such as "metatarsal" (which is easily misrecognized as "finger bone"), the possibility of correct recognition can be improved as much as possible with the assistance of the image content of the inserted image (such as "metatarsal is the long bone connecting the ankle and the toes of the foot"), and the possibility of recognizing and transcribing "metatarsal" in the first voice segment as other easily confused words such as "finger bone" is reduced. Of course, the above example is only one possible example of transcribing the first voice segment based on the recognition content, and other possible ways are not limited herein and will not be exemplified one by one. In addition, the speaker label to which the second subtext belongs in the continued recognition process can reuse the original recognition method (such as through voiceprint recognition); or, the speaker label to which the second subtext belongs can be output simultaneously in the process of recognizing and transcribing the reference recognition content. For example, the prompt instruction can be constructed based on the first voice segment and the recognition content of the inserted image, and the prompt instruction is used to instruct the large language model to recognize the content and the speaker label of the first voice segment with the assistance of the recognition content of the inserted image. The output content of the large language model in response to the prompt instruction can be obtained as the second subtext corresponding to the first voice segment and the speaker label thereof.It should be noted that the prompt instruction can still contain other contents, such as still containing the reference voice of several speakers, so that the large language model discriminates the speaker label to which the first voice segment belongs with the assistance of the reference voice of several speakers.
[0031] In one implementation scenario, as previously described, in the case of real-time operation for speaker modification, the voice segment to which the first subtext to which the speaker modification is directed belongs can be obtained as a first target segment, and the speaker label for which the speaker modification takes effect can be obtained as a target label. On this basis, in response to determining that the second voice segment currently to be transcribed and the first target segment belong to the same speaker in the process of continuing real-time transcription, the speaker label to which the second subtext corresponding to the second voice segment belongs is set as the target label. It should be noted that the voiceprint features of the second voice segment and the voiceprint features of the first target segment can be extracted respectively, and if the similarity (such as cosine similarity) between the voiceprint features is higher than a set threshold, it can be determined that the second voice segment and the first target segment belong to the same speaker. Of course, the above example is only one possible example of determining whether two voice segments belong to the same speaker in actual application, and other possible ways are not limited here, and will not be exemplified one by one. Please continue to refer to Figure 2 In the case of real-time operation of "speaker modification" for the first subtext "Yes, here we also thank the organizers", the voice segment to which the first subtext "Yes, here we also thank the organizers" belongs can be obtained as a first target segment, and the speaker label for which the real-time operation of "speaker modification" takes effect can be obtained as a target label, such as modifying the original speaker label "Li Si" to "Zhao Liu", then the target label is "Zhao Liu". On this basis, the second voice segment "Then, our meeting today officially begins" currently to be transcribed is first determined whether it belongs to the same speaker as the first target segment "Yes, here we also thank the organizers", if it is determined that it belongs to the same speaker, the speaker label "Wang Wu" to which the second voice segment "Then, our meeting today officially begins" belongs can be modified to the target label "Zhao Liu"; otherwise, if it is determined that it does not belong to the same speaker, the speaker label to which the second voice segment belongs can be maintained unchanged, that is, in the above example, the speaker label "Wang Wu" to which the second voice segment "Then, our meeting today officially begins" belongs can be maintained unchanged. Of course, the above example is only one possible example in actual application, and other possible cases are not limited here, and will not be exemplified one by one. The above method, in response to determining that the second voice segment currently to be transcribed and the first target segment belong to the same speaker in the process of continuing real-time transcription, sets the speaker label to which the second subtext corresponding to the second voice segment belongs as the target label, which can as much as possible improve the accuracy of the speaker label.
[0032] In one implementation scenario, as previously described, in the case that the real-time operation is transcription modification, a reference word pair formed by the original word of the real-time transcription and the target word on which the transcription modification takes effect can be obtained as reference data. On this basis, in response to detecting that the output word of the real-time transcription matches the original word in the reference word pair in the process of continuing the real-time transcription, the output word can be modified to the target word in the reference word pair. Please continue to refer to Figure 2 , if the speaker "Li Si" has the speech content "Yes, here we also thank the organizers" followed by the name of the organizers, such as "ABCD Company", but the recognition is wrong, and it is actually "EFGH Company" with similar or identical pronunciation. The original word "ABCD Company" of the real-time operation "transcription modification" for the real-time transcription, and the target word "EFGH Company" on which the real-time operation "transcription modification" takes effect, form a reference word pair. It should be noted that the reference data can include at least one reference word pair. In addition, in the above example, "ABCD Company" and "EFGH Company" are only exemplary company names given for ease of description, and are not limited. Based on this, if it is detected that the output word of the real-time transcription contains the original word "ABCD Company" in the aforementioned reference word pair during the process of continuing the transcription, it can be modified to the target word "EFGH Company". Of course, the above example is only one possible example of continuing real-time transcription when the real-time operation is transcription modification in actual application, and other possible cases are not limited herein and will not be exemplified one by one. The above method, in response to detecting that the output word of the real-time transcription matches the original word in the reference word pair in the process of continuing the real-time transcription, modifies the output word to the target word in the reference word pair, which can automatically modify the subsequent transcription according to the original word on which the real-time operation "transcription modification" takes effect and the target word on which the real-time operation "transcription modification" takes effect.
[0033] In one implementation scenario, as previously described, in the case that the real-time operation is note recording, the recording content and recording time of the note recording can be obtained as reference data. On this basis, in response to the third voice segment to be transcribed being adjacent to the recording time in the process of continuing the real-time transcription, the third voice segment can be transcribed based on the recording content to obtain the second subtext corresponding to the third voice segment. As one possible implementation, a prompt instruction can be constructed based on the recording content and the third voice segment, and the prompt instruction is used to instruct the large language model to refer to the recording content to transcribe the third voice. The output content of the large language model in response to the prompt instruction is taken as the second subtext corresponding to the third voice segment. Please continue to refer to Figure 2If the recording content "Today's topics: 1, thinking chain, 2,..." is detected when the speaker "Wang Wu" speaks "So, our meeting today officially begins", then if the speaker "Wang Wu" continues to speak "The first topic of today is thinking chain" in the process of continuing real-time transcription, the above continued speech content can be transcribed based on the recording content "Today's topics: 1, thinking chain, 2,..." to obtain the corresponding second subtext "The first topic of today is thinking chain", so as to avoid misidentification of key elements in the speech content (such as "thinking chain" misidentified as "thinking love"). Of course, the above example is only one possible example of transcribing the third speech segment based on the recording content in actual application, and other possible cases are not limited here, and are not exemplified one by one. In addition, the speaker label to which the second subtext belongs in the process of continuing recognition can reuse the original recognition method (such as through voiceprint recognition, etc.); or, the speaker label to which the second subtext belongs can also be output in the process of recognizing and transcribing the reference recognition content; for example, a prompt instruction can be constructed based on the third speech segment and the recording content, and the prompt instruction is used to instruct the large language model to recognize the content and the speaker label of the third speech segment under the assistance of the recording content, and then the output content of the large language model in response to the prompt instruction can be obtained as the second subtext corresponding to the third speech segment and the speaker label thereof. It should be noted that the prompt instruction can also continue to include other content, such as it can also continue to include the reference speech of several speakers, so that the large language model discriminates the speaker label to which the third speech segment belongs under the assistance of the reference speech of several speakers. The above method, in response to the third speech segment adjacent to the recording time in the process of continuing real-time transcription, transcribes the third speech segment based on the recording content to obtain the second subtext corresponding to the third speech segment, which can refer to the recording content in the process of continuing recognition and transcription, and helps to improve the accuracy of continuing recognition and transcription.
[0034] It should be noted that the above examples are only several possible examples of continuing transcription recognition when the real-time operations are image insertion, speaker modification, transcription modification, and note recording in actual application, and other possible cases are not limited here, and are not exemplified one by one.
[0035] In one implementation scenario, in response to the summarization operation, real-time operations belonging to the target type are screened from the voice stream from starting recording to triggering the summarization operation as target operations. On this basis, the latest transcribed content can be summarized based on the validity data of the target operations to obtain the summary content of the voice stream from starting recording to triggering the summarization operation. It should be noted that the latest transcribed content is the transcribed content of the voice stream from starting recording to triggering the summarization operation. The above manner combines the validity data of each real-time operation belonging to the target type to summarize the latest transcribed content to obtain the summary content of the voice stream from starting recording to triggering the summarization operation, which can improve the accuracy and integrity of voice summarization.
[0036] In one specific implementation scenario, please refer to Figure 2 In the recording interface, a related control for triggering the summarization operation can be provided. If a triggering action on the related control is detected in the recording interface, it can be determined that the summarization operation is triggered. In addition, the triggering time of the summarization operation can be stopping inputting the voice stream, ending recording, etc. The triggering time of the summarization operation is not limited herein.
[0037] In one specific implementation scenario, the target type includes at least one of image insertion and note recording. For example, the target type can include image insertion; or the target type can include note recording; or the target type can include image insertion and note recording, which is not limited herein. In addition, as described previously, in the case where the target type includes image insertion, the validity data can include the recognized content of the inserted image and the insertion position, and in the case where the target type includes note recording, the validity data can include the recorded content of the note and the recording time. For details, please refer to the foregoing related description, which is not repeated herein.
[0038] In one specific implementation scenario, in the case where the target type includes image insertion, when summarizing the latest transcribed content, the recognized content can be inserted in the latest transcribed content, and the insertion position of the recognized content is the insertion position of the inserted image. For example, the insertion position of the inserted image is between the speech content of the speaker "Li Si" "Yes, here we also thank the organizers" and the speech content of the speaker "Wang Wu" "So, our meeting today officially begins", and the recognized content of the inserted image can also be located between "Yes, here we also thank the organizers" and "So, our meeting today officially begins" in the latest transcribed content. Of course, the above example is only one possible example when the target type includes image insertion, and other possible cases are not limited here, and will not be exemplified one by one. On this basis, the latest transcribed content with the inserted recognized content can be summarized, such as using a pre-trained language model such as BERT (Bidirectional Encoder Representations from Transformers, Transformer-based encoder representation) or a large language model such as Llama to summarize, which is not limited here, and thus the summary content of the voice stream from the start of recording to the triggering of the summarization operation can be obtained.
[0039] In one specific implementation scenario, in the case where the target type includes note recording, when summarizing the latest transcribed content, the recorded content can be inserted in the latest transcribed content, and the insertion position of the recorded content is the recording time corresponding to the recorded content. For example, when recording between the speech content of the speaker "Li Si" "Yes, here we also thank the organizers" and the speech content of the speaker "Wang Wu" "So, our meeting today officially begins", the recorded content can be inserted between "Yes, here we also thank the organizers" and "So, our meeting today officially begins" in the latest transcribed content, and then summarization can be performed (for specific implementation, please refer to the foregoing description), and thus the summary content of the voice stream from the start of recording to the triggering of the summarization operation can be obtained.
[0040] In one implementation scenario, in response to detecting the overall transcription triggered by the voice stream, the voice data of the voice stream from the start of recording to the triggering of the overall transcription can be obtained as the to-be-identified voice, and the overall transcription is performed on the to-be-identified voice based on the reference data of the overall transcription, to obtain third identification text. It should be noted that the third identification text includes a plurality of third subtexts and speaker labels to which each third subtext belongs, the reference data of the overall transcription includes the effective data of each real-time operation triggered in the process of the voice stream from the start of recording to the triggering of the overall transcription, and the transcription content in the recording interface is updated based on the third identification text. In addition, the trigger timing of the "overall transcription" in the embodiments of the present disclosure can be the stop of voice stream input, the end of recording, etc., and the trigger timing of the "overall transcription" is not limited herein. The specific implementation of the "overall transcription" can be referred to the following disclosed embodiments, and will not be described herein. The above manner, in response to detecting the overall transcription triggered by the voice stream, obtains the voice data of the voice stream from the start of recording to the triggering of the overall transcription as the to-be-identified voice, performs the overall transcription on the to-be-identified voice based on the reference data of the overall transcription, obtains the third identification text, and can perform the overall transcription on the voice data from the start of recording to the triggering of the overall transcription based on the reference data of the overall transcription to obtain the third identification text, which is helpful to further implement more accurate overall transcription on the basis of real-time transcription.
[0041] The above scheme performs real-time transcription based on the voice stream to obtain the first identification text, and the first identification text includes a plurality of first subtexts and speaker labels to which each first subtext belongs. Then, based on the first identification text, the transcription content in the recording interface is updated in real time, so that in response to the real-time operation triggered on the recording interface, the effective data of the real-time operation is obtained as the reference data of real-time transcription, and then the voice stream is continuously transcribed in real time based on the reference data to obtain the second identification text, and the second identification text includes a plurality of second subtexts and speaker labels to which each second subtext belongs. The transcription content in the recording interface is updated based on the second identification text. Therefore, on the one hand, the real-time operation triggered on the recording interface can be accepted during real-time transcription, which is helpful to meet individual needs. On the other hand, after accepting the real-time operation, the real-time transcription of the subsequent voice stream is further fed back according to the effective data of the real-time operation. Compared with conventional recognition transcription technology, the accuracy of transcription recording can be improved as much as possible. Therefore, the accuracy of transcription recording can be improved as much as possible on the premise of meeting individual needs.
[0042] Please refer to Figure 3 , Figure 3 is a flowchart of another embodiment of the voice recording method of the present application. Specifically, it can include the following steps: Step S31: performing real-time transcription based on the voice stream to obtain first identification text.
[0043] In the embodiments of the present disclosure, the first recognized text includes a plurality of first subtexts and a speaker label to which each first subtext belongs. For details, please refer to the foregoing disclosed embodiments, which will not be repeated here.
[0044] Step S32: Based on the first recognized text, the transcription content in the recording interface is updated in real time.
[0045] For details, please refer to the foregoing disclosed embodiments, which will not be repeated here.
[0046] Step S32: Obtain the effective data triggered by the real-time operation of the recording interface.
[0047] Specifically, the real-time operation can include but is not limited to image insertion, speaker modification, transcription modification, note recording, etc. For specific meanings and corresponding effective data, please refer to the foregoing disclosed embodiments, which will not be repeated here.
[0048] Step S34: In response to detecting the trigger of the overall transcription on the voice stream, the voice data of the voice stream from the start of recording to the trigger of the overall transcription is obtained as the voice to be recognized.
[0049] In the embodiments of the present disclosure, the trigger time of the overall transcription can be any time of the voice stream. That is, the trigger time of the overall transcription can be an intermediate time of the voice stream (i.e., the voice stream is neither stopped inputting nor ended recording), or the voice stream is stopped inputting or ended recording, etc. Here, the trigger time of the overall transcription is not limited. In addition, the overall transcription can be implemented by a deep learning speech recognition model, or can also be implemented by a large language model, and the implementation of the overall transcription is not limited here.
[0050] In one implementation scenario, the recording interface can be provided with a related control for triggering the overall transcription. When the trigger action of the above-mentioned related control in the recording interface is detected, the overall transcription is determined to be triggered.
[0051] In one implementation scenario, when the overall transcription on the voice stream is detected, the voice data of the voice stream from the start of recording to the trigger of the overall transcription can be obtained as the voice to be recognized. Please refer to Figure 2, the voice stream starts recording is initiated from the speech content of the speaker "Zhang San" "Ladies and gentlemen, today we are gathered together", if the output of the speech content of the speaker "Wang Wu" "So, our meeting today officially begins" triggers the overall forwarding, then in this example, the voice to be identified includes the voice data from the start of "Ladies and gentlemen, today we are gathered together" to the output of "So, our meeting today officially begins" triggers the overall transcription. Of course, the above example is only one possible example of the voice to be identified in the actual application process, and other possible cases are not limited, and will not be exemplified one by one.
[0052] Step S35: Overall transcription of the voice to be identified based on the reference data of the overall transcription, to obtain a third recognition text.
[0053] In the embodiments of the present disclosure, the third recognition text can include several third subtexts and the speaker labels to which each third subtext belongs, the reference data of the overall transcription can include the effective data of each real-time operation triggered in the process of the voice stream from starting recording to triggering the overall transcription, and the transcription content in the recording interface can be updated based on the third recognition text. It should be noted that the specific meaning of the third recognition text can be referred to the related description of the second recognition text in the foregoing disclosed embodiments, and the specific meaning of updating the transcription content in the recording interface based on the third recognition text can be referred to the related description of updating the transcription content in the recording interface based on the second recognition text in the foregoing disclosed embodiments, which will not be repeated here.
[0054] In one implementation scenario, as described above, in the case that the respective real-time operations triggered in the process of the voice stream from starting recording to triggering the overall transcription include image insertion, the reference data of the overall transcription can include the recognition content and the insertion position of the inserted image. It should be noted that the specific meaning and acquisition method of the recognition content and the insertion position can be referred to the related description in the foregoing disclosed embodiments, which will not be repeated here. On this basis, the voice segment to which the subtext adjacent to the insertion position in the recording interface belongs can be selected as the second target segment during the overall transcription, and then in response to the fourth voice segment to be transcribed currently in the process of the overall transcription belonging to the second target segment, the fourth voice segment can be transcribed based on the recognition content to obtain the third subtext corresponding to the fourth voice segment. The above-mentioned manner can assist to improve the accuracy of the overall transcription by referring to the recognition content and the insertion position of the inserted image in the process of the overall transcription.
[0055] In a specific implementation scenario, the subtext adjacent to the insertion position can include the first subtext, the second subtext in the foregoing disclosed embodiments, or both the first subtext and the second subtext, which is not limited here.
[0056] In a specific implementation scenario, the subtext adjacent to the insertion position can include the subtext before the insertion position, the subtext after the insertion position, or the subtext before and after the insertion position, which is not limited herein.
[0057] In a specific implementation scenario, when the number of subtexts between the subtext and the insertion position in the recording interface does not exceed the number threshold (for example, 2, 3, 4, etc.), it can be determined that the subtext in the recording interface is adjacent to the insertion position, and the specific way of determining whether the subtext in the recording interface is adjacent to the insertion position is not limited herein.
[0058] In a specific implementation scenario, as a possible implementation, the phonetic segments to which the subtexts before the insertion position and between the insertion position and the number threshold can be selected as the second target segment, and the phonetic segments to which the subtexts after the insertion position and between the insertion position and the number threshold can also be selected as the second target segment. Of course, the above example is only one possible example of selecting the first target segment in actual application, and the specific content of the second target segment is not limited herein.
[0059] In a specific implementation scenario, after determining the second target segment, if it is detected that the fourth phonetic segment to be transcribed currently belongs to the second target segment, the fourth phonetic segment can be transcribed based on the recognition content of the inserted image, that is, the third subtext corresponding to the fourth phonetic segment can be obtained. For example, a prompt instruction can be constructed based on the recognition content and the fourth phonetic segment, and the prompt instruction is used to instruct the large language model to transcribe the fourth phonetic segment with reference to the recognition content, and then the output content of the large language model in response to the prompt instruction can be obtained as the third subtext corresponding to the fourth phonetic segment. It should be noted that the specific way of transcribing the fourth phonetic segment based on the recognition content can refer to the related description of transcribing the first phonetic segment based on the recognition content in the foregoing disclosed embodiments, which will not be repeated here.
[0060] In one implementation scenario, as mentioned above, in the case that the real-time operation triggered in the process from starting recording to triggering the overall transcription respectively contains speaker modification, the reference data of the overall transcription can contain the speech segment to which the subtext targeted by the speaker modification belongs and the speaker label under which the speaker modification takes effect. On this basis, the speech segment to which the subtext targeted by the speaker modification belongs can be selected as the third target segment, and the speaker label under which the speaker modification takes effect can be selected as the target label. It should be noted that the specific meaning and acquisition method of the third target segment can refer to the related description of the first target segment in the foregoing disclosed embodiments, and the specific meaning and acquisition method of the target label can refer to the related description of the target label in the foregoing disclosed embodiments, which will not be described here. After determining the third target segment and the target label, in response to determining that the fifth speech segment to be transcribed currently belongs to the same speaker as the third target segment in the process of the overall transcription, the speaker label corresponding to the third subtext to which the fifth speech segment belongs can be set as the target label. It should be noted that the voiceprint features of the fifth speech segment and the voiceprint features of the third target segment can be extracted, and then it can be determined whether the fifth speech segment and the third target segment belong to the same speaker based on the similarity (such as cosine similarity) between the voiceprint features of the fifth speech segment and the third target segment. For example, when the similarity is higher than a similarity threshold, it can be determined that the fifth speech segment and the third target segment belong to the same speaker, otherwise it can be determined that the fifth speech segment and the third target segment do not belong to the same speaker. In addition, each speech segment to be transcribed in the speech stream can be sequentially taken as the fifth speech segment in the process of the overall transcription. The above-mentioned manner can automatically correct the speaker label of the speech segment belonging to the same speaker to the target label in the process of the overall transcription in the case that the real-time operation contains speaker modification.
[0061] In one implementation scenario, as mentioned above, in the case that the real-time operation triggered in the process from the start of the voice stream to the triggering of the overall transcription includes transcription modification, the reference data of the overall transcription can include the original sentence of the real-time transcription targeted by the transcription modification and the target sentence on which the transcription modification takes effect. On this basis, the original sentence of the real-time transcription targeted by the transcription modification and the target sentence on which the transcription modification takes effect can be selected as a reference sentence pair. It should be noted that the specific meanings and obtaining manners of the original sentence targeted by the transcription modification, the target sentence on which the transcription modification takes effect, and the reference sentence pair can be referred to the related descriptions in the foregoing disclosed embodiments, which will not be described here again. In addition, the reference data of the overall transcription can specifically include at least one reference sentence pair. After obtaining the reference sentence pair, in response to detecting that the output sentence of the overall transcription matches the original sentence in the reference sentence pair in the process of the overall transcription, the output sentence is modified to the target sentence in the reference sentence pair. Illustratively, when the output sentence of the overall transcription is the same as the original sentence in the reference sentence pair, it can be determined that the output sentence of the overall transcription matches the original sentence in the reference sentence pair; or when the output sentence of the overall transcription is different from the original sentence in the reference sentence pair but the pronunciation is similar, it can be determined that the output sentence of the overall transcription matches the original sentence in the reference sentence pair. Of course, the above examples are only several possible examples of determining whether the output sentence of the overall transcription matches the original sentence in the reference sentence pair, and the specific manner of determining whether the output sentence matches the original sentence is not limited here and will not be exemplified one by one. The above manner can automatically correct the output sentence matching the original sentence to the target sentence in the case that the real-time operation includes transcription modification in the process of the overall transcription.
[0062] In one implementation scenario, in the case that the respective real-time operation triggered in the process from the start of the voice stream to the triggering of the overall transcription includes note recording as described above, the reference data of the overall transcription can include the recording content and the recording time of the note recording. On this basis, in response to the sixth voice segment currently to be transcribed being adjacent to the recording time in the process of the overall transcription, the sixth voice segment can be transcribed based on the recording content to obtain a third subtext corresponding to the sixth voice segment. It should be noted that in the case that the time interval between the sixth voice segment currently to be transcribed and the recording time does not exceed the time threshold, it can be determined that the sixth voice segment is adjacent to the recording time. In addition, the sixth voice segment can be located before the recording time, and the sixth voice segment can also be located after the recording time, that is, for the voice segment before or after the recording time, it can be detected whether it is adjacent to the recording time. Of course, in some implementation examples, it can also be detected whether the sixth voice segment located before the recording time is adjacent to the recording time, or in other implementation examples, it can also be detected whether the sixth voice segment located after the recording time is adjacent to the recording time. Here, whether the sixth voice segment adjacent to the recording time is located before or after the recording time is not limited. In addition, the implementation manner of transcribing the sixth voice segment based on the recording content can refer to the related description of transcribing the third voice segment based on the recording content in the foregoing disclosed embodiments, and will not be repeated here. The above-mentioned manner can automatically transcribe the voice segment adjacent to the recording time in the voice stream according to the recording content in the case that the real-time operation includes note recording in the process of the overall transcription.
[0063] It should be noted that the above examples are only several possible examples of overall recognition in the case that the real-time operation is image insertion, speaker modification, transcription modification, and note recording in the actual application process, and other possible cases are not limited here. For example, the process from the start of the voice stream to the triggering of the overall transcription can also include two or more real-time operations, such as image insertion and speaker modification, or speaker modification and transcription modification, or transcription modification and note recording, or speaker modification, transcription modification, and note recording, and here will not be exemplified one by one. In addition, if the process from the start of the voice stream to the triggering of the overall transcription includes two or more real-time operations, the specific implementation of the overall transcription can refer to the foregoing description, and will not be repeated here.
[0064] The above scheme is based on real-time transcription of the voice stream to obtain the first recognized text, and the first recognized text includes a plurality of first subtexts and a speaker label to which each first subtext belongs. Based on the first recognized text, the transcription content in the recording interface is updated in real time, and the effective data of the real-time operation triggered in the recording interface is obtained, so as to respond to the detection of the overall transcription triggered by the voice stream. The voice data from the start of recording to the triggering of the overall transcription is obtained as the recognized voice, and then the overall transcription of the recognized voice is performed based on the reference data of the overall transcription to obtain the third recognized text. The third recognized text includes a plurality of third subtexts and a speaker label to which each third subtext belongs. The reference data of the overall transcription includes the effective data of each real-time operation triggered in the process of the voice stream from the start of recording to the triggering of the overall transcription. The transcription content in the recording interface is updated based on the third recognized text. On the one hand, the real-time operation triggered in the recording interface can be accepted during the overall transcription, which helps to meet the individual needs. On the other hand, after accepting the real-time operation, the overall transcription of the voice stream is further fed back according to the effective data of the real-time operation. Compared with the conventional recognition and transcription technology, the accuracy of the transcription record can be improved as much as possible. Therefore, the accuracy of the transcription record can be improved as much as possible on the premise of meeting the individual needs.
[0065] Please refer to Figure 4 , Figure 4 is a frame schematic diagram of an embodiment of the voice recording device. The voice recording device 40 includes a first transcription module 41, an interface updating module 42, a reference obtaining module 43, and a second transcription module 44. The first transcription module 41 is configured to perform real-time transcription based on the voice stream to obtain the first recognized text. The first recognized text includes a plurality of first subtexts and a speaker label to which each first subtext belongs. The interface updating module 42 is configured to update the transcription content in the recording interface in real time based on the first recognized text. The reference obtaining module 43 is configured to obtain the effective data of the real-time operation as the reference data of the real-time transcription in response to the real-time operation triggered in the recording interface. The second transcription module 44 is configured to continue real-time transcription of the voice stream based on the reference data to obtain the second recognized text. The second recognized text includes a plurality of second subtexts and a speaker label to which each second subtext belongs. The transcription content in the recording interface is updated based on the second recognized text.
[0066] The above scheme, the voice recording device 40 performs real-time transcription based on the voice stream to obtain the first recognized text, and the first recognized text includes a plurality of first subtexts and a speaker label to which each first subtext belongs. Based on the first recognized text, the transcription content in the recording interface is updated in real time, so as to obtain the effective data of the real-time operation triggered in the recording interface as the reference data of real-time transcription, and then continue to perform real-time transcription on the voice stream based on the reference data to obtain the second recognized text. The second recognized text includes a plurality of second subtexts and a speaker label to which each second subtext belongs. The transcription content in the recording interface is updated based on the second recognized text. On the one hand, the real-time operation triggered in the recording interface can be accepted during real-time transcription, which helps to meet individual needs. On the other hand, after accepting the real-time operation, the real-time transcription of the subsequent voice stream is further fed back according to the effective data of the real-time operation. Compared with conventional recognition and transcription technology, the accuracy of transcription recording can be improved as much as possible. Therefore, the accuracy of transcription recording can be improved as much as possible on the premise of meeting individual needs.
[0067] In some disclosed embodiments, when the real-time operation is image insertion, the reference acquisition module 43 is specifically configured to acquire the recognized content of the inserted image and the insertion position as the reference data; and the second transcription module 44 is specifically configured to, in response to the current first voice segment to be transcribed being adjacent to a reference voice segment during continuous real-time transcription, transcribe the first voice segment based on the recognized content to obtain a second subtext corresponding to the first voice segment; wherein the reference voice segment is a voice segment to which the first subtext located at the insertion position belongs.
[0068] In some disclosed embodiments, when the real-time operation is speaker modification, the reference acquisition module 43 is specifically configured to acquire a voice segment to which the first subtext to which the speaker modification is directed belongs as a first target segment, and acquire a speaker label to which the speaker modification is effective as a target label; and the second transcription module 44 is specifically configured to, in response to determining that the current second voice segment to be transcribed and the first target segment belong to the same speaker during continuous real-time transcription, set the speaker label to which the second subtext corresponding to the second voice segment belongs as the target label.
[0069] In some disclosed embodiments, when the real-time operation is transcription modification, the reference acquisition module 43 is specifically configured to acquire a reference word pair formed by an original word to which the transcription modification is directed and a target word to which the transcription modification is effective as the reference data; and the second transcription module 44 is specifically configured to, in response to detecting that the output word of the real-time transcription matches the original word in the reference word pair during continuous real-time transcription, modify the output word to the target word in the reference word pair.
[0070] In some disclosed embodiments, in the case of recording a note in real-time operation, the reference obtaining module 43 is specifically configured to obtain the recording content and the recording time of the note as the reference data; and the second transcription module 44 is specifically configured to, in response to the third voice segment currently to be transcribed being adjacent to the recording time in the process of continuing real-time transcription, transcribe the third voice segment based on the recording content to obtain the second subtext corresponding to the third voice segment.
[0071] In some disclosed embodiments, the voice recording device 40 comprises an operation screening module configured to, in response to an abstract operation, screen real-time operations belonging to a target type as target operations in the process of recording the voice stream from the start of recording to triggering the abstract operation; and the voice recording device 40 comprises a content abstracting module configured to abstract the latest transcribed content based on the effective data of the target operations to obtain the summary content of the voice stream from the start of recording to triggering the abstract operation.
[0072] In some disclosed embodiments, the target type comprises at least one of image insertion and note recording; wherein, in the case of the target type comprising image insertion, the effective data comprises the recognized content of the inserted image and the insertion position, and in the case of the target type comprising note recording, the effective data comprises the recording content and the recording time of the note.
[0073] In some disclosed embodiments, the voice recording device 40 comprises a voice obtaining module configured to, in response to detecting a whole transcription triggered for the voice stream, obtain voice data of the voice stream from the start of recording to triggering the whole transcription as voice to be recognized; and the voice recording device 40 comprises a whole transcription module configured to, based on the reference data of the whole transcription, transcribe the voice to be recognized to obtain a third recognized text; wherein the third recognized text comprises a plurality of third subtexts and a speaker label to which each third subtext respectively belongs, the reference data of the whole transcription comprises effective data of each real-time operation triggered in the process of recording the voice stream from the start of recording to triggering the whole transcription, and the transcribed content in the recording interface is updated based on the third recognized text.
[0074] In some disclosed embodiments, in the case of each real-time operation triggered in the process of recording the voice stream from the start of recording to triggering the whole transcription comprising image insertion, the reference data of the whole transcription comprises the recognized content of the inserted image and the insertion position, and the whole transcription module is specifically configured to select a voice segment in the recording interface adjacent to the insertion position as a second target segment; in response to a fourth voice segment currently to be transcribed in the process of the whole transcription belonging to the second target segment, transcribe the fourth voice segment based on the recognized content to obtain a third subtext corresponding to the fourth voice segment.
[0075] In some disclosed embodiments, in a case where each real-time operation triggered in the process of the voice stream from starting recording to triggering the overall transcription contains speaker modification, the reference data of the overall transcription contains a speech segment to which the subtext targeted by the speaker modification belongs and a speaker label under which the speaker modification takes effect, and the overall transcription module is specifically configured to select the speech segment to which the subtext targeted by the speaker modification belongs as a third target segment, and select the speaker label under which the speaker modification takes effect as a target label; in response to determining that a fifth speech segment currently to be transcribed and the third target segment belong to the same speaker in the process of the overall transcription, setting the speaker label corresponding to the third subtext to which the fifth speech segment belongs as the target label.
[0076] In some disclosed embodiments, in a case where each real-time operation triggered in the process of the voice stream from starting recording to triggering the overall transcription contains transcription modification, the reference data of the overall transcription contains an original sentence of the real-time transcription targeted by the transcription modification and a target sentence under which the transcription modification takes effect, and the overall transcription module is specifically configured to select the original sentence of the real-time transcription targeted by the transcription modification and the target sentence under which the transcription modification takes effect as a reference sentence pair; in response to detecting that an output sentence of the overall transcription matches the original sentence in the reference sentence pair in the process of the overall transcription, modifying the output sentence to the target sentence in the reference sentence pair.
[0077] In some disclosed embodiments, in a case where each real-time operation triggered in the process of the voice stream from starting recording to triggering the overall transcription contains note recording, the reference data of the overall transcription contains recording content of the note recording and a recording time, and the overall transcription module is specifically configured to, in response to a sixth speech segment currently to be transcribed being adjacent to the recording time in the process of the overall transcription, transcribe the sixth speech segment based on the recording content to obtain a third subtext corresponding to the sixth speech segment.
[0078] In some disclosed embodiments, the real-time operation contains at least one of image insertion, speaker modification, transcription modification, and note recording; and / or, the time axis of the voice stream and a recording interface are displayed on the same screen, the transcription content in the recording interface is associated with timestamp information of a corresponding speech segment, and in response to a point selection operation triggered on the time axis, the voice stream and the transcription content in the recording interface jump to a corresponding time of the point selection operation; and / or, in a case where the real-time operation contains note recording, the recording content of the note recording and the transcription content are displayed in separate areas in the recording interface.
[0079] Please refer to Figure 5 , Figure 5is a frame schematic diagram of another embodiment of the voice recording device of the present application. The voice recording device 50 comprises: a real-time transcription module 51, an interface updating module 52, a reference obtaining module 53, a voice obtaining module 54, and a whole transcription module 55. The real-time transcription module 51 is configured to perform real-time transcription based on a voice stream to obtain first recognized text. The first recognized text comprises a plurality of first subtexts and a speaker label of each first subtext. The interface updating module 52 is configured to update the transcription content in the recording interface in real time based on the first recognized text. The reference obtaining module 53 is configured to obtain effective data triggered by real-time operations of the recording interface. The voice obtaining module 54 is configured to, in response to detecting whole transcription triggered on the voice stream, obtain voice data of the voice stream from the start of recording to the triggering of the whole transcription as to-be-recognized voice. The whole transcription module 55 is configured to perform whole transcription on the to-be-recognized voice based on reference data of the whole transcription to obtain third recognized text. The third recognized text comprises a plurality of third subtexts and a speaker label of each third subtext. The reference data of the whole transcription comprises effective data of each real-time operation triggered in the process of the voice stream from the start of recording to the triggering of the whole transcription. The transcription content in the recording interface is updated based on the third recognized text.
[0080] In the above scheme, the voice recording device 50 performs real-time transcription based on a voice stream to obtain first recognized text. The first recognized text comprises a plurality of first subtexts and a speaker label of each first subtext. The transcription content in the recording interface is updated in real time based on the first recognized text. Effective data triggered by real-time operations of the recording interface is obtained. In response to detecting whole transcription triggered on the voice stream, voice data of the voice stream from the start of recording to the triggering of the whole transcription is obtained as to-be-recognized voice. The to-be-recognized voice is further transcribed based on reference data of the whole transcription to obtain third recognized text. The third recognized text comprises a plurality of third subtexts and a speaker label of each third subtext. The reference data of the whole transcription comprises effective data of each real-time operation triggered in the process of the voice stream from the start of recording to the triggering of the whole transcription. The transcription content in the recording interface is updated based on the third recognized text. On the one hand, the real-time operations triggered on the recording interface can be accepted during the whole transcription, which helps to meet individual needs. On the other hand, after accepting the real-time operations, the whole transcription of the voice stream is further fed back according to the effective data of the real-time operations. Compared with conventional recognition and transcription technologies, the accuracy of transcription recording can be improved as much as possible. Therefore, the accuracy of transcription recording can be improved as much as possible on the premise of meeting individual needs.
[0081] Please refer to Figure 6 , Figure 6is a schematic diagram of a framework of an embodiment of the electronic device. The electronic device 60 at least includes a memory 61 and a processor 62 coupled with each other. The memory 61 at least stores program instructions. The processor 62 is configured to execute the program instructions to implement the steps in any of the voice recording method embodiments. For details, refer to the foregoing embodiments, which will not be repeated here. As a possible example, the electronic device 60 can include but is not limited to a mobile phone, a tablet computer, a learning machine, a smart large screen, a server, and the like. The specific type of the electronic device 60 is not limited here.
[0082] Specifically, the processor 62 is configured to control itself and the memory 61 to implement the steps in any of the voice recording method embodiments. The processor 62 can also be referred to as a CPU (Central Processing Unit). The processor 62 can be an integrated circuit chip with processing capability. The processor 62 can also be a general purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 62 can be jointly implemented by an integrated circuit chip.
[0083] According to the scheme, the electronic device 60 performs real-time transcription based on the voice stream to obtain the first recognized text, and the first recognized text includes a plurality of first subtexts and a speaker label to which each first subtext belongs. Then, the transcription content in the recording interface is updated in real time based on the first recognized text, so that the effective data of the real-time operation triggered on the recording interface is obtained as reference data for real-time transcription. Then, the second recognized text is obtained by continuing real-time transcription of the voice stream based on the reference data, and the second recognized text includes a plurality of second subtexts and a speaker label to which each second subtext belongs. The transcription content in the recording interface is updated based on the second recognized text. On the one hand, the real-time operation triggered on the recording interface can be accepted during real-time transcription, which helps to meet individual needs. On the other hand, the real-time operation is accepted, and the effective data of the real-time operation is further used to benefit the real-time transcription of the subsequent voice stream. Compared with conventional recognition and transcription technologies, the accuracy of the transcription record can be improved as much as possible. Therefore, the accuracy of the transcription record can be improved as much as possible on the premise of meeting individual needs. In addition, the first recognized text is obtained by performing real-time transcription based on the voice stream, and the first recognized text includes a plurality of first subtexts and a speaker label to which each first subtext belongs. Then, the transcription content in the recording interface is updated in real time based on the first recognized text, and the effective data of the real-time operation triggered on the recording interface is obtained. Then, the third recognized text is obtained by performing overall transcription on the to-be-recognized voice based on the reference data of the overall transcription, and the third recognized text includes a plurality of third subtexts and a speaker label to which each third subtext belongs. The reference data of the overall transcription includes the effective data of each real-time operation triggered during the process of recording the voice stream from the start to the triggering of the overall transcription. The transcription content in the recording interface is updated based on the third recognized text. On the one hand, the real-time operation triggered on the recording interface can be accepted during overall transcription, which helps to meet individual needs. On the other hand, the real-time operation is accepted, and the effective data of the real-time operation is further used to benefit the overall transcription of the voice stream. Compared with conventional recognition and transcription technologies, the accuracy of the transcription record can be improved as much as possible. Therefore, the accuracy of the transcription record can be improved as much as possible on the premise of meeting individual needs.
[0084] Please refer to Figure 7 , Figure 7 is a framework schematic diagram of an embodiment of the computer readable storage medium of the present application. The computer readable storage medium 70 stores program instructions 71 capable of being executed by a processor, and the program instructions 71 are used to implement the steps in any of the voice recording method embodiments described above.
[0085] According to the scheme, the computer readable storage medium 70 performs real-time transcription based on the voice stream to obtain the first recognized text, and the first recognized text includes a plurality of first subtexts and speaker labels respectively belonging to each first subtext. Then, the transcription content in the recording interface is updated in real time based on the first recognized text, so that the effective data of the real-time operation triggered on the recording interface is obtained as reference data for real-time transcription. Then, the voice stream is further transcribed in real time based on the reference data to obtain the second recognized text, and the second recognized text includes a plurality of second subtexts and speaker labels respectively belonging to each second subtext. The transcription content in the recording interface is updated based on the second recognized text. Therefore, on the one hand, the real-time operation triggered on the recording interface can be accepted during real-time transcription, which helps to meet individual needs. On the other hand, the real-time operation is accepted, and the effective data of the real-time operation is further used to benefit the real-time transcription of the subsequent voice stream. Compared with conventional recognition and transcription technologies, the accuracy of the transcription record can be improved as much as possible. Therefore, the accuracy of the transcription record can be improved as much as possible on the premise of meeting individual needs. In addition, the voice stream is transcribed in real time to obtain the first recognized text, and the first recognized text includes a plurality of first subtexts and speaker labels respectively belonging to each first subtext. Then, the transcription content in the recording interface is updated in real time based on the first recognized text, and the effective data of the real-time operation triggered on the recording interface is obtained, so that the voice data from the start of recording to the triggering of the overall transcription of the voice stream is obtained as the recognized voice in response to detecting the overall transcription triggered on the voice stream. Then, the overall transcription is performed on the recognized voice based on the reference data of the overall transcription to obtain the third recognized text, and the third recognized text includes a plurality of third subtexts and speaker labels respectively belonging to each third subtext. The reference data of the overall transcription includes the effective data of each real-time operation triggered in the process from the start of recording to the triggering of the overall transcription of the voice stream. The transcription content in the recording interface is updated based on the third recognized text. Therefore, on the one hand, the real-time operation triggered on the recording interface can be accepted during the overall transcription, which helps to meet individual needs. On the other hand, the real-time operation is accepted, and the effective data of the real-time operation is further used to benefit the overall transcription of the voice stream. Compared with conventional recognition and transcription technologies, the accuracy of the transcription record can be improved as much as possible. Therefore, the accuracy of the transcription record can be improved as much as possible on the premise of meeting individual needs.
[0086] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, details are not repeated here.
[0087] The above description of each embodiment tends to emphasize the differences between each embodiment, and the same or similar parts can be mutually referred to. For brevity, details are not repeated here.
[0088] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other manners. For example, the division of the apparatus embodiments is merely illustrative, and the division of the modules or units can be different, for example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0089] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.
[0090] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0091] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various program codes that can be stored in the medium.
[0092] If the technical solution of the present application involves personal information, the product applying the technical solution of the present application has clearly informed the personal information processing rules before processing the personal information and obtained the personal independent consent. If the technical solution of the present application involves sensitive personal information, the product applying the technical solution of the present application has obtained the personal independent consent before processing the sensitive personal information and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that it has entered the personal information collection range and will collect personal information. If the individual voluntarily enters the collection range, it is considered to agree to collect personal information. Or on the device for processing personal information, through the pop-up information or by asking the individual to upload his / her personal information, the individual's authorization is obtained under the condition of using obvious mark / information to inform the personal information processing rules. The personal information processing rules can include personal information processor, personal information processing purpose, processing method and personal information type, etc.
Claims
1. A voice recording method, characterized in that: include: Performing real-time transcription based on the speech stream to obtain a first recognized text; wherein the first recognized text includes a plurality of first subtexts and speaker labels to which each of the first subtexts belongs; Based on the first recognized text, updating the transcription content in the record interface in real time; In response to a real-time operation triggered on the recording interface, obtaining validation data of the real-time operation as reference data for the real-time transcription; The voice stream is continued to be transcribed in real time based on the reference data to obtain a second recognized text; wherein, the second recognized text includes a plurality of second subtexts and speaker labels to which each second subtext belongs, and the transcription content in the recording interface is updated based on the second recognized text.
2. The method according to claim 1, characterized in that In a case where the real-time operation is image insertion, obtaining validation data of the real-time operation as reference data includes: Acquire the identification content and insertion position of the inserted image as the reference data; The step of continuing to transcribe the voice stream in real time based on the reference data to obtain a second recognized text includes: In response to the first speech segment to be transcribed being adjacent to a reference speech segment during the process of continuing the real-time transcription, the first speech segment is transcribed based on the recognition content to obtain the second sub-text corresponding to the first speech segment; wherein the reference speech segment is the speech segment to which the first sub-text located at the insertion position belongs.
3. The method according to claim 1, characterized in that In the case where the real-time operation is speaker modification, obtaining the effective data of the real-time operation as reference data includes: Obtaining the speech segment to which the first subtext targeted by the speaker modification belongs as the first target segment, and obtaining the speaker label for which the speaker modification takes effect as the target label; The step of continuing to transcribe the voice stream in real time based on the reference data to obtain a second recognized text includes: In response to determining that the second speech segment to be transcribed and the first target segment belong to the same speaker during the process of continuing the real-time transcription, the speaker tag to which the second subtext corresponding to the second speech segment belongs is set as the target tag.
4. The method according to claim 1, wherein In the case where the real-time operation is transcription and modification, obtaining the effective data of the real-time operation as reference data includes: Acquiring a reference phrase pair formed by the transcription modification for the original phrase transcribed in real time and the target phrase on which the transcription modification is effective as the reference data; The step of continuing to transcribe the voice stream in real time based on the reference data to obtain a second recognized text includes: In response to detecting that an output word of the real-time transcription matches the original word in the reference word pair during the continuous real-time transcription, the output word is modified to the target word in the reference word pair.
5. The method according to claim 1, wherein In the case where the real-time operation is a note record, obtaining the effective data of the real-time operation as reference data includes: Obtaining the content and recording time of the note as the reference data; The step of continuing to transcribe the voice stream in real time based on the reference data to obtain a second recognized text includes: In response to the third speech segment to be transcribed being adjacent to the recording time during the process of continuing the real-time transcription, the third speech segment is transcribed based on the recorded content to obtain the second subtext corresponding to the third speech segment.
6. The method according to claim 1, characterized in that The method further comprises: In response to a summary operation, screening, as target operations, real-time operations belonging to a target type from the real-time operations in the voice stream from the start of recording to the triggering of the summary operation; The latest transcribed content is summarized based on the effective data of the target operation to obtain the minutes content of the voice stream from the start of recording to the triggering of the summary operation.
7. The method according to claim 6, characterized in that The target type includes at least one of image insertion and note recording; Among them, when the target type includes the image insertion, the effective data includes the identification content and insertion position of the inserted image; when the target type includes the note record, the effective data includes the record content and record time of the note record.
8. The method according to claim 1, characterized in that The method further comprises: In response to detecting that the overall transcription is triggered for the voice stream, obtaining voice data of the voice stream from the start of recording to the triggering of the overall transcription as the voice to be recognized; The speech to be recognized is transcribed as a whole based on the reference data of the overall transcription to obtain a third recognized text; wherein, the third recognized text includes a plurality of third sub-texts and speaker labels to which each of the third sub-texts belongs, and the reference data of the overall transcription includes the effective data of each of the real-time operations triggered in the process from the start of recording of the speech stream to the triggering of the overall transcription, and the transcription content in the recording interface is updated based on the third recognized text.
9. The method according to claim 8, characterized in that In a case where each of the real-time operations triggered during the process from the start of recording of the voice stream to the triggering of the overall transcription includes image insertion, the reference data for the overall transcription includes the identification content and insertion position of the inserted image, and the overall transcription of the to-be-recognized voice based on the reference data for the overall transcription to obtain a third recognized text includes: selecting the speech segment to which the subtext adjacent to the insertion position in the recording interface belongs as the second target segment; In response to the fourth speech segment to be transcribed belonging to the second target segment during the overall transcription process, the fourth speech segment is transcribed based on the recognition content to obtain the third subtext corresponding to the fourth speech segment.
10. The method according to claim 8, characterized in that In a case where each of the real-time operations triggered during the process from the start of recording of the voice stream to the triggering of the overall transcription includes a speaker modification, the reference data of the overall transcription includes the voice segment to which the subtext targeted by the speaker modification belongs and the speaker tag for which the speaker modification takes effect, and the overall transcription of the to-be-recognized voice based on the reference data of the overall transcription to obtain a third recognized text includes: Selecting the speech segment to which the subtext targeted by the speaker modification belongs as the third target segment, and selecting the speaker label for which the speaker modification takes effect as the target label; In response to determining during the overall transcription process that the fifth speech segment to be transcribed and the third target segment belong to the same speaker, the speaker label to which the fifth speech segment corresponds to the third subtext is set as the target label.
11. The method according to claim 8, characterized in that In a case where each of the real-time operations triggered during the process from the start of recording of the voice stream to the triggering of the overall transcription includes a transcription modification, the reference data of the overall transcription includes the original words and sentences for the real-time transcription for which the transcription modification is applied and the target words and sentences for which the transcription modification is effective, and the overall transcription of the to-be-recognized voice based on the reference data of the overall transcription to obtain a third recognized text includes: selecting an original phrase for the real-time transcription and a target phrase for which the transcription modification is effective as a reference phrase pair; In response to detecting during the overall transcription that the output word of the overall transcription matches the original word in the reference word pair, the output word is modified to the target word in the reference word pair.
12. The method according to claim 8, characterized in that In a case where each of the real-time operations triggered during the process from the start of recording of the voice stream to the triggering of the overall transcription includes a note record, the reference data of the overall transcription includes the recorded content and recording time of the note record, and the overall transcription of the to-be-recognized voice based on the overall transcription reference data to obtain a third recognized text includes: In response to the sixth speech segment to be transcribed being adjacent to the recording time during the overall transcription process, the sixth speech segment is transcribed based on the recorded content to obtain the third subtext corresponding to the sixth speech segment.
13. The method according to any one of claims 1 to 12, characterized in that The real-time operation includes at least one of image insertion, speaker modification, transcription modification, and note recording; And / or, the timeline of the voice stream and the recording interface are displayed on the same screen, the transcribed content in the recording interface is associated with the timestamp information of the corresponding voice segment, and in response to a click operation triggered on the timeline, the voice stream and the transcribed content in the recording interface jump to the time corresponding to the click operation; And / or, in a case where the real-time operation includes note recording, the recording content and the transcribed content of the note recording are displayed in partitions in the recording interface.
14. A voice recording method, characterized in that: include: Performing real-time transcription based on the speech stream to obtain a first recognized text; wherein the first recognized text includes a plurality of first subtexts and speaker labels to which each of the first subtexts belongs; Based on the first recognized text, updating the transcription content in the record interface in real time; Acquire effective data of the real-time operation triggered on the recording interface; In response to detecting that the overall transcription is triggered for the voice stream, obtaining voice data of the voice stream from the start of recording to the triggering of the overall transcription as the voice to be recognized; The speech to be recognized is transcribed as a whole based on the reference data of the overall transcription to obtain a third recognized text; wherein, the third recognized text includes a plurality of third sub-texts and speaker labels to which each of the third sub-texts belongs, and the reference data of the overall transcription includes the effective data of each of the real-time operations triggered in the process from the start of recording of the speech stream to the triggering of the overall transcription, and the transcription content in the recording interface is updated based on the third recognized text.
15. A voice recording device, characterized in that: include: A first transcription module is configured to perform real-time transcription based on the speech stream to obtain a first recognized text; wherein the first recognized text includes a plurality of first subtexts and speaker labels to which each of the first subtexts belongs; An interface updating module, configured to update the transcription content in the recording interface in real time based on the first recognized text; a reference acquisition module, configured to, in response to a real-time operation triggered on the recording interface, acquire effective data of the real-time operation as reference data for the real-time transcription; The second transcription module is used to continue to transcribe the voice stream in real time based on the reference data to obtain a second recognized text; wherein the second recognized text includes a plurality of second subtexts and speaker labels to which each second subtext belongs, and the transcription content in the recording interface is updated based on the second recognized text.
16. A voice recording device, characterized in that: include: A real-time transcription module is configured to perform real-time transcription based on the speech stream to obtain a first recognized text; wherein the first recognized text includes a plurality of first subtexts and speaker labels to which each of the first subtexts belongs; An interface updating module, configured to update the transcription content in the recording interface in real time based on the first recognized text; A reference acquisition module is used to obtain the effective data of the real-time operation triggered on the recording interface; a speech acquisition module, configured to, in response to detecting that the overall transcription is triggered on the speech stream, acquire speech data of the speech stream from the start of recording to the triggering of the overall transcription as speech to be recognized; The overall transcription module is used to perform overall transcription on the speech to be recognized based on the reference data of the overall transcription to obtain a third recognition text; wherein, the third recognition text includes a plurality of third sub-texts and speaker labels to which each of the third sub-texts belongs, and the reference data of the overall transcription includes the effective data of each of the real-time operations triggered in the process from the start of recording of the speech stream to the triggering of the overall transcription, and the transcription content in the recording interface is updated based on the third recognition text.
17. An electronic device, characterized in that: The device comprises at least a memory and a processor, wherein the memory stores at least program instructions, and the processor is configured to execute the program instructions to implement the voice recording method according to any one of claims 1 to 14.
18. A computer-readable storage medium, characterized in that Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the voice recording method according to any one of claims 1 to 14.