Speech transcription method, device, equipment and storage medium
By obtaining the audio storage address in a network-restricted environment and pulling the audio from the storage address for speech transcription, the problem of large audio files being unable to be transcribed under network restrictions is solved, and speech transcription is successfully achieved in a network-restricted environment.
Patent Information
- Application Number
- CN202210468744.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-04-29
AI Technical Summary
In an intranet environment, speech transcription of large audio files cannot be achieved due to network limitations.
In a first environment with network restrictions, voice transcription is performed by obtaining the storage address of the audio to be transcribed and pulling the audio from the storage address, and the audio is stored and transcribed in a second environment without network restrictions.
It realizes the speech transcription of large audio files under network restrictions and improves the success rate of speech transcription in network-restricted environments.
Smart Images

Figure CN114898747B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to the field of speech technology, and more particularly to a speech transcription method, apparatus, device, and storage medium. Background Art
[0002] In an intranet environment, there are certain restrictions on the connection duration and request size between services. In related technologies, when converting large audio files to text in this situation, the audio file is directly uploaded for speech transcription and then the corresponding transcription results are obtained. However, speech transcription of large audio files cannot be achieved in environments with network restrictions. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, and storage medium for speech transcription.
[0004] According to one aspect of the present disclosure, a speech transcription method is provided, including: determining a first duration occupied by uploading audio to be transcribed under a first network environment; when the first duration is greater than a preset duration, obtaining a storage address of the audio to be transcribed from a received transcription request, wherein the storage address is a storage address corresponding to a storage space that receives and stores the audio to be transcribed under a second network environment, and a first data transmission speed under the first network environment is less than a second data transmission speed under the second network environment; pulling the corresponding audio to be transcribed from the storage address; and performing speech transcription on the audio to be transcribed.
[0005] Optionally, pulling the corresponding audio to be transcribed from the storage address includes: calling a transcription interface corresponding to the transcription request, and pulling the audio to be transcribed from the storage address in the transcription interface.
[0006] Optionally, performing voice transcription on the audio to be transcribed includes: obtaining a processing status of the audio to be transcribed; and when the processing status indicates that all transcription is completed, determining that the transcription of the audio to be transcribed is completed.
[0007] Optionally, after pulling the audio to be transcribed from the storage address in the transcription interface, the method further includes: receiving an event identifier returned by the transcription interface, wherein the audio to be transcribed each time the voice transcription is performed corresponds one-to-one to the event identifier.
[0008] Optionally, voice transcription is performed on the audio to be transcribed, including: receiving a polling request for the transcription details interface initiated by the target object through an event identifier; returning the processing status of the audio to be transcribed in the transcription details interface; and when the processing status is processing completed, returning the transcription result to the target object.
[0009] Optionally, when the processing status is one of the following situations, the polling request for the transcription details interface initiated by the target object through the event identifier is refused, including: the processing of the audio to be transcribed fails, the processing of the audio to be transcribed succeeds, and the polling request meets the first preset condition of the polling.
[0010] Optionally, it also includes: when the audio to be transcribed in the storage address is multiple audio clips, obtaining the identifiers of the multiple audio clips respectively; pulling the corresponding audio to be transcribed from the storage address includes: pulling multiple audio clips from the storage address in sequence, and when the identifier of the pulled audio clip is an end identifier, determining that the audio to be transcribed has been pulled.
[0011] Optionally, pulling the corresponding audio to be transcribed from the storage address includes: when the audio to be transcribed in the storage address is a single continuous audio, pulling the audio to be transcribed from the storage address, and determining that the audio to be transcribed has been read if the audio to be transcribed meets any of the following conditions: the amount of data of the audio to be transcribed that has been read reaches a preset data amount; the reading time of the audio to be transcribed is greater than the preset time.
[0012] Optionally, voice transcription is performed on the audio to be transcribed, including: grouping the audio to be transcribed according to a preset method to obtain multiple groups of sub-audios to be transcribed, wherein the preset method is to group the audio to be transcribed according to the unit data amount in streaming reading, and there is a sequence between the multiple groups of sub-audios to be transcribed; reading the multiple groups of sub-audios to be transcribed in sequence according to the sequence; voice transcription is performed on the multiple groups of sub-audios to be transcribed to obtain multiple groups of transcription results corresponding to the multiple groups of sub-audios to be transcribed.
[0013] According to another aspect of the present disclosure, an interactive method for speech transcription is provided, comprising: displaying a human-computer interaction interface, wherein a first area is provided in the human-computer interaction interface, the first area being used to display audio to be transcribed pulled from a storage address, the storage address being the storage address of the audio to be transcribed in the storage space obtained from the received transcription request; in response to a trigger instruction of a target control in the first area, the audio to be transcribed is transcribed using the above-mentioned speech transcription method, and the transcription result is displayed.
[0014] Optionally, a second area is provided in the human-computer interaction interface, and the second area is used to display configuration properties for configuring the transcription process of the audio to be transcribed.
[0015] According to another aspect of the present disclosure, a speech transcription device is provided, including: a determination module, used to determine a first duration occupied by uploading audio to be transcribed under a first network environment; an acquisition module, used to obtain the storage address of the audio to be transcribed from a received transcription request when the first duration is greater than a preset duration, wherein the storage address is a storage address in which the audio to be transcribed is received and stored by a storage space under a second network environment, and a first data transmission speed under the first network environment is less than a second data transmission speed under the second network environment; a pulling module, used to pull the corresponding audio to be transcribed from the storage address; and a transcription module, used to perform speech transcription on the audio to be transcribed.
[0016] According to another aspect of the present disclosure, an interactive device for speech transcription is provided, including: a first display module for displaying a human-computer interaction interface, a first area being provided in the human-computer interaction interface, the first area being used to display audio to be transcribed pulled from a storage address, the storage address being the storage address of the audio to be transcribed in the storage space obtained from the received transcription request; a processing module for responding to a trigger instruction of a target control in the first area, transcribing the audio to be transcribed using the above-mentioned speech transcription method; and a second display module for displaying the transcription result.
[0017] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.
[0018] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned speech transcription method.
[0019] According to yet another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the above method when executed by a processor.
[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0022] Figure 1 is a flow chart of a speech transcription method according to the first embodiment of the present disclosure;
[0023] Figure 2 is a flowchart for determining the completion of transcription of audio to be transcribed according to the second embodiment of the present disclosure;
[0024] Figure 3 is a flowchart of obtaining a processing status of audio to be transcribed according to the second embodiment of the present disclosure;
[0025] Figure 4 This is a flowchart of determining to pull all the audio to be transcribed according to the second embodiment of the present disclosure;
[0026] Figure 5 is a flowchart of performing voice transcription on multiple groups of audios to be transcribed according to the second embodiment of the present disclosure;
[0027] Figure 6 is a structural diagram of a speech transcription device according to the third embodiment of the present disclosure;
[0028] Figure 7a is a schematic diagram of an interactive interface for speech transcription according to the fourth embodiment of the present disclosure;
[0029] Figure 7b is a flowchart of speech transcription according to the fourth embodiment of the present disclosure;
[0030] Figure 7c is a structural diagram of an interactive device for speech transcription according to a fifth embodiment of the present disclosure;
[0031] Figure 8 3 is a block diagram of an electronic device for implementing the speech transcription method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0033] First, some nouns or terms that appear in the description of the embodiments of the present disclosure are subject to the following explanations:
[0034] Uniform Resource Locator (URL): It is a method of indicating the location of information on the World Wide Web service program of the Internet.
[0035] Weak network environments or network-restricted environments include, but are not limited to: communication networks deployed on high-speed vehicles, various Wi-Fi hotspots, communication networks deployed far from urban areas or in special areas, etc.
[0036] In some companies' intranet environments, there are certain restrictions on the connection duration and request size between services. In this situation, when transcribing large audio files to text offline, simply uploading the large audio file for speech-to-text conversion is not feasible due to network limitations. Conventional speech transcription solutions typically require directly uploading the audio file and then obtaining the corresponding transcription results.
[0037] In order to solve the above technical problems, the embodiments of the present disclosure provide corresponding solutions, which are described in detail below.
[0038] Example 1
[0039] The embodiments of the present disclosure provide an embodiment of a method for speech transcription. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] Figure 1 is a flow chart of a speech transcription method according to an embodiment of the present disclosure, such as Figure 1 As shown, the method includes the following steps:
[0041] Step S102, determining a first duration of time taken to upload the audio to be transcribed under a first network environment;
[0042] In some embodiments of the present disclosure, the first network environment may refer to the environment in which the current client is located as a network-constrained environment or a weak network environment, and the client is a client for implementing an audio transcription service. Assuming that the network speed under the first network environment is 1M / second, if the file size of the audio to be transcribed exceeds 60M, it means that it takes at least 1 minute to complete uploading the audio to be transcribed to the voice transcription service side for voice transcription service. The 1 minute here corresponds to the first duration of step S102, and in many environments, an interface may not be allowed to take 1 minute, so audio files above 60M can be considered as large audio files. In this case, it is necessary to obtain the storage address of the audio to be transcribed, pull the corresponding audio to be transcribed from the storage address, and complete the audio-to-text operation.
[0043] It should be noted that the above is only an example of a case where the audio file to be transcribed may be a large audio file. The specific value of the large audio file is not limited. Whether it is a large audio file can be determined according to actual circumstances.
[0044] Step S104: when the first duration is greater than the preset duration, obtaining a storage address of the audio to be transcribed from the received transcription request, wherein the storage address is a storage address corresponding to a storage space that receives and stores the audio to be transcribed under the second network environment, wherein a first data transmission speed under the first network environment is less than a second data transmission speed under the second network environment;
[0045] In step S104, it is assumed that the maximum duration allowed for a request by an interface is 5 seconds. The maximum duration allowed for a request by the interface here is the preset duration. Taking the above example, it is assumed that the network speed under the first network environment is 1M / second, and the file size of the audio to be transcribed is 60M. The corresponding first duration is 1 minute, which is far longer than the preset duration. In this case, it is necessary to obtain the storage address of the audio to be transcribed from the received transcription request, pull the corresponding audio to be transcribed from the storage address, and complete the audio-to-text operation.
[0046] It should be noted that the second network environment in step S104 is an environment without network restrictions, and the network speed in the second network environment is higher than the network speed in the first network environment, that is, the second data transmission speed is higher than the first data transmission speed. Before obtaining the storage address of the audio to be transcribed from the received transcription request, the user first uploads the audio to be transcribed to a storage space in the second network environment and obtains the storage address of the audio to be transcribed in the storage space, and the storage space is accessible.
[0047] Step S106, pulling the corresponding audio to be transcribed from the storage address;
[0048] For example, the corresponding audio data is pulled from the storage space corresponding to the storage address. The storage space may be a storage medium on the network side or a virtual storage space.
[0049] Step S108: performing voice transcription on the audio to be transcribed.
[0050] In an embodiment of the present disclosure, through the above steps, in a first network environment, when the first duration required to upload the audio to be transcribed is greater than the preset duration, it is determined that the audio file to be transcribed is a large audio file in the first network environment. At this time, the voice transcription service side needs to obtain the storage address of the audio to be transcribed from the received transcription request, pull the corresponding audio to be transcribed from the storage address, perform voice-to-text operations on the pulled audio to be transcribed, and save the corresponding transcription results, thereby achieving the purpose of pulling the corresponding audio to be transcribed from the storage address of the audio to be transcribed, so as to solve the technical effect of not being able to convert large audio into text in an environment with network restrictions.
[0051] Example 2
[0052] This embodiment is a further limitation of the speech transcription method in Example 1. Specifically, in step S106 of the above-mentioned speech transcription method, the corresponding audio to be transcribed is pulled from the storage address, which specifically includes the following steps: calling the transcription interface corresponding to the transcription request, passing the storage address as a parameter into the transcription interface, and pulling the audio to be transcribed from the storage address in the transcription interface.
[0053] In the disclosed embodiment, the audio to be transcribed is stored in a storage space, such as a network disk, a NAS storage space built on an intranet, an NFS storage space, etc. The storage space provides a storage address, i.e., a URL, and users can download the corresponding audio to be transcribed through the URL.
[0054] It's important to note that the storage format for the transcribed audio is not limited; it's determined by the storage space. For example, when storing in HDFS, the audio file might be broken into multiple small blocks, each containing multiple data copies, while when storing in NFS, it might simply store a single complete audio file. Regardless of the storage format used, the only requirement is that users can access the transcribed audio through the URL.
[0055] In another optional embodiment, since the user has a request time limit when directly initiating a request to the speech transcription service side, it is impossible to directly upload a large audio file. Therefore, when calling the transcription interface, the user passes the URL to the speech transcription service side through the interface, because the speech transcription service side generally allows a longer request time, and the request from the speech transcription service side is also safer. Then the speech transcription service side can download the corresponding audio file according to this URL.
[0056] It should be noted that in scenarios other than speech transcription services, the above method can be used when it is impossible to directly upload the corresponding service request to the service side in the current network environment, and there is no limitation at this time.
[0057] When the user calls the transcription interface, the transcription interface passes the storage address URL of the audio to be transcribed in the parameter. In this solution, the parameter is an HTTP parameter included in the request body during the HTTP request. In the embodiment of the present disclosure, the request parameter only includes the URL of the audio to be transcribed. It should be noted that the entire request parameter is in JSON format and placed in the body of the HTTP request. The parameter can be in the following format:
[0058] {
[0059] "audioFileUrl": "https: / / baidu.com / file / audio / sample.wav"
[0060] }
[0061] In the disclosed embodiment, the essence of the transcription interface is a network service interface, which can obtain a specific result after inputting specific content (such as the storage address URL of the audio to be transcribed) through an HTTP request, and the result can be whether the request is successful. It should be noted that the transcription interface in the disclosed embodiment is a custom interface, and the main function of the interface is to pull the corresponding audio to be transcribed according to the storage address of the audio to be transcribed.
[0062] In step S108 of the above-mentioned speech transcription method, the audio to be transcribed is speech transcribed, such as Figure 2 The flowchart shown specifically includes the following steps:
[0063] Step S202, obtaining the processing status of the audio to be transcribed;
[0064] Step S204: When the processing status is all completed, determine that the transcription of the audio to be transcribed is completed and save the transcription result.
[0065] The transcription results can be stored in a network disk, NAS storage space built on the intranet, NFS storage space, etc. When the processing status of the audio to be transcribed is fully completed, the user can obtain the transcription results of all the audio to be transcribed.
[0066] It should be noted that before executing the methods of Example 1 and Example 2, the user must first upload and store the audio to be transcribed in an environment without network restrictions and in an accessible state. Specifically, the user uploads the audio to be transcribed to a storage space in his own way, and can access the corresponding audio to be transcribed through the URL corresponding to the storage space. In this way, when the user requests the transcription interface, the URL is passed to the voice transcription service side, and the voice transcription service side then pulls and downloads the audio to be transcribed according to the URL. After obtaining the audio to be transcribed, the voice-to-text transcription work is started.
[0067] In the above-mentioned speech transcription method, after pulling the audio to be transcribed from the storage address in the transcription interface, the method further includes the following steps: receiving an event identifier returned by the transcription interface, wherein each time the audio to be transcribed is performed, the audio to be transcribed corresponds to the event identifier. When a user polls the processing status of the audio to be transcribed, the corresponding speech transcription event can be found by the event identifier, and the corresponding processing status can be accurately obtained to avoid confusion.
[0068] After the transcription interface receives the storage address URL of the audio to be transcribed that the user passed in as a parameter, the transcription interface returns an event identifier, i.e., an event id. In the disclosed embodiment, one audio to be transcribed corresponds to one url, and one event id necessarily corresponds to a transcription event of the audio to be transcribed, but one audio to be transcribed may correspond to multiple event ids. For example, when the same audio to be transcribed is used for speech-to-text conversion, the speech-to-text conversion is performed at two different times yesterday and today, resulting in two different event ids, event 1 and event 2. It should also be noted that an audio to be transcribed may contain multiple audio clips, i.e., each audio clip in the multiple audio clips is a part of the audio to be transcribed, and a complete audio to be transcribed is composed of multiple audio clips. These multiple audio clips are divided in a certain way when they are stored in the storage space. Therefore, when the speech transcription service side pulls the corresponding audio to be transcribed, all audio clips corresponding to the audio to be transcribed are pulled. Alternatively, the multiple audio clips contained in an audio to be transcribed may also be unrelated, such as multiple independent audio clips that are uploaded and transcribed in batches.
[0069] In the above-mentioned speech transcription method, the audio to be transcribed includes at least one of the following parameter information: file format, audio duration, number of channels, sampling rate, and file size. By obtaining the file size in the audio parameter information, it can be determined whether the audio to be transcribed can be directly uploaded to the service side under the current network environment. If the current environment does not meet the requirements for directly uploading the audio to the service side, the user needs to upload the audio to a storage space in an environment without network restrictions and obtain the corresponding URL. The user can obtain the corresponding audio file from the storage space corresponding to the URL by passing the URL as a parameter to the interface.
[0070] Since audio files have their own parameter information, no additional storage is required. For example, the audio format and duration of an MP3 file can be obtained directly from the file header. General parameter information includes:
[0071] File format: such as mp3, wav, wma, etc.;
[0072] Audio duration: total audio time, usually in seconds;
[0073] Number of channels: refers to the number of channels of data contained in the audio;
[0074] Sample rate: The number of audio data points per second.
[0075] In step S108 of the above-mentioned speech transcription method, the audio to be transcribed is speech transcribed, such as Figure 3 The flowchart shown specifically includes the following steps:
[0076] Step S302: receiving a polling request for a transcription details interface initiated by the target object through an event identifier;
[0077] Step S304: Return the processing status of the audio to be transcribed in the transcription details interface;
[0078] Step S306: When the processing status is complete, the transcription result is returned to the target object.
[0079] Since the transcription process is asynchronous, in order to facilitate users to obtain the transcription progress in a timely manner, users can request to poll the transcription details interface through the event id to obtain the transcription progress and corresponding transcription results.
[0080] It should be noted that after making an audio transcription request, the transcribed text result will not be obtained immediately. Instead, the transcription event needs to be continuously polled through the asynchronous event id. Only after confirming that the transcription event processing is completed can the user obtain the final transcription result.
[0081] In the above-mentioned speech transcription method, when the processing status is one of the following situations, the polling request for the transcription details interface initiated by the target object through the event identifier is refused to be received, including: the processing of the audio to be transcribed fails, the processing of the audio to be transcribed is successful, and the polling request meets the first preset condition for polling. The first preset condition can be that the polling request reaches the maximum duration of the polling, or it can be that the polling request reaches the maximum number of polling times.
[0082] In another optional embodiment, when the processing status is one of the following: failed processing of the audio to be transcribed, successful processing of the audio to be transcribed, or the polling request meets the first preset condition for polling, the user no longer polls; when the processing status is not one of the above, the user continues to poll for the transcription event. By determining the current conditions for polling, polling can facilitate the user to obtain transcription progress in a timely manner, and resources can be saved when polling is not performed.
[0083] In the above-mentioned speech transcription method, Figure 4The flowchart shown in FIG. 1 further includes the following steps: when the audio to be transcribed at the storage address is a plurality of audio segments, respectively obtaining identifiers of the plurality of audio segments, wherein the plurality of audio segments are obtained by dividing the audio to be transcribed according to a first preset method; and then pulling the corresponding audio to be transcribed from the storage address, specifically including the following steps:
[0084] Step S402: Pull multiple audio clips from the storage address in sequence;
[0085] Step S404: When the identifier of the pulled audio segment is an end identifier, it is determined that the audio to be transcribed has been pulled completely.
[0086] In the above steps, the first preset mode is determined according to the maximum timeout duration and the first data transmission speed of allowing to upload audio to be transcribed under the first network environment. Specifically, if the maximum timeout duration of allowing to upload audio to be transcribed under the first network environment is 5 seconds, and the first data transmission speed is 1M / s, then it can be determined that the first preset mode for treating audio to be transcribed is divided to be no more than 5s×1M / s=5M at each division, that is, assuming that the file size of audio to be transcribed is 60M, the audio file can be first split into 15 audio clips by 4M size, because these 15 audio clips are obtained by splitting a complete audio file, therefore these 15 audio clips are stored in a url. When an audio to be transcribed is split into multiple audio clips and uploaded to a storage space, the time of uploading can be saved, and the requirement to network speed is reduced simultaneously.
[0087] These 15 audio clips correspond to a mark in the storage address url, such as mark 1, mark 2, ... mark 15. The allocation of the mark can be allocated after intercepting 4M of the 60M audio in the order of playback. Therefore, when pulling the audio to be transcribed from the storage address, it is necessary to obtain the marks of multiple audio clips, and then determine whether the marks of multiple audio clips are end marks. When the mark is the end mark, it can be determined that all the audio to be transcribed has been pulled. In the above example, mark 15 can be considered as the end mark. When the mark of a section of audio is identified as mark 15, it is determined that all the audio files of the audio to be transcribed have been pulled to the end. When the mark of the audio clip is not identified as mark 15, it means that all the audio files of the audio to be transcribed have not been pulled to the end. Therefore, when a number is used as the mark, the mark with the largest value in the mark can be used as the end mark. The mark can also be represented by other means, which is not limited this time. By obtaining the marks of multiple audio clips corresponding to the audio to be transcribed, it is possible to accurately know whether the audio currently pulled is the last audio, and it is also possible to determine whether all audio files have been pulled.
[0088] In the above-mentioned speech transcription method, the corresponding audio to be transcribed is pulled from the storage address, and another situation is also included: that is, when the audio to be transcribed in the storage address is a single continuous audio, the audio to be transcribed is directly pulled from the storage address, and when the audio to be transcribed meets any of the following conditions, it is determined that the reading of the audio to be transcribed is completed: the data volume of the read audio to be transcribed reaches the preset data volume; the reading time of the audio to be transcribed is greater than the preset time.
[0089] Before pulling the audio to be transcribed, it is necessary to first obtain the data amount of the audio to be transcribed, which is the preset data amount. In the process of pulling the audio to be transcribed, the data amount of the audio to be transcribed that has been read is monitored in real time. When the data amount of the audio to be transcribed that has been read reaches the preset data amount, it is determined that the reading of the audio to be transcribed is completed; in another optional embodiment, whether the reading of the audio to be transcribed is completed can be determined by the reading time. Specifically, the preset time required to read the audio to be transcribed is determined according to the network speed in the current environment and the data amount of the audio to be transcribed, and the reading time for reading the audio to be transcribed is obtained in real time. When the reading time is greater than or equal to the preset time, it is determined that the reading of the audio to be transcribed is completed.
[0090] It's important to note that a single continuous audio file can be understood as a complete audio file stored in the storage space without splitting the audio file. This simple storage method eliminates the need to set an identifier for the audio file to be transcribed, simplifying the retrieval process.
[0091] From the above description, it can be seen that in the speech transcription method in the embodiment of the present disclosure, the audio to be transcribed in the storage address url includes two forms: one is a single continuous audio, that is, one url corresponds to a complete audio file, and the audio at this time is not split or other operations; the other is to split a complete audio into multiple audio segments and store them in the url, that is, one audio to be transcribed corresponds to multiple audio segments, and at this time, one url corresponds to multiple audio segments.
[0092] In step S108 of the above-mentioned speech transcription method, the audio to be transcribed is speech transcribed, such as Figure 5 The flowchart shown specifically includes the following steps:
[0093] Step S502: Grouping the audio to be transcribed according to a preset method to obtain multiple groups of sub-audios to be transcribed, wherein the preset method is to group the audio to be transcribed according to a unit data amount during streaming reading, and there is a sequence between the multiple groups of sub-audios to be transcribed;
[0094] Step S504, reading multiple groups of sub-audios to be transcribed in order;
[0095] Step S506 , performing voice transcription on the multiple groups of sub-audios to be transcribed, and obtaining multiple groups of transcription results corresponding to the multiple groups of sub-audios to be transcribed.
[0096] In above-mentioned step S502, preset mode refers to the mode of reading audio file, can be streaming reading audio file, for example, every 128 bytes in audio file to be transcribed is taken as a group, grouped and sent to speech transcription service, and after multiple groups of sub-audio to be transcribed are carried out speech transcription, corresponding speech transcription result is obtained.It should be noted that, when audio to be transcribed is a single continuous audio or multiple audio clips, audio to be transcribed can be grouped according to preset mode, multiple groups of sub-audio to be transcribed are obtained, by above-mentioned steps S504 to step S506, corresponding multiple groups of transcription results are obtained.Through this reading mode, the time of reading can be reduced and the efficiency of reading can be improved.
[0097] In the speech transcription method of the embodiment of the present disclosure, the transcription results are saved according to the second time length during the speech transcription of multiple groups of sub-audios to be transcribed, so that the transcription results can be automatically saved.
[0098] It should be noted that the second duration may be a fixed frequency during the speech transcription process, such as saving the transcription result once every 1 second. The specific frequency may be determined according to actual conditions and is not limited here.
[0099] In step S102 of the above-mentioned speech transcription method, determining the first duration of time taken to upload the audio to be transcribed under the first network environment specifically includes the following steps: obtaining a first data transmission speed and a file size of the audio to be transcribed; dividing the file size of the audio to be transcribed by the first data transmission speed to obtain the first duration of time taken to upload the audio to be transcribed.
[0100] For example, if the first data transmission speed is 1M / s and the file size of the audio to be transcribed is 60M, the first duration is 60M / 1(M / s)=60 seconds. According to the above steps, it is possible to know when to pull the corresponding audio to be transcribed from the storage address, thereby quickly implementing the voice transcription service, without the audio to be transcribed being unable to be uploaded to the voice transcription service side due to a long request time, and the audio to be transcribed being unable to be transcribed in time.
[0101] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0102] Example 3
[0103] Figure 6 is a structural diagram of a speech transcription device according to an embodiment of the present disclosure, such as Figure 6The device includes the following modules:
[0104] A determination module 602 is configured to determine a first duration of time taken to upload the audio to be transcribed under a first network environment;
[0105] an acquisition module 604, configured to acquire, from the received transcription request, a storage address of the audio to be transcribed when the first duration is greater than a preset duration, wherein the storage address is a storage address corresponding to a storage space that receives and stores the audio to be transcribed under the second network environment, and a first data transmission speed under the first network environment is less than a second data transmission speed under the second network environment;
[0106] A pulling module 606 is used to pull the corresponding audio to be transcribed from the storage address;
[0107] The transcription module 608 is used to perform speech transcription on the audio to be transcribed.
[0108] The pulling module 606 in the above-mentioned speech transcription device is also used to implement the following process: calling the transcription interface corresponding to the transcription request, and pulling the audio to be transcribed from the storage address in the transcription interface.
[0109] The transcription module 608 in the speech transcription device is further configured to obtain a processing status of the audio to be transcribed; and when the processing status indicates that all transcription is completed, it is determined that the transcription of the audio to be transcribed is completed.
[0110] The pulling module 606 in the above-mentioned speech transcription device is further used to implement the following process: receiving the event identifier returned by the transcription interface, wherein the audio to be transcribed each time the speech transcription is performed corresponds to the event identifier one by one.
[0111] The transcription module 608 in the speech transcription device is also used to receive a polling request for the transcription details interface initiated by the target object through an event identifier; return the processing status of the audio to be transcribed in the transcription details interface; and return the transcription result to the target object when the processing status is completed.
[0112] In the speech transcription device, when the processing status is one of the following situations, the polling request for the transcription details interface initiated by the target object through the event identifier is rejected, including: the processing of the audio to be transcribed fails, the processing of the audio to be transcribed is successful, and the polling request meets the first preset condition for polling. The first preset condition can be that the polling request reaches the maximum duration of the polling, or it can be that the polling request reaches the maximum number of polling times.
[0113] In the speech transcription device, it is also used to implement the following process: when the audio to be transcribed in the storage address is multiple audio segments, the identifiers of the multiple audio segments are obtained respectively; pulling the corresponding audio to be transcribed from the storage address includes: pulling multiple audio segments from the storage address in sequence, and when the identifier of the pulled audio segment is an end identifier, determining that the audio to be transcribed has been pulled.
[0114] The pulling module 606 in the speech transcription device is also used to implement the following process: when the audio to be transcribed in the storage address is a single continuous audio, the audio to be transcribed is pulled from the storage address, and when the audio to be transcribed meets one of the following conditions, it is determined that the reading of the audio to be transcribed is completed: the data volume of the audio to be transcribed that has been read reaches the preset data volume; the reading time of the audio to be transcribed is greater than the preset time.
[0115] The transcription module 608 in the speech transcription device is also used to implement the following process: grouping the audio to be transcribed according to a preset method to obtain multiple groups of sub-audios to be transcribed, wherein the preset method is to group the audio to be transcribed according to the unit data amount in streaming reading, and there is a sequence between the multiple groups of sub-audios to be transcribed; reading the multiple groups of sub-audios to be transcribed in sequence according to the sequence; performing speech transcription on the multiple groups of sub-audios to be transcribed to obtain multiple groups of transcription results corresponding to the multiple groups of sub-audios to be transcribed.
[0116] Optionally, during the process of voice transcription of multiple groups of sub-audios to be transcribed, the transcription results are saved according to the second time length.
[0117] The determination module 602 in the speech transcription device is used to implement the following process: obtaining a first data transmission speed and a file size of the audio to be transcribed; dividing the file size of the audio to be transcribed by the first data transmission speed to obtain a first time duration occupied by uploading the audio to be transcribed.
[0118] It should be noted that Figure 6 The speech transcription device shown is used to perform Figures 1 to 5 The speech transcription method shown in the figure, therefore the relevant explanations in the above speech transcription method are also applicable to the speech transcription device, and will not be repeated here.
[0119] Example 4
[0120] According to an embodiment of the present disclosure, an interactive method for speech transcription is provided, specifically comprising: displaying a human-computer interaction interface, wherein a first area is provided in the human-computer interaction interface, and the first area is used to display audio to be transcribed pulled from a storage address, and the storage address is the storage address of the audio to be transcribed in the storage space obtained from the received transcription request; in response to a trigger instruction of a target control in the first area, the speech transcription method in the above-mentioned embodiment 1 and embodiment 2 is used to transcribe the audio to be transcribed, and the transcription result is displayed.
[0121] In the above-mentioned interactive method for speech transcription, a second area is further provided in the human-computer interaction interface, and the second area is used to display configuration properties for configuring the transcription process of the audio to be transcribed.
[0122] Corresponding to an interactive method for speech transcription in an embodiment of the present disclosure, Figure 7a A schematic diagram of an interactive interface for speech transcription is provided. The upper part of the interface is the display position of the address bar. The address bar includes at least a transcription interface and a URL address corresponding to the audio to be transcribed. It should be noted that the URL address corresponding to the audio to be transcribed may not be displayed in the address bar in the embodiment of the present disclosure. The URL address can be obtained through the background. The speech transcription service side pulls the corresponding audio to be transcribed from the storage space storing the audio to be transcribed through the URL address corresponding to the audio to be transcribed, and displays it in the first area of the human-computer interaction interface, that is, the right part of the interface. After pulling all the audio to be transcribed, the user clicks the transcription button in the lower right corner to perform speech transcription on the audio to be transcribed. After the speech transcription service side completes the transcription of all the audio to be transcribed, the audio transcription result is displayed in the first area for the user to view the transcription result. It should be noted that the audio transcription result can be directly displayed in the first area and overwrite the original audio to be transcribed; it can also be displayed in a new interface for the convenience of user viewing.
[0123] exist Figure 7a The second area of the schematic diagram shown is used to display the configuration properties for configuring the transcription process of the audio to be transcribed, such as the transcription language, the field to which the audio to be transcribed belongs, etc., which are not limited here. In actual applications, the audio-to-text interface may be expanded according to actual conditions.
[0124] According to an embodiment of the present disclosure, Figure 7a In the interactive interface diagram for speech transcription shown in FIG, the specific speech transcription process diagram is as follows Figure 7b shown.
[0125] like Figure 7bAs shown, in step 701, the user first needs to upload the audio to be transcribed to a storage space in an environment without network restrictions, and it must be in an accessible state. The storage space can be a network disk, a NAS storage space built on an intranet, an NFS storage space, etc. The storage space provides a storage address, i.e., a URL. After the user uploads the audio to be transcribed to the storage space, the user will obtain the URL of the audio to be transcribed in the storage space. When the network environment does not allow direct uploading of large audio files, the audio file can be pulled and downloaded by obtaining the URL address of the audio file in the storage space, thereby performing the speech-to-text operation.
[0126] It's important to note that the storage format for the transcribed audio is not limited; it's determined by the storage space. For example, when storing in HDFS, the audio file might be broken into multiple small blocks, each containing multiple data copies, while when storing in NFS, it might simply store a single complete audio file. Regardless of the storage format used, the only requirement is that users can access the transcribed audio through the URL.
[0127] In addition, since audio files carry relevant parameter information, no additional storage is required. For example, the audio format and duration of MP3 files can be obtained directly from the file header. General parameter information includes: file format: such as MP3, WAV, WMA, etc.; audio duration: the total audio time, generally in seconds; number of channels: refers to how many channels the audio contains; sampling rate: the number of audio data points per second.
[0128] After the user uploads the audio to be transcribed to the storage space, in order to determine which method to use for the audio transcription service, it is necessary to obtain the current network speed and the file size of the audio to be transcribed. According to the current network speed, assuming that the network speed is 1M / second, if the file size of the audio to be transcribed is only 2M, it only takes 2 seconds to upload the audio file to the voice transcription service side. In this case, the corresponding audio to be transcribed does not need to be obtained through the URL of the audio to be transcribed, and directly uploading the audio file to the voice transcription service side does not need to occupy too much time of the interface and will not exceed the request time requested by the user.
[0129] If the file size of the audio to be transcribed exceeds 60MB, it will take at least one minute to upload the audio. However, in many environments, it may not be possible for an interface to take up to one minute. Therefore, audio files over 60MB are considered large. In this case, you need to obtain the storage address of the audio to be transcribed before completing the audio-to-text operation.
[0130] In step 702, the user initiates an audio transcription request to the voice transcription service side through the user-side service interface. The user passes the URL of the audio to be transcribed as a parameter to the transcription interface. This parameter is an HTTP parameter included in the request body during an HTTP request, and in the embodiment of the present disclosure, the request parameter only includes the URL of the audio to be transcribed. It should be noted that the entire request parameter is in JSON format and placed in the body of the HTTP request. The parameter can be in the following format:
[0131] {
[0132] "audioFileUrl": "https: / / baidu.com / file / audio / sample.wav"
[0133] }
[0134] In the disclosed embodiment, the essence of the transcription interface is a network service interface, which can obtain a specific result after inputting specific content (such as the storage address URL of the audio to be transcribed) through an HTTP request, and the result can be whether the request is successful. It should be noted that the transcription interface in the disclosed embodiment is a custom interface, and the main function of the interface is to pull the corresponding audio to be transcribed according to the storage address of the audio to be transcribed.
[0135] Step 703: After the audio transcription service side receives the audio transcription request initiated by the user through the user-side service interface, it pulls the transcribed audio from the storage space according to the URL passed in by the user in the transcription interface. After the operation of pulling the audio to be transcribed is completed, the transcription interface returns an event identifier, i.e., an event ID, where each speech transcription event corresponds to an event identifier.
[0136] It should be noted that one audio to be transcribed corresponds to one url, and one event id must correspond to one audio to be transcribed, but one audio to be transcribed may correspond to multiple event ids. For example, when using the same audio to be transcribed for speech-to-text conversion, the speech-to-text conversion is performed at two different times yesterday and today, and two different event ids, event 1 and event 2, will be obtained. It should also be noted that an audio to be transcribed may contain multiple audio clips, that is, each of the multiple audio clips is a part of the audio to be transcribed, and a complete audio to be transcribed is composed of multiple audio clips. These multiple audio clips have been divided in a certain way when stored in the storage space. Therefore, when the speech transcription service side pulls the corresponding audio to be transcribed, all audio clips corresponding to the audio to be transcribed will be pulled.
[0137] When pulling transcribed audio from the storage space according to the URL passed in by the user in the transcription interface, the audio to be transcribed in the storage space includes two forms: one is a single continuous audio, that is, one URL corresponds to a complete audio file, and the audio at this time is not split or other operations; the other is to split a complete audio into multiple audio segments and store them in the URL, that is, one audio to be transcribed corresponds to multiple audio segments, and at this time, one URL corresponds to multiple audio segments.
[0138] If you need to perform a speech-to-text operation on a 60M audio file, you can first split the audio file into 15 audio segments of 4M in size. Since these 15 audio segments are obtained by splitting a complete audio file, these 15 audio segments are stored in a URL. The split size here is determined based on the maximum timeout duration allowed for uploading the audio to be transcribed to the speech transcription service side under the current network environment and the current network speed. Specifically, if the maximum timeout duration allowed for uploading the audio to be transcribed under the current network environment is 5 seconds, and the current network speed is 1M / s, it can be determined that when splitting the audio to be transcribed, each split should not exceed 5s×1M / s=5M.
[0139] It should be noted that when the audio to be transcribed in the storage address corresponds to multiple audio clips, the conditions for ending the pull need to be determined. Therefore, when the audio to be transcribed is split into multiple audio clips and stored in the storage space, an identifier needs to be set for each split audio clip, and the last identifier serves as the end identifier of the audio to be transcribed.
[0140] Take the above-mentioned splitting of the 60M audio file into 15 audio clips as an example, these 15 audio clips respectively correspond to an identifier in the storage address url, such as identifier 1, identifier 2, ... identifier 15, and the allocation of the identifier can be allocated after intercepting 4M of the 60M audio in the order of playback. Therefore, when pulling the audio to be transcribed from the storage address, it can be determined by the identifiers of multiple audio clips whether the audio to be transcribed has been completely pulled, that is, it is necessary to determine whether the identifiers of multiple audio clips are end identifiers. When the identifiers are end identifiers, it can be determined that all the audio to be transcribed has been pulled. In the above example, identifier 15 can be considered as the end identifier. When the identifier of the audio clip is identified as identifier 15, it is determined that all the audio files of the audio to be transcribed have been pulled. After the operation of pulling the audio to be transcribed is completed, the transcription interface returns an event identifier, i.e., event id.
[0141] In step 704, the speech transcription service side performs speech transcription based on the pulled audio to be transcribed. During this process, the audio transcription service side groups the audio to be transcribed according to a preset method to obtain multiple groups of sub-audios to be transcribed, wherein the preset method is to group the audio to be transcribed according to the unit data volume in streaming reading, and there is a sequence between the multiple groups of sub-audios to be transcribed; the multiple groups of sub-audios to be transcribed are read in sequence according to the sequence; speech transcription is performed on the multiple groups of sub-audios to be transcribed to obtain multiple groups of transcription results corresponding to the multiple groups of sub-audios to be transcribed.
[0142] Specifically, the above-mentioned preset method refers to a method of reading audio files, which can be streaming reading of audio files, such as taking every 128 bytes in the audio file to be transcribed as a group, sending them to the speech transcription service in groups, and performing speech transcription on multiple groups of sub-audios to be transcribed to obtain the corresponding speech transcription results.
[0143] In step 705, the audio transcription service continuously saves the transcription results during the voice transcription process of the multiple audio groups to be transcribed. The transcription results can be stored in a network disk, a NAS storage space built on the intranet, an NFS storage space, etc.
[0144] Step 706: After all the multiple audios to be transcribed are voice-transcribed, the transcription process ends.
[0145] Step 707: After initiating the audio transcription request, the user initiates a polling request from the user-side service interface through the event ID. The polling request is used to poll the processing status of the voice transcription event, including the transcription progress and the corresponding transcription results. It should be noted that the object of the polling request initiated by the user can be the transcription details interface, or other user-defined interfaces that can separately poll the transcription progress and the corresponding transcription results. This is not limited here.
[0146] When the processing status is one of the following situations, refuse to receive the polling request for the transcription details interface initiated by the target object through the event identifier, including: the processing of the audio to be transcribed fails, the processing of the audio to be transcribed is successful, and the polling request meets the first preset condition for polling. The first preset condition can be that the polling request reaches the maximum duration of the polling, or it can be that the polling request reaches the maximum number of polling times.
[0147] In another optional embodiment, when the processing status is one of the following: failure to process the audio to be transcribed, success in processing the audio to be transcribed, or the polling request meets the first preset condition for polling, the user will no longer poll; when the processing status is not one of the above, the user will continue to poll for the transcription event.
[0148] Step 708: When the processing status of the audio to be transcribed is completed, the final transcription result is obtained from the storage space where the transcription result is stored, and the final audio transcription result is displayed to the user.
[0149] It should be noted that after making an audio transcription request, the transcribed text result will not be obtained immediately. Instead, the processing status of the transcription event needs to be continuously polled through the asynchronous event ID. Only after confirming that the transcription event has been successfully processed can the final transcription result be displayed to the user.
[0150] Example 5
[0151] According to an embodiment of the present disclosure, the present disclosure also provides an interactive device for speech transcription, such as Figure 7c As shown in the structural diagram, the device includes:
[0152] A first display module 712 is configured to display a human-computer interaction interface, wherein the human-computer interaction interface includes a first area for displaying the audio to be transcribed pulled from a storage address, where the storage address is the storage address of the audio to be transcribed in the storage space obtained from the received transcription request;
[0153] A processing module 714 is configured to transcribe the audio to be transcribed using the above-mentioned speech transcription method in response to a trigger instruction of a target control in the first area;
[0154] The second display module 716 is used to display the transcription result.
[0155] It should be noted that Figure 7c The interactive device for speech transcription shown is used to execute the above-mentioned interactive method for speech transcription, so the relevant explanations in the above-mentioned interactive method for speech transcription are also applicable to the interactive device for speech transcription, and will not be repeated here.
[0156] Example 6
[0157] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0158] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0159] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0160] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0161] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the speech transcription method. For example, in some embodiments, the speech transcription method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the audio-to-text method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the audio-to-text method in any other suitable manner (e.g., via firmware).
[0162] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0163] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0164] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0165] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0166] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0167] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0168] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0169] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A speech transcription method, comprising: Determining a first duration occupied by uploading the audio to be transcribed under a first network environment, where the first duration is an upload duration; When the first duration is greater than a preset duration, obtaining a storage address of the audio to be transcribed from the received transcription request, wherein the storage address is a storage address corresponding to a storage space that receives and stores the audio to be transcribed under a second network environment, and a first data transmission speed under the first network environment is less than a second data transmission speed under the second network environment; Pull the corresponding audio to be transcribed from the storage address; Perform voice transcription on the audio to be transcribed.
2. The method according to claim 1, wherein Pulling the corresponding audio to be transcribed from the storage address includes: The transcription interface corresponding to the transcription request is called, and the audio to be transcribed is pulled from the storage address in the transcription interface.
3. The method according to claim 2, wherein: Performing voice transcription on the audio to be transcribed includes: Obtain the processing status of the audio to be transcribed; when the processing status indicates that all transcription is completed, determine that the transcription of the audio to be transcribed is completed.
4. The method according to claim 2, wherein: After pulling the audio to be transcribed from the storage address in the transcription interface, the method further includes: Receive the event identifier returned by the transcription interface, wherein the audio to be transcribed each time the voice transcription is performed corresponds to the event identifier one by one.
5. The method according to claim 4, wherein Performing voice transcription on the audio to be transcribed includes: Receiving a polling request for a transcription details interface initiated by a target object through the event identifier; Return the processing status of the audio to be transcribed in the transcription details interface; When the processing status is processing completed, the transcription result is returned to the target object.
6. The method according to claim 5, wherein: When the processing status is one of the following situations, refuse to receive the polling request for the transcription details interface initiated by the target object through the event identifier, including: the audio to be transcribed processing fails, the audio to be transcribed processing succeeds, and the polling request meets the first preset condition for polling.
7. The method according to claim 1, wherein Also includes: When the audio to be transcribed in the storage address is a plurality of audio segments, obtaining identifiers of the plurality of audio segments respectively; Pulling the corresponding audio to be transcribed from the storage address includes: Pull the multiple audio segments from the storage address in sequence, and when the identifier of the pulled audio segment is an end identifier, determine that the audio to be transcribed has been pulled.
8. The method according to claim 1, wherein Pulling the corresponding audio to be transcribed from the storage address includes: In the case that the audio to be transcribed in the storage address is a single continuous audio, the audio to be transcribed is pulled from the storage address, and when the audio to be transcribed meets any of the following conditions, it is determined that the reading of the audio to be transcribed is completed: the amount of data of the read audio to be transcribed reaches a preset data amount; the reading time of the audio to be transcribed is greater than the preset time.
9. The method according to claim 7 or 8, wherein Performing voice transcription on the audio to be transcribed includes: The audio to be transcribed is grouped according to a preset method to obtain multiple groups of sub-audios to be transcribed, wherein the preset method is to group the audio to be transcribed according to the unit data amount in streaming reading, and there is a sequence between the multiple groups of sub-audios to be transcribed; Reading the plurality of sub-audios to be transcribed in sequence according to the sequence; Perform voice transcription on the multiple groups of sub-audios to be transcribed to obtain multiple groups of transcription results corresponding to the multiple groups of sub-audios to be transcribed.
10. An interactive method for speech transcription, comprising: Displaying a human-computer interaction interface, wherein the human-computer interaction interface is provided with a first area, the first area being used to display the audio to be transcribed pulled from a storage address, the storage address being the storage address of the audio to be transcribed in the storage space obtained from the received transcription request; In response to a trigger instruction of a target control in the first area, the audio to be transcribed is transcribed using the method described in any one of claims 1 to 9, and the transcription result is displayed.
11. The method according to claim 10, wherein: A second area is provided in the human-computer interaction interface, and the second area is used to display configuration properties for configuring the transcription process of the audio to be transcribed.
12. A speech transcription device comprising: A determination module, configured to determine a first duration occupied by uploading the audio to be transcribed under a first network environment, wherein the first duration is an upload duration; an acquisition module, configured to acquire, from the received transcription request, a storage address of the audio to be transcribed when the first duration is greater than a preset duration, wherein the storage address is a storage address at which the audio to be transcribed is received and stored by a storage space under a second network environment, and a first data transmission speed under the first network environment is less than a second data transmission speed under the second network environment; A pulling module, configured to pull the corresponding audio to be transcribed from the storage address; The transcription module is used to perform voice transcription on the audio to be transcribed.
13. An interactive device for speech transcription, comprising: A first display module is configured to display a human-computer interaction interface, wherein the human-computer interaction interface is provided with a first area, wherein the first area is configured to display the audio to be transcribed pulled from a storage address, wherein the storage address is a storage address of the audio to be transcribed in a storage space obtained from the received transcription request; a processing module, configured to transcribe the audio to be transcribed using the method according to any one of claims 1 to 9 in response to a trigger instruction of a target control in the first area; The second display module is used to display the transcription results.
14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
16. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Data storage method, data storage device, equipment and storage medium
CN108520763A
Voice processing method and device and device for voice processing
CN111696550A