Audio transcription method and device based on intelligent analysis, equipment and medium
By employing multi-level file verification and intelligent analysis methods, the problem of low efficiency in traditional audio transcription has been solved, achieving efficient and accurate audio transcription and multi-angle analysis, thereby improving the efficiency and accuracy of audio transcription.
Patent Information
- Application Number
- CN202511489117.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-03-03
AI Technical Summary
Existing audio transcription methods based on traditional neural networks have weak generalization ability and poor performance in parsing long-term speech text, resulting in low audio transcription efficiency.
Intelligent analysis methods such as multi-level file verification, temporal attention convolution feature extraction, and stylized keyword decoding are employed to improve the accuracy and efficiency of audio transcription. These methods include file format and integrity verification, audio segmentation, progress polling, text summary feature activation, and stylized keyword decoding.
File verification ensures the processability of audio files, improves transcription efficiency and accuracy, enables global transcription, enhances semantic understanding and multi-angle analysis, and improves the intuitiveness and flexibility of audio transcription.
Smart Images

Figure CN121600931A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio recognition technology, and in particular to an audio transcription method, apparatus, device, and medium based on intelligent analysis. Background Technology
[0002] Audio transcription refers to the process of converting the speech content in audio into text, and then further analyzing and processing the transcribed text to extract key information, identify semantic patterns, or perform structured processing. Audio transcription analysis is widely used in fields such as meeting minutes, customer service dialogue analysis, and medical analysis.
[0003] Existing audio transcription methods are mostly based on traditional neural networks. They use traditional neural network models such as Hidden Markov Models to transcribe and analyze audio text. In practical applications, audio transcription methods based on traditional neural networks rely heavily on decoding rules, have weak generalization ability, and perform poorly on long-term speech texts, which may result in low efficiency during audio transcription. Summary of the Invention
[0004] This invention provides an artificial intelligence-based audio transcription method, apparatus, computer device, and medium based on intelligent analysis. It improves the security of audio analysis by using a multi-level file verification method, enhances the intuitiveness of audio transcription by using a polling progress query method, improves the global and local audio features by using a temporal attention convolution feature extraction method, thereby improving the accuracy of audio transcription, and realizes multi-angle text summarization analysis by using a stylized keyword decoding method, thus solving the technical problem of low efficiency in audio transcription.
[0005] Firstly, an audio transcription method based on intelligent analysis is provided, including: The file to be transcribed is subjected to file format and integrity verification to obtain file verification results. Based on the file verification results and the file size of the file to be transcribed, the file to be transcribed is sliced and uploaded to obtain the uploaded audio file. The uploaded audio file is segmented to obtain a segmented audio sequence. The speech text of the segmented audio sequence is extracted to obtain the primary transcribed text. The analysis transcription progress and real-time transcription progress of the uploaded audio file are calculated. The analysis transcription progress and the real-time transcription progress are filtered for maximum values to obtain the transcription progress. The primary transcribed text is then updated to standard transcribed text based on the transcription progress. Feature activation is performed on the text context features of the standard transcribed text to obtain text summary features; Obtain the style keywords input by the user, and perform feature style decoding on the text summary features based on the style keywords to obtain the decoded text summary; The standard transcribed text and the decoded text summary are combined and spliced to obtain an audio transcribed file.
[0006] Secondly, an audio transcription device based on intelligent analysis is provided, comprising: The file verification module is used to perform file format verification and integrity verification on the file to be transcribed, obtain the file verification result, and upload the file slices of the file to be transcribed based on the file verification result and the file size of the file to be transcribed, thereby obtaining the uploaded audio file. The text recognition module is used to perform audio segmentation on the uploaded audio file to obtain a segmented audio sequence, and extract the speech text of the segmented audio sequence to obtain the primary transcribed text; The progress polling module is used to calculate the analysis transcription progress and real-time transcription progress of the uploaded audio file, perform maximum value filtering on the analysis transcription progress and the real-time transcription progress to obtain the transcription progress, and update the primary transcribed text to standard transcribed text according to the transcription progress. The summary encoding module is used to activate the text context features of the standard transcribed text to obtain text summary features; The style decoding module is used to obtain the style keywords input by the user, and perform feature style decoding on the text summary features based on the style keywords to obtain the decoded text summary; The result generation module is used to combine and splice the standard transcribed text and the decoded text summary to obtain an audio transcribed file.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described audio transcription method based on intelligent analysis.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described audio transcription method based on intelligent analysis.
[0009] In the above-mentioned intelligent analysis-based audio transcription method, apparatus, computer equipment, and storage medium, the file format and integrity of the file to be transcribed are verified to obtain a file verification result. Based on the file verification result and the file size of the file to be transcribed, the file to be transcribed is sliced and uploaded to obtain an uploaded audio file. This ensures that the file to be transcribed is a processable audio file and enables efficient uploading of the file to be transcribed, improving the efficiency of audio transcription analysis. By segmenting the uploaded audio file to obtain segmented audio sequences, and extracting the speech text from the segmented audio sequences, segmented text recognition can be achieved to obtain primary transcribed text. Speech transcription can be performed by combining the local features and global context features of the uploaded audio file, improving the accuracy of audio transcription. By calculating the analyzed transcription progress and real-time transcription progress of the uploaded audio file, and performing maximum filtering on the analyzed transcription progress and real-time transcription progress, the transcription progress is obtained. The latest transcription progress can be obtained through polling and dual analysis methods, thereby improving the intuitiveness of audio transcription. By updating the primary transcribed text to standard transcribed text based on the transcription progress, global transcription of the audio file can be achieved, improving the efficiency of audio transcription analysis.
[0010] By activating the text context features of the standard transcribed text, the semantic understanding of the standard transcribed text can be enhanced by combining bidirectional context features, thereby obtaining concise text summary features and improving the efficiency of audio transcription analysis. By decoding the text summary features according to the style keywords, a decoded text summary is obtained, enabling customized role management and persistent operations, achieving multi-role and multi-angle transcribed text analysis, thus improving the flexibility of transcribed text analysis. By combining and splicing the standard transcribed text and the decoded text summary, an audio transcription file is obtained, clearly displaying the audio transcription results and providing different types of summary text based on different keywords, thereby improving the efficiency of audio transcription analysis. Therefore, the intelligent analysis-based audio transcription method and system proposed in this invention can solve the problem of low efficiency in audio transcription. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating an audio transcription method based on intelligent analysis in one embodiment of the present invention; Figure 2 yes Figure 1 A schematic diagram of a specific implementation method for step S20; Figure 3 yes Figure 1 A schematic diagram of a specific implementation method for step S30; Figure 4 This is a schematic diagram of an audio transcription device based on intelligent analysis in one embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Detailed Implementation The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] Please see Figure 1 As shown, Figure 1 A flowchart illustrating an audio transcription method based on intelligent analysis provided in an embodiment of the present invention includes the following steps: S10: Perform file format and integrity checks on the file to be transcribed to obtain the file verification results. Based on the file verification results and the file size of the file to be transcribed, upload the file slices to obtain the uploaded audio file.
[0014] Specifically, the file to be transcribed refers to an audio file that needs to be transcribed. The file to be transcribed can be an audio recording of a meeting or a call recording of a negotiation. By transing the file to be transcribed into text format and extracting text summaries according to different styles, the efficiency of content organization can be improved.
[0015] Specifically, in order to ensure that the uploaded file for transcription is an analyzable audio file, the file to be transcribed needs to be verified. The transcription of the file to be transcribed needs to be performed through a backend server. Therefore, the verified file to be transcribed needs to be uploaded to the backend server. The uploaded audio file is the file to be transcribed in the backend server.
[0016] In this embodiment of the invention, the process of performing format and integrity checks on the file to be transcribed to obtain the file verification result includes: The file size is obtained by performing a size analysis on the file to be transcribed. The file extensions of the files to be transcribed are extracted to obtain the file extensions. The header features of the file to be transcribed are extracted to obtain the file header features; The format of the file to be transcribed is validated based on the file size, the file extension, and the file header features to obtain the format validation result. The integrity of the file to be transcribed is verified, and the integrity verification result is obtained. A file verification result is generated based on the format verification result and the integrity verification result.
[0017] In detail, the getsize function can be used for size statistics, regular expressions can be used for file extension extraction, and the Linux file command can be used for header feature extraction. The file size refers to the capacity of the file to be transcribed, the file extension is a mechanism used by the operating system to identify the file format, and the file header features refer to the key characters in the header of the file to be transcribed that identify the file type.
[0018] Specifically, when the file extension is mped, the corresponding file header feature can be FF FB or FF F3, and when the file extension is jpeg, the corresponding file header feature can be FF D8 FF E0 / FF D8 FF DB.
[0019] In detail, the step of performing format verification on the file to be transcribed based on the file size, the file extension, and the file header features to obtain the format verification result means determining the file type of the file to be transcribed based on the file extension, determining whether the file size is a valid size for the file type, determining whether the file header features correspond to the file type, and determining whether the file type is an audio type. If all of the above are true, the format verification result is set as successful; otherwise, the format verification result is set as failed.
[0020] Specifically, the integrity check refers to determining whether the file to be transcribed is complete and whether there are any missing files. Integrity check can be performed using hash verification. The integrity check result includes successful check and failed check.
[0021] In detail, generating a file verification result based on the format verification result and the integrity verification result means setting the format verification result to successful when both the format verification result and the integrity verification result are successful, and otherwise setting the format verification result to failed.
[0022] Specifically, the step of uploading slices of the file to be transcribed based on the file verification result and the file size to obtain the uploaded audio file means uploading slices of the file to be transcribed according to the file size when the file verification result is successful.
[0023] In this embodiment of the invention, by performing file verification and uploading on the file to be transcribed, an uploaded audio file is obtained. This ensures that the file to be transcribed is a processable audio file and enables efficient uploading of the file to be transcribed, thereby improving the efficiency of audio transcription analysis.
[0024] S20: Perform audio segmentation on the uploaded audio file to obtain a segmented audio sequence, extract the speech text of the segmented audio sequence, and obtain the primary transcribed text.
[0025] In detail, the primary transcribed text refers to the audio dialogue text corresponding to the transcribed audio portion of the uploaded audio file, and the primary transcribed text includes the dialogue text content in the uploaded audio file that needs to be recorded and analyzed.
[0026] In this embodiment of the invention, reference is made to Figure 2 As shown, the extraction of speech text from the segmented audio sequence to obtain primary transcribed text includes: S21. Extract audio features from the segmented audio sequence feature by feature to obtain an audio feature sequence; S22. Perform temporal attention convolution on the audio feature sequence to obtain audio temporal features; S23. Decode the audio timing features to obtain decoded text; S24. Perform context repair on the decoded text to obtain the primary transcribed text.
[0027] In detail, the audio segmentation can be performed using the Voice Activity Detection (VAD) method. The feature-by-feature extraction of the segmented audio sequence refers to selecting segmented audio segments in the segmented audio sequence one by one and extracting audio features from the segmented audio segments. Mel Frequency Cepstral Coefficients (MFCC) or continuous wavelet transform can be used to extract audio features segment by segment.
[0028] Specifically, the step of performing temporal attention convolution on the audio feature sequence to obtain audio temporal features includes: The audio feature sequence is pre-convolved with temporal convolution to obtain convolutional audio features; The convolutional audio features are subjected to primary feedforward activation to obtain feedforward audio features; The feedforward audio features are globally enhanced using a multi-head attention mechanism to obtain attention-based audio features. The attention audio features are subjected to depthwise separable convolution to obtain depth audio features; Secondary feedforward activation is performed on the depth audio features to obtain activated audio features; The activated audio features are subjected to residual connections to obtain audio temporal features.
[0029] Specifically, the temporal convolution preprocessing includes sequentially performing convolution operations, activation operations, and layer normalization operations on the audio feature sequence. Here, two convolutional layers can be used for convolution operations, and the ReLU activation algorithm can be used for activation operations.
[0030] In detail, a primary feedforward activation can be performed using a feedforward network and a Swish activation function. The secondary feedforward activation has the same network layer structure as the primary feedforward activation. The secondary feedforward activation can further enhance the modeling ability of global features and improve the nonlinear expression ability of temporal information.
[0031] Specifically, the global enhancement refers to using a multi-head global attention mechanism to calculate the feature weights of each audio feature in the feedforward audio features, and using the feature weights to perform feature weighting updates on the feedforward audio features to obtain attention audio features.
[0032] In detail, the depthwise convolution can extract local temporal features, improve the ability to model short-term features, and thus enhance the perception of phoneme-level information.
[0033] Specifically, the Connectionist Temporal Classification (CTC) model can be used for text decoding, which means decoding each feature in the standard audio features into the corresponding text. The context repair refers to repairing and updating the content of the decoded text based on the text context features of the decoded text, such as using the Transformer model to correct spelling errors in the decoded text.
[0034] In this embodiment of the invention, by performing segment-by-segment text recognition on the uploaded audio file, a primary transcribed text is obtained. This allows for speech transcription by combining the local features and global contextual features of the uploaded audio file, thereby improving the accuracy of audio transcription.
[0035] S30: Calculate the analysis transcription progress and real-time transcription progress of the uploaded audio file, perform maximum value filtering on the analysis transcription progress and the real-time transcription progress to obtain the transcription progress, and update the primary transcribed text to standard transcribed text according to the transcription progress.
[0036] Specifically, the transcription progress refers to the progress of the uploaded audio file in being transcribed into the primary transcribed text, and the transcription progress refers to the percentage of the uploaded audio file that has been transcribed relative to the total content of the uploaded audio file.
[0037] In this embodiment of the invention, reference is made to Figure 3 As shown, the calculation of the analysis transcription progress and real-time transcription progress of the uploaded audio file includes: S31. Obtain the real-time transcription duration of the uploaded audio file according to the preset polling interval, and obtain the remaining file size and initial file size of the uploaded audio file respectively; S32. Analyze the total upload time of the uploaded audio file based on the initial file size to obtain the analysis transcription time; S33. Calculate the transcription analysis progress based on the real-time transcription duration and the transcription analysis duration; S34. Calculate the real-time transcription progress based on the remaining file size and the initial file size.
[0038] In detail, the polling interval refers to the fixed time interval corresponding to the repeated query. The polling interval can be 1 second. The real-time transcription duration refers to the time elapsed from the start of transcription of the uploaded audio file to the present. The real-time transcription duration can be calculated by the difference between the real-time time and the transcription start time. The real-time time refers to the current time, and the transcription start time refers to the timestamp when the transcription work begins.
[0039] Specifically, the remaining file size refers to the file size of the remaining uploaded audio files that have not been transcribed, and the initial file size is the same as the file size of the file to be transcribed.
[0040] In detail, the step of analyzing the total upload time of the uploaded audio file based on the initial file size to obtain the analysis transcription time includes: obtaining a preset transcription speed; and obtaining the analysis transcription time of the uploaded audio file by dividing the initial file size by the transcription speed. The transcription speed refers to the average transcription speed calculated based on the previous transcription history, combined with the file size of each file to be transcribed and the total time taken when the transcription is completely successful. For example, if a 47.38MB audio file requires 69 seconds to process, the transcription speed is 0.687MB / s.
[0041] Specifically, calculating the analysis transcription progress based on the real-time transcription duration and the analysis transcription duration means dividing the analysis transcription duration by the real-time transcription duration to obtain the analysis transcription progress.
[0042] In detail, calculating the real-time transcription progress based on the remaining file size and the initial file size means calculating the ratio of the remaining file size to the initial file size and using the ratio as the real-time transcription progress.
[0043] Specifically, the step of filtering the analyzed transcription progress and the real-time transcription progress to obtain the transcription progress means taking the maximum value of the analyzed transcription progress and the real-time transcription progress as the transcription progress, and updating the transcription progress using the upload.onprogress method of XMLHttpRequest.
[0044] Specifically, updating the primary transcribed text to standard transcribed text based on the transcription progress includes: determining whether the transcription progress is at a preset end progress threshold; if not, returning to the step of obtaining the real-time transcription duration of the uploaded audio file according to a preset polling interval; if yes, updating the primary transcribed text to standard transcribed text.
[0045] Specifically, the completion progress threshold refers to the transcription progress being 100%, and the standard transcribed text refers to all the primary transcribed text corresponding to the uploaded audio file after transcription is completed.
[0046] In this embodiment of the invention, the transcription progress is obtained by polling the uploaded audio file. The latest transcription progress can be obtained through polling and dual analysis, thereby improving the intuitiveness of audio transcription. By updating the primary transcribed text to standard transcribed text according to the transcription progress, global transcription of the audio file can be achieved, improving the efficiency of audio transcription analysis.
[0047] S40: Perform feature activation on the text context features of the standard transcribed text to obtain text summary features.
[0048] In detail, the text summarization feature refers to the text-hidden representation of the summary content of the standard transcribed text, i.e., feature vector encoding.
[0049] Specifically, the generated standard transcribed text may contain a lot of content, and directly analyzing the standard transcribed text is inefficient. Therefore, it is necessary to extract the text summary of the standard transcribed text and calculate the text context features.
[0050] In this embodiment of the invention, the step of performing feature activation on the text context features of the standard transcribed text to obtain text summary features includes: The standard transcribed text is segmented to obtain a sequence of transcribed text words; The transcribed text word sequence is subjected to word embedding operation to obtain the transcribed word feature sequence; Position encoding is performed on the transcribed word feature sequence to obtain the transcribed word encoding sequence; The text context features are obtained by using a bidirectional attention mechanism to extract dependency features from the transcribed word encoding sequence. The text context features are subjected to summarization feedforward activation to obtain text summarization features.
[0051] Specifically, bidirectional maximum matching can be used for text segmentation, which refers to splitting text into individual words. Word embedding refers to converting discrete words into continuous vector representations, and methods such as FastText and ELMo can be used for word embedding operations.
[0052] In detail, the positional encoding is used to supplement the sequence information of each transcribed word feature in the transcribed word feature sequence, and the positional encoding can be performed using sine coding or relative positional coding methods.
[0053] Specifically, the bidirectional attention mechanism refers to simultaneously considering the forward and backward information of the input sequence to improve the model's ability to understand the context. It can utilize the encoding of the bidirectional Transformer model to achieve dependent feature extraction. The summary feedforward activation refers to using a feedforward network to activate the text context features.
[0054] In this embodiment of the invention, by encoding the summary features of the standard transcribed text, the semantic understanding of the standard transcribed text can be enhanced by combining bidirectional contextual features, thereby obtaining concise text summary features and improving the efficiency of audio transcription analysis.
[0055] S50: Obtain the style keywords input by the user, and perform feature style decoding on the text summary features based on the style keywords to obtain the decoded text summary.
[0056] Specifically, the style keywords refer to the role-identifying keywords or language styles of the decoded text summary to be generated. For example, the style keywords can be "CEO", "product", "designer", "operations", "user interview", "testing", and "administrator".
[0057] Specifically, the decoded text summary refers to the text summary of the standard transcribed text under the style keyword. For example, when the file to be transcribed is a recording of a seminar on a certain product, and the style keyword is "designer", the decoded text summary is a text content summary of the product design part of the transcribed file.
[0058] In this embodiment of the invention, the step of performing feature style decoding on the text summary features based on the style keywords to obtain the decoded text summary includes: The style keywords are embedded to obtain style word features; The text summary features are cross-embedded using the style word features to obtain fused summary features; The fused summary features are decoded to obtain the decoded text summary.
[0059] In detail, the word embedding operation refers to encoding the style keywords into feature word vectors. The method of the word embedding operation is the same as that in step S40 above, and will not be repeated here.
[0060] Specifically, the feature cross-embedding refers to using a cross-attention mechanism to perform feature cross-fusion of the style word features and the text summary features. That is, attention weights are calculated using the keyword features and value features of the style word features and the query features of the text summary features, and the style word features and text summary features are weighted and fused according to the attention weights. The feature decoding includes the steps of feedforward activation, residual connection and feature activation of the fused summary features.
[0061] In this embodiment of the invention, by performing style decoding on the text summary features based on the style keywords, a decoded text summary is obtained, which enables customized role management and persistence operations, and realizes multi-role and multi-angle transcriptional text analysis, thereby improving the flexibility of transcriptional text analysis.
[0062] S60: Combine and splice the standard transcribed text and the decoded text summary to obtain an audio transcribed file.
[0063] Specifically, the audio transcription file includes the standard transcribed text and the decoded text summary. Generating the audio transcription file based on the standard transcribed text and the decoded text summary means combining and splicing the standard transcribed text and the decoded text summary to obtain the audio transcription file.
[0064] In this embodiment of the invention, by generating an audio transcription file based on the standard transcribed text and the decoded text summary, the results of audio transcription can be clearly displayed, and different types of summary text can be provided according to different keywords, thereby improving the efficiency of audio transcription analysis.
[0065] As can be seen, in the above scheme, by verifying the file format and integrity of the file to be transcribed, a file verification result is obtained. Based on the file verification result and the file size of the file to be transcribed, the file to be transcribed is sliced and uploaded to obtain an uploaded audio file. This ensures that the file to be transcribed is a processable audio file and enables efficient uploading of the file to be transcribed, improving the efficiency of audio transcription analysis. By segmenting the uploaded audio file to obtain segmented audio sequences, and extracting the speech text from the segmented audio sequences, segment-by-segment text recognition can be achieved to obtain primary transcribed text. Speech transcription can be performed by combining the local features and global context features of the uploaded audio file, improving the accuracy of audio transcription. By calculating the analyzed transcription progress and real-time transcription progress of the uploaded audio file, and performing maximum filtering on the analyzed transcription progress and real-time transcription progress, the transcription progress is obtained. The latest transcription progress can be obtained through polling and dual analysis, thereby improving the intuitiveness of audio transcription. By updating the primary transcribed text to standard transcribed text based on the transcription progress, global transcription of the audio file can be achieved, improving the efficiency of audio transcription analysis.
[0066] By activating the text context features of the standard transcribed text, the semantic understanding of the standard transcribed text can be enhanced by combining bidirectional context features, thereby obtaining concise text summary features and improving the efficiency of audio transcription analysis. By decoding the text summary features according to the style keywords, a decoded text summary is obtained, enabling customized role management and persistent operations, achieving multi-role and multi-angle transcribed text analysis, thus improving the flexibility of transcribed text analysis. By combining and splicing the standard transcribed text and the decoded text summary, an audio transcription file is obtained, clearly displaying the audio transcription results and providing different types of summary text based on different keywords, thereby improving the efficiency of audio transcription analysis. Therefore, the intelligent analysis-based audio transcription method and system proposed in this invention can solve the problem of low efficiency in audio transcription.
[0067] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0068] In one embodiment, an audio transcription device based on intelligent analysis is provided, which corresponds one-to-one with the audio transcription method based on intelligent analysis described in the above embodiments. For example... Figure 4 As shown, the audio transcription device based on intelligent analysis includes a file verification module 101, a text recognition module 102, a progress polling module 103, a summary encoding module 104, a style decoding module 105, and a result generation module 106. Detailed descriptions of each functional module are as follows: The file verification module 101 is used to perform file format verification and integrity verification on the file to be transcribed, obtain the file verification result, and upload the file slices of the file to be transcribed according to the file verification result and the file size of the file to be transcribed, thereby obtaining the uploaded audio file. The text recognition module 102 is used to perform audio segmentation on the uploaded audio file to obtain a segmented audio sequence, and extract the speech text of the segmented audio sequence to obtain primary transcribed text. The progress polling module 103 is used to calculate the analysis transcription progress and real-time transcription progress of the uploaded audio file, perform maximum value filtering on the analysis transcription progress and the real-time transcription progress to obtain the transcription progress, and update the primary transcribed text to standard transcribed text according to the transcription progress. The summarization encoding module 104 is used to activate the text context features of the standard transcribed text to obtain text summary features; The style decoding module 105 is used to obtain the style keywords input by the user, and perform feature style decoding on the text summary features according to the style keywords to obtain the decoded text summary. The result generation module 106 is used to combine and splice the standard transcribed text and the decoded text summary to obtain an audio transcribed file.
[0069] In one embodiment, when the file verification module 101 performs file format verification and integrity verification on the file to be transcribed and obtains the file verification result, it is used to: The file size is obtained by performing a size analysis on the file to be transcribed. The file extensions of the files to be transcribed are extracted to obtain the file extensions. The header features of the file to be transcribed are extracted to obtain the file header features; The format of the file to be transcribed is validated based on the file size, the file extension, and the file header features to obtain the format validation result. The integrity of the file to be transcribed is verified, and the integrity verification result is obtained. A file verification result is generated based on the format verification result and the integrity verification result.
[0070] In one embodiment, when the text recognition module 102 extracts the speech text from the segmented audio sequence to obtain the primary transcribed text, it is used to: The uploaded audio file is segmented to obtain a segmented audio sequence; Audio features are extracted from the segmented audio sequence feature by feature; Perform temporal attention convolution on the audio feature sequence to obtain audio temporal features; The standard audio features are then decoded to obtain the decoded text. The decoded text is then subjected to context repair to obtain the primary transcribed text.
[0071] In one embodiment, when the text recognition module 102 performs temporal attention convolution on the audio feature sequence to obtain audio temporal features, it is used to: The audio feature sequence is pre-convolved with temporal convolution to obtain convolutional audio features; The convolutional audio features are subjected to primary feedforward activation to obtain feedforward audio features; The feedforward audio features are globally enhanced using a multi-head attention mechanism to obtain attention-based audio features. The attention audio features are subjected to depthwise separable convolution to obtain depth audio features; Secondary feedforward activation is performed on the depth audio features to obtain activated audio features; The activated audio features are subjected to residual connections to obtain audio temporal features.
[0072] In one embodiment, when the progress polling module 103 performs the calculation of the analysis transcription progress and real-time transcription progress of the uploaded audio file, it is used to: The real-time transcription duration of the uploaded audio file is obtained according to a preset polling interval, and the remaining file size and initial file size of the uploaded audio file are obtained respectively. Based on the initial file size, the total upload time of the uploaded audio file is analyzed to obtain the transcription analysis time. The analysis transcription progress is calculated based on the real-time transcription duration and the analysis transcription duration. The real-time transcription progress is calculated based on the remaining file size and the initial file size.
[0073] In one embodiment, when the summarization encoding module 104 performs feature activation on the text context features of the standard transcribed text to obtain text summary features, it is used to: The standard transcribed text is segmented to obtain a sequence of transcribed text words; The transcribed text word sequence is subjected to word embedding operation to obtain the transcribed word feature sequence; Position encoding is performed on the transcribed word feature sequence to obtain the transcribed word encoding sequence; The text context features are obtained by using a bidirectional attention mechanism to extract dependency features from the transcribed word encoding sequence. The text context features are subjected to summarization feedforward activation to obtain text summarization features.
[0074] In one embodiment, when the style decoding module 105 performs feature style decoding on the text summary features based on the style keywords to obtain a decoded text summary, it is used to: The style keywords are embedded to obtain style word features; The text summary features are cross-embedded using the style word features to obtain fused summary features; The fused summary features are decoded to obtain the decoded text summary.
[0075] Specific limitations regarding the intelligent analysis-based audio transcription device can be found in the limitations of the intelligent analysis-based audio transcription method described above, and will not be repeated here. Each module in the aforementioned intelligent analysis-based audio transcription device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independent of the processor in a computer device, or stored in software in the memory of a computer device, so that the processor can call and execute the corresponding operations of each module.
[0076] In one embodiment, a computer device is provided, the internal structure of which can be shown in the following diagram. Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps of an audio transcription method based on intelligent analysis.
[0077] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: The file to be transcribed is subjected to file format and integrity verification to obtain file verification results. Based on the file verification results and the file size of the file to be transcribed, the file to be transcribed is sliced and uploaded to obtain the uploaded audio file. The uploaded audio file is segmented to obtain a segmented audio sequence. The speech text of the segmented audio sequence is extracted to obtain the primary transcribed text. The analysis transcription progress and real-time transcription progress of the uploaded audio file are calculated. The analysis transcription progress and the real-time transcription progress are filtered for maximum values to obtain the transcription progress. The primary transcribed text is then updated to standard transcribed text based on the transcription progress. Feature activation is performed on the text context features of the standard transcribed text to obtain text summary features; Obtain the style keywords input by the user, and perform feature style decoding on the text summary features based on the style keywords to obtain the decoded text summary; The standard transcribed text and the decoded text summary are combined and spliced to obtain an audio transcribed file.
[0078] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: The file to be transcribed is subjected to file format and integrity verification to obtain file verification results. Based on the file verification results and the file size of the file to be transcribed, the file to be transcribed is sliced and uploaded to obtain the uploaded audio file. The uploaded audio file is segmented to obtain a segmented audio sequence. The speech text of the segmented audio sequence is extracted to obtain the primary transcribed text. The analysis transcription progress and real-time transcription progress of the uploaded audio file are calculated. The analysis transcription progress and the real-time transcription progress are filtered for maximum values to obtain the transcription progress. The primary transcribed text is then updated to standard transcribed text based on the transcription progress. Feature activation is performed on the text context features of the standard transcribed text to obtain text summary features; Obtain the style keywords input by the user, and perform feature style decoding on the text summary features based on the style keywords to obtain the decoded text summary; The standard transcribed text and the decoded text summary are combined and spliced to obtain an audio transcribed file.
[0079] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0080] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0081] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0082] The above-described embodiments are merely illustrative of the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention. It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use.
Claims
1. An audio transcription method based on intelligent analysis, characterized in that, The method includes: The file to be transcribed is subjected to file format and integrity verification to obtain file verification results. Based on the file verification results and the file size of the file to be transcribed, the file to be transcribed is sliced and uploaded to obtain the uploaded audio file. The uploaded audio file is segmented to obtain a segmented audio sequence. The speech text of the segmented audio sequence is extracted to obtain the primary transcribed text. The analysis transcription progress and real-time transcription progress of the uploaded audio file are calculated. The analysis transcription progress and the real-time transcription progress are filtered for maximum values to obtain the transcription progress. The primary transcribed text is then updated to standard transcribed text based on the transcription progress. Feature activation is performed on the text context features of the standard transcribed text to obtain text summary features; Obtain the style keywords input by the user, and perform feature style decoding on the text summary features based on the style keywords to obtain the decoded text summary; The standard transcribed text and the decoded text summary are combined and spliced to obtain an audio transcribed file.
2. The audio transcription method based on intelligent analysis as described in claim 1, characterized in that, The file format and integrity checks of the file to be transcribed are performed to obtain the file verification results, including: The file size is obtained by performing a size analysis on the file to be transcribed. The file extensions of the files to be transcribed are extracted to obtain the file extensions. The header features of the file to be transcribed are extracted to obtain the file header features; The format of the file to be transcribed is validated based on the file size, the file extension, and the file header features to obtain the format validation result. The integrity of the file to be transcribed is verified, and the integrity verification result is obtained. A file verification result is generated based on the format verification result and the integrity verification result.
3. The audio transcription method based on intelligent analysis as described in claim 1, characterized in that, The extraction of speech text from the segmented audio sequence to obtain primary transcribed text includes: Audio features are extracted from the segmented audio sequence feature by feature; Perform temporal attention convolution on the audio feature sequence to obtain audio temporal features; The audio timing features are then decoded to obtain the decoded text. The decoded text is then subjected to context repair to obtain the primary transcribed text.
4. The audio transcription method based on intelligent analysis as described in claim 3, characterized in that, The step of performing temporal attention convolution on the audio feature sequence to obtain audio temporal features includes: The audio feature sequence is pre-convolved with temporal convolution to obtain convolutional audio features; The convolutional audio features are subjected to primary feedforward activation to obtain feedforward audio features; The feedforward audio features are globally enhanced using a multi-head attention mechanism to obtain attention-based audio features. The attention audio features are subjected to depthwise separable convolution to obtain depth audio features; Secondary feedforward activation is performed on the depth audio features to obtain activated audio features; The activated audio features are subjected to residual connections to obtain audio temporal features.
5. The audio transcription method based on intelligent analysis as described in claim 1, characterized in that, The calculation of the analysis transcription progress and real-time transcription progress of the uploaded audio file includes: The real-time transcription duration of the uploaded audio file is obtained according to a preset polling interval, and the remaining file size and initial file size of the uploaded audio file are obtained respectively. Based on the initial file size, the total upload time of the uploaded audio file is analyzed to obtain the transcription analysis time. The analysis transcription progress is calculated based on the real-time transcription duration and the analysis transcription duration. The real-time transcription progress is calculated based on the remaining file size and the initial file size.
6. The audio transcription method based on intelligent analysis as described in claim 1, characterized in that, The text context features of the standard transcribed text are activated to obtain text summarization features, including: The standard transcribed text is segmented to obtain a sequence of transcribed text words; The transcribed text word sequence is subjected to word embedding operation to obtain the transcribed word feature sequence; Position encoding is performed on the transcribed word feature sequence to obtain the transcribed word encoding sequence; Dependency features are extracted from the transcribed word coding sequence using a bidirectional attention mechanism to obtain text context features; The text context features are subjected to summarization feedforward activation to obtain text summarization features.
7. The audio transcription method based on intelligent analysis as described in claim 1, characterized in that, The step of performing feature style decoding on the text summary features based on the style keywords to obtain the decoded text summary includes: The style keywords are embedded to obtain style word features; The text summary features are cross-embedded using the style word features to obtain fused summary features; The fused summary features are decoded to obtain the decoded text summary.
8. An audio transcription device based on intelligent analysis, characterized in that, include: The file verification module is used to perform file format verification and integrity verification on the file to be transcribed, obtain the file verification result, and upload the file slices of the file to be transcribed based on the file verification result and the file size of the file to be transcribed, thereby obtaining the uploaded audio file. The text recognition module is used to perform audio segmentation on the uploaded audio file to obtain a segmented audio sequence, and extract the speech text of the segmented audio sequence to obtain the primary transcribed text; The progress polling module is used to calculate the analysis transcription progress and real-time transcription progress of the uploaded audio file, perform maximum value filtering on the analysis transcription progress and the real-time transcription progress to obtain the transcription progress, and update the primary transcribed text to standard transcribed text according to the transcription progress. The summary encoding module is used to activate the text context features of the standard transcribed text to obtain text summary features; The style decoding module is used to obtain the style keywords input by the user, and perform feature style decoding on the text summary features based on the style keywords to obtain the decoded text summary; The result generation module is used to combine and splice the standard transcribed text and the decoded text summary to obtain an audio transcribed file.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the audio transcription method based on intelligent analysis as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the audio transcription method based on intelligent analysis as described in any one of claims 1 to 7.