Audio processing method and related device

By extracting audio frame features and determining personalized subsequence lengths, the processing performance problem caused by fixed segmentation of audio signals is solved, and the efficiency and accuracy of audio processing are improved.

CN120690177APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510897268.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing technology uses a unified fixed-length segmentation for audio signals of different lengths, which affects the audio task processing performance. In particular, the computational complexity is high for long audio and the context information is insufficient for short audio, resulting in low processing efficiency.

Method used

By extracting features from audio frames, an audio feature sequence is generated. The subsequence length is personalized based on the total length and divided into multiple sub-feature sequences to adapt to audio signals of different lengths and improve processing performance.

Benefits of technology

The audio processing performance is improved. Through personalized subsequence length segmentation, it adapts to audio signals of different lengths, improving processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690177A_ABST
    Figure CN120690177A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method and a related device. The method comprises the following steps: performing feature extraction on an audio frame of an audio signal to obtain an audio frame feature; generating an audio feature sequence by using the audio frame features; based on the total length of the audio feature sequence, the subsequence length of the audio feature sequence is determined, and the total length represents the number of audio frame features contained in the audio feature sequence; and based on the sub-sequence length, segmenting the audio feature sequence into a plurality of first sub-feature sequences, and processing the audio signal according to the plurality of first sub-feature sequences. According to the invention, the processing performance of audio task processing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to artificial intelligence technology, and in particular to an audio processing method and related devices. Background Art

[0002] With the widespread application of audio recognition technology, there is a need to process audio signals of different lengths in real-world scenarios. When processing audio signals, it is often necessary to segment the audio feature sequence of the audio signal, and then perform subsequent task processing based on the sub-feature sequences obtained by segmentation, such as speech-to-text transcription. However, in related technologies, most audio signals of different lengths are segmented using a unified fixed length, which affects the processing performance of audio task processing. Summary of the Invention

[0003] The embodiments of the present application provide an audio processing method and related devices, which can improve the processing performance of audio task processing.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] This embodiment of the present application provides an audio processing method, the method comprising:

[0006] Perform feature extraction on the audio frame of the audio signal to obtain audio frame features;

[0007] generating an audio feature sequence using the audio frame features;

[0008] Determining a subsequence length of the audio feature sequence based on a total length of the audio feature sequence, wherein the total length represents the number of audio frame features included in the audio feature sequence;

[0009] Based on the subsequence length, the audio feature sequence is divided into a plurality of first sub-feature sequences, and the audio frame features are processed according to the first sub-feature sequences.

[0010] The present invention provides an audio processing device, including:

[0011] A feature extraction module is used to extract features from audio frames of audio signals to obtain audio frame features;

[0012] A sequence generation module, configured to generate an audio feature sequence using the audio frame features;

[0013] a length determination module, configured to determine a subsequence length of the audio feature sequence based on a total length of the audio feature sequence, wherein the total length represents the number of audio frame features included in the audio feature sequence;

[0014] A sequence processing module is configured to divide the audio feature sequence into a plurality of first sub-feature sequences based on the sub-sequence lengths, and process the audio signal according to the plurality of first sub-feature sequences.

[0015] An embodiment of the present application provides an electronic device, comprising:

[0016] a memory for storing computer-executable instructions or computer programs;

[0017] The processor is configured to implement the audio processing method provided in the embodiment of the present application when executing the computer-executable instructions or computer program stored in the memory.

[0018] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the audio processing method provided in the embodiment of the present application when executed by a processor.

[0019] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the audio processing method provided in the embodiment of the present application is implemented.

[0020] The embodiments of the present application have the following beneficial effects: first, feature extraction is performed on the audio frames of the audio signal, and the audio frame features extracted from the audio frames are used to construct an audio feature sequence of the speech signal; then, a personalized subsequence length is determined for the audio feature sequence, so that a subsequence length adapted to audio signals of different lengths can be obtained; and the audio feature sequence is segmented based on the subsequence length to obtain a subfeature sequence matching the length of the audio signal, that is, the length of the obtained subfeature sequence is more reasonable; finally, the audio signal is processed based on the subfeature sequence with a more reasonable length, thereby improving the processing performance of the audio processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is a schematic diagram of an application environment of the audio processing method provided in an embodiment of the present application;

[0022] Figure 2 This is a flow diagram of the audio processing method provided in the embodiment of the present application. Figure 1 ;

[0023] Figure 3 This is a flow diagram of the audio processing method provided in the embodiment of the present application. Figure 2 ;

[0024] Figure 4 This is a flow diagram of the audio processing method provided in the embodiment of the present application. Figure 3 ;

[0025] Figure 5 This is a flow diagram of the audio processing method provided in the embodiment of the present application. Figure 4 ;

[0026] Figure 6 Schematic diagram of the text conversion process of a speech signal provided in an embodiment of the present application;

[0027] Figure 7 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0029] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0030] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0031] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0032] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0033] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0034] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0035] 1) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0036] 2) Audio signals refer to the acoustic representation of human speech, which is an electrical or digital signal obtained through sensors such as microphones. Audio signals contain the content of human speech, including semantics, speaker identity, emotion, and speech rate.

[0037] 3) Audio frames are formed by segmenting a continuous audio signal into overlapping or non-overlapping segments of fixed length. Each segment is called an audio frame. Audio frames serve as the basic unit of audio signal analysis, enabling feature extraction and other processing.

[0038] 4) Audio feature sequence is a structured data representation obtained by extracting features from audio frames frame by frame, so each feature in the audio feature sequence corresponds to an audio frame. For example, if the audio signal is framed, the audio frame sequence obtained is represented as {x1, x2, ..., x T}, then the audio feature sequence can be expressed as {f1,f2,…,f T}, where f t is the corresponding audio frame x t characteristics.

[0039] With the widespread application of audio recognition technology, practical scenarios require the processing of audio signals of varying lengths. For example, long audio signals, such as meeting and lecture recordings, need to be transcribed from speech to text, while short audio signals, such as voice commands and short conversations, need to be transcribed from speech to text. When processing audio signals, it is often necessary to segment the audio feature sequence of the audio signal, and then perform subsequent processing tasks, such as speech to text transcription, based on the sub-feature sequences obtained by segmentation.

[0040] However, in the related art, for audio signals of different lengths, most of them are segmented with a unified fixed length. For example, a 1-hour audio signal A and a 5-minute audio signal B are both segmented in units of 10s. In this way, a large number of sub-feature sequences will be obtained for long audio, and the more sub-feature sequences there are, the higher the computational complexity of processing the audio signal, which results in long audio taking up more time to complete processing, low processing efficiency, and the more sub-feature sequences there are, the greater the memory usage will be, further affecting processing efficiency. For short audio, only a few sub-feature sequences may be obtained, making the sub-feature sequences of short audio too sparse, that is, insufficient contextual information, making it difficult to accurately perform subsequent task processing, resulting in reduced processing accuracy.

[0041] To sum up, in the related art, if the audio feature sequences of audio signals of different lengths are segmented according to the same fixed rules, the length of the obtained sub-feature sequences will not match the length of the audio signals, thereby affecting the processing performance of the audio task processing.

[0042] In addition, in related technologies, there is also the problem of low adaptability of audio processing models to multi-language tasks and multi-domain tasks.

[0043] In response to at least one of the above-mentioned problems existing in the related art, an embodiment of the present application provides an audio processing method and related apparatus, the method comprising: extracting features from audio frames of an audio signal to obtain audio frame features; generating an audio feature sequence using the audio frame features; determining the subsequence length of the audio feature sequence based on the total length of the audio feature sequence, wherein the total length represents the number of audio features contained in the audio feature sequence; dividing the audio feature sequence into multiple first sub-feature sequences based on the subsequence length, and processing the audio signal based on the multiple first sub-feature sequences. In this way, a subsequence length that is compatible with audio signals of different lengths can be obtained, and the audio feature sequence can be divided based on the subsequence length to obtain a sub-feature sequence that matches the length of the audio signal, that is, the length of the obtained sub-feature sequence is more reasonable, and the audio signal is processed based on the sub-feature sequence with a more reasonable length, thereby improving the processing performance of the audio processing.

[0044] In order to better understand the audio processing method and related devices provided in the embodiments of the present application, the application environment applicable to the embodiments of the present application is described below.

[0045] See also Figure 1 , Figure 1Schematic diagram of an application environment of the audio processing method provided in the embodiment of the present application. As an implementation method, the audio processing method of the embodiment of the present application can be applied to an electronic device, wherein, wherein, the electronic device can be such as Figure 1 The server 110 shown in FIG. 1 can be connected to the terminal 120 via a network 130. The network 130 is used to provide a medium for a communication link between the server 110 and the terminal 120. The network 130 can include various connection types, such as wired communication links, wireless communication links, etc., which are not limited in this embodiment of the present application.

[0046] It should be understood that Figure 1 The server 110, terminal 120, and network 130 are merely illustrative. Any number of servers, networks, and terminals may be used as needed. For example, the server 110 may be a physical server or a server cluster consisting of multiple servers, and the terminal 120 may be a smartphone, tablet computer, desktop computer, laptop computer, smartwatch, or other device. It will be appreciated that in embodiments of the present application, multiple terminals 120 may be allowed to access the server 110 simultaneously.

[0047] The embodiment of the present application can be implemented by a server. For example, the server 110 will obtain multiple first sub-feature sequences for the audio signal that needs to be processed based on the audio processing method provided in the embodiment of the present application, so that feature processing can be performed on the audio signal based on the multiple first sub-feature sequences.

[0048] The audio processing method provided in the embodiments of the present application can be applied to various audio signal processing scenarios, such as generating transcripts from conference recordings, analyzing the sentiment of telephone voices, etc. Below, the scenarios in which the audio processing method provided in the embodiments of the present application can be applied are described.

[0049] 1) Transcript generation scenario. For example, for a conference recording, the server divides the conference recording into multiple first sub-feature sequences according to the audio processing method provided in the embodiment of the present application, and then generates a transcript of the conference recording based on the multiple first sub-feature sequences.

[0050] 2) Emotional analysis scenario, for example, for the user's telephone voice, the server repeatedly divides the telephone voice into multiple first sub-feature sequences according to the audio processing proposed in the embodiment of the present application, and then determines the emotional category of the telephone voice based on the multiple first sub-feature sequences, such as happiness, anger, etc.

[0051] The following describes the audio processing method provided by the embodiment of the present application. As previously mentioned, the electronic device that implements the audio processing method of the embodiment of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.

[0052] See also Figure 2 , Figure 2 This is a flow diagram of the audio processing method provided in the embodiment of the present application. Figure 1 , will combine Figure 2 The steps shown are explained, Figure 2 The main body of the step is the electronic device.

[0053] Step 101: extract features from audio frames of an audio signal to obtain audio frame features.

[0054] The present embodiment is implemented in the context of audio signal processing to generate a first sub-feature sequence for the audio signal for subsequent task processing. After acquiring the audio signal to be processed, the electronic device first segments the audio signal into multiple audio frames. The electronic device then extracts features from the obtained audio frames, and the extracted features are referred to as audio frame features.

[0055] In the embodiments of this application, an audio signal refers to an electrical or digital signal, the acoustic representation of human speech, obtained through a sensor such as a microphone. An audio signal can include a conference recording, a customer service recording, or a voice message on a social network. An audio frame is a continuous audio signal segmented into overlapping or non-overlapping segments of fixed length. This means that audio frames can be overlapping or non-overlapping.

[0056] Electronic devices can divide audio signals into audio frames by using a sliding window segmentation method. First, the electronic device obtains parameters such as frame length, frame shift step, and window type, and then slides the window on the time axis according to the frame shift step, and uses the sampling points in the window as an audio frame. Thus, each time the window is slid, an audio frame is obtained. In this way, the continuous audio signal can be divided into several audio frames. If the length of the end is less than an audio frame, it is padded with zeros to ensure the length of all audio frames. Before starting the window sliding, the electronic device can also resample the audio signal to a standard frequency (such as 16kHz) to eliminate the sampling rate differences caused by different devices. The electronic device can normalize the amplitude of the audio signal to the range of [-1,1] to prevent the amplitude of the audio signal from affecting the stability of the frame division.

[0057] The electronic device can perform a short-time Fourier transform (STFT) on the audio frame to convert the time domain signal contained in the audio frame into a frequency domain energy distribution, and then calculate the Mel-frequency cepstral coefficients based on the frequency domain energy spectrum, and use the obtained Mel-frequency cepstral coefficients as the audio frame features of the audio frame. More specifically, the electronic device can map the linear frequency to the Mel scale to determine 20-40 triangular filters, and then pass the frequency domain energy spectrum through the obtained triangular filters to calculate the energy sum of each frequency band, and then take the natural logarithm of the energy sum of each frequency band to obtain a Log-Mel spectrum, and then perform a discrete cosine transform on the Log-Mel spectrum, and use the first 12-13 dimensional coefficients as the Mel-frequency cepstral coefficients.

[0058] Electronic devices can also calculate the short-time energy and zero-crossing rate of audio frames and use them as audio frame features. Short-time energy refers to the sum of the energy of all sampling points in an audio frame, reflecting the strength of the audio frame. The zero-crossing rate indicates the number of times the waveform of a speech frame crosses the zero point, reflecting the frequency characteristics of the speech frame.

[0059] The short-time energy can be calculated by the formula. For example, for an audio frame x[n] (n=0, 1, ..., N-1) containing N sampling points, its short-time energy E can be obtained by the formula of the form (1):

[0060]

[0061] The zero crossing rate can be calculated by the formula. For example, for an audio frame x[n] (n=0, 1, ..., N-1) containing N sampling points, its zero crossing rate ZCR can be obtained by the formula of form (2):

[0062]

[0063] Where sign(x) is the sign function, which takes the value 1 when x>0 and takes the value -1 otherwise. The denominator 2(N-1) is the normalization coefficient.

[0064] Step 102: Generate an audio feature sequence using audio frame features.

[0065] It should be noted that after the electronic device segments the audio signal, it can obtain multiple audio frames. For each audio frame, the electronic device will perform feature extraction to obtain corresponding audio frame features, thereby obtaining multiple audio frame features. The electronic device then arranges the multiple audio frame features in the time sequence of the audio frames, or in reverse time sequence. The resulting sequence is the audio feature sequence.

[0066] Step 103: Determine the subsequence length of the audio feature sequence based on the total length of the audio feature sequence.

[0067] After obtaining the audio feature sequence of the audio signal, the electronic device will count the number of audio frame features contained in the audio feature sequence, and use the number obtained by counting as the total length of the audio feature sequence. Thus, the total length of the audio feature sequence represents the number of audio frame features contained in the audio feature sequence. Afterwards, the electronic device will calculate the length of the segmentation adapted to the audio feature sequence of the audio signal based on the total length of the audio feature sequence, that is, the subsequence length. Thus, in the embodiment of the present application, the subsequence length is obtained by performing personalized operations on the audio feature sequence, so that audio feature sequences of different lengths will have different subsequence lengths. In this way, the rationality of the segmentation length of the audio feature sequence can be guaranteed, so that the audio feature sequence can be segmented based on a more personalized and reasonable segmentation length in the future, so as to obtain a first sub-feature sequence with a more reasonable length.

[0068] Exemplarily, in an embodiment of the present application, it is possible to achieve a subsequence length of 20 audio frames for long audio, such as an audio signal of 1000 audio frames, and a subsequence length of 5 audio frames for short audio, such as an audio signal of 50 audio frames. Compared with the related art, in which 10 audio frames are used as the sub-length sequence for audio signals of 1000 audio frames and audio signals of 50 audio frames, the personalized subsequence length in the embodiment of the present application is more adapted to the total length of the audio signal, so that personalized segment lengths can be obtained for different audio signals, making the subsequence more reasonable.

[0069] See also Figure 3 , Figure 3 This is a flow diagram of the audio processing method provided in the embodiment of the present application. Figure 2 In some embodiments of the present application, Figure 2 Step 103 in the above, i.e., determining the subsequence length of the audio feature sequence based on the total length of the audio feature sequence, can be implemented by the following process:

[0070] Step 1031: Determine a first sub-length determination method that matches the audio feature sequence based on the total length of the audio feature sequence.

[0071] In the embodiment of the present application, different sub-length determination methods can be used, that is, there can be multiple different calculation methods for the subsequence length, and the subsequence lengths obtained by using different sub-length determination methods may also be different. Therefore, in the embodiment of the present application, the electronic device can assign an appropriate sub-length determination method to the audio feature sequence based on the total length of the audio feature sequence, so that the subsequence length can be calculated for the audio feature sequence based on the appropriate sub-length determination method, thereby enabling a personalized subsequence length to be obtained for the audio feature sequence.

[0072] In some embodiments of the present application, Figure 3 Step 1031, i.e., determining the first sub-length determination method that matches the audio feature sequence based on the total length of the audio feature sequence, can be achieved by the following processing: determining the length interval to which the total length belongs from the preset first length interval and the preset second length interval, where the left boundary of the first length interval is greater than the right boundary of the second length interval; and determining the second sub-sequence length determination method corresponding to the length interval to which the total length belongs from multiple second sub-sequence length determination methods as the first sub-length determination method.

[0073] In an embodiment of the present application, different sub-length determination methods are preset for different length intervals, so that the electronic device can determine a suitable sub-length determination method for the audio feature sequence according to the total length of the audio feature sequence. The electronic device first compares the total length of the audio signal with the left boundary and the right boundary of the first length interval and the left boundary and the right boundary of the second length interval, thereby determining the length interval in which the total length falls. This length interval is the length interval to which the total length of the audio signal belongs, and then directly uses the second sub-length determination method corresponding to the length interval to which the total length belongs among the preset multiple second sub-length determination methods as the first sub-length determination method. Among them, the left boundary refers to the starting value of the length interval, that is, the minimum value of the numerical range corresponding to the length interval, and the right boundary refers to the end of the length interval, that is, the maximum value of the numerical range corresponding to the length interval.

[0074] It should be noted that the left boundary of the first length interval is greater than the right boundary of the second length interval, that is, each value in the first length interval will be greater than the maximum value of the second length interval. For example, if the number of audio frame features, that is, the number of audio frames, is taken as a unit, the first length interval can be [2000, 10000], and the second length interval can be [1, 1999]. It can be seen that if the length interval to which the total length of the audio signal belongs is the first length interval, it means that the audio signal is long audio, and its subsequence length needs to be determined according to the subsequence length determination method corresponding to long audio; if the length interval to which the total length of the audio signal belongs is the second length interval, it means that the audio signal is short audio, and its subsequence length needs to be determined according to the subsequence length determination method corresponding to short audio.

[0075] In this embodiment of the present application, the maximum value of the subsequence length can be selected as the right boundary value of the second length interval, so that the left boundary value of the first length interval needs to be greater than the maximum value of the subsequence length. For example, if the maximum value of the subsequence length is C max , then the left boundary value of the first length interval needs to be greater than C max .

[0076] It can be understood that in an embodiment of the present application, the electronic device can determine the most appropriate first sub-length determination method for the audio feature sequence based on the total length of the audio feature sequence, so that the most appropriate sub-length sequence can be determined for the audio feature sequence based on the most appropriate first sub-length determination method in the future, thereby obtaining a personalized sub-length sequence.

[0077] In other embodiments of the present application, Figure 3 Step 1031, i.e., determining the first sub-length determination method that matches the audio feature sequence based on the total length of the audio feature sequence, can also be achieved by the following processing: using a determination method prediction model, based on the total length of the audio feature sequence and the audio feature sequence, predicting the first sub-length determination method that matches the audio feature sequence.

[0078] The electronic device can also use a determination method prediction model to read the total length of the audio feature sequence and the audio feature sequence itself, so as to predict the word length determination method for the audio feature sequence based on the read content, and determine the prediction result of the determination method prediction model as the first sub-length determination method. The determination method prediction model here can be a large language model that has been pre-trained, or a large model obtained by targeted fine-tuning on the large language model, and the embodiments of the present application are not limited thereto.

[0079] Step 1032: Determine the subsequence length using the first subsequence length determination method.

[0080] After obtaining the first sub-length determination method, the electronic device processes the audio feature sequence according to the specific processing procedure provided by the first sub-length determination method to obtain a sub-sequence length of the audio feature sequence. The first sub-length determination method may include first obtaining a first ratio parameter and then determining the sub-sequence length by combining the first ratio parameter and the total length.

[0081] In some embodiments of the present application, Figure 3 Step 1032 in the above, i.e., determining the subsequence length by using the first sub-length determination method, can be implemented by the following processing: obtaining a first proportion parameter, the first proportion parameter including: a first sub-parameter and a second sub-parameter, the first sub-parameter being greater than the second sub-parameter; in response to the length interval to which the total length belongs being the first length interval, adjusting the total length using the first sub-parameter, and determining the minimum value between the adjusted length and the first preset length as the subsequence length, wherein the first preset length is the maximum value of the subsequence length; in response to the length interval to which the total length belongs being the second length interval, adjusting the total length using the second sub-parameter, and determining the maximum value between the adjusted length and the second preset length as the subsequence length, wherein the second preset length is the minimum value of the subsequence length.

[0082] It should be noted that the first scale parameter is a parameter that controls the proportional relationship between the subsequence length and the total length. That is, the first scale parameter represents the proportional relationship between the subsequence length and the total length. In the embodiment of the present application, the electronic device may determine the first scale parameter based on the complexity of the audio signal, or may directly determine a manually preset scale parameter as the first scale parameter.

[0083] In some embodiments of the present application, Figure 3 Step 1031 in the above, i.e., obtaining the first proportional parameter, can be implemented by the following processing: obtaining the complexity of the audio signal; and using a preset parameter corresponding to the complexity among multiple preset parameters as the first proportional parameter.

[0084] That is to say, the electronic device stores multiple different preset parameters in advance, and establishes a mapping relationship between these preset parameters and different levels of complexity. The electronic device can first analyze the complexity of the audio signal to obtain the complexity of the audio signal, and then obtain the preset parameters of the complexity corresponding to the audio signal from the multiple preset parameters, and use the obtained parameters as the first proportional parameters.

[0085] It should be noted that the complexity of an audio signal is a comprehensive indicator of the audio signal, which is used to measure the comprehensive manifestation of the information density and noise interference of the audio signal. Here, information density refers to the amount of effective information contained in the audio signal per unit time, which can reflect the richness of the details of the audio signal in time or space; noise interference refers to the presence of non-target, irregular random signals in the audio signal, which will mask the effective information in the audio signal and reduce the quality of the audio signal. Therefore, both information density and noise interference will affect the complexity of the audio signal. For example, for two audio signals with the same information density, the audio signal with stronger noise interference will be more difficult to extract effective information for processing, and the processing difficulty will be greater, that is, the audio signal will have a higher complexity; while for two audio signals with noise interference, the audio signal with greater information density will be more difficult to process, and thus will have a higher complexity. It can be seen that the complexity of an audio signal can characterize the difficulty of processing an audio signal.

[0086] The complexity of the audio signal can be calculated in real time or manually set. If the complexity of the audio signal is calculated in real time, the electronic device can obtain the complexity of the audio signal through real-time calculation. If the complexity of the audio signal is manually set, the electronic device can directly obtain the complexity of the audio signal from the storage space.

[0087] In some embodiments of the present application, obtaining the complexity of the audio signal in the above content can be achieved through the following processing: performing speech rate analysis on the audio signal to obtain the speech rate information of the audio signal; performing noise intensity analysis on the audio signal to obtain the noise intensity information of the audio signal; and calculating the complexity of the audio signal based on the speech rate information and the noise intensity information.

[0088] It should be noted that the information density of an audio signal is proportional to its speech rate. Therefore, electronic devices can determine the information density of an audio signal by analyzing its speech rate. Similarly, the noise intensity of an audio signal is proportional to its interference intensity. Therefore, electronic devices can determine its interference intensity by analyzing its noise intensity. Subsequently, electronic devices can calculate the complexity of the audio signal by combining the speech rate and noise intensity information.

[0089] The speech rate analysis of the audio signal in the embodiment of the present application can be achieved through the following processing: the electronic device extracts features such as zero-crossing rate, fundamental frequency variance, Mel-frequency cepstral coefficients, spectral entropy, etc. for all audio frames of the audio signal, and then provides these features to a machine learning model for estimating the speech rate of the audio, such as a support vector regression model, a convolutional neural network model, etc., and determines the output of these models as the speech rate information of the audio signal; or the electronic device performs band-pass filtering on the audio signal, and calculates the enhanced syllable peak of the Teager Energy Operator (TEO), and then determines the local maximum of the TEO energy envelope, and counts the number of peaks per unit time to obtain the speech rate information.

[0090] The noise intensity analysis of the audio signal in the embodiment of the present application can be achieved through the following processing: the electronic device first calculates the short-time energy for each audio frame of the audio signal, divides the audio frame into silent frames and non-silent frames according to the short-time energy and the energy threshold, and extracts all silent frames, then calculates the noise energy for the silent frames, and then calculates the overall energy of the audio signal, subtracts the overall energy from the noise energy to obtain the effective audio energy of the non-silent frames, and then calculates the signal-to-noise ratio of the effective audio energy and the noise energy, and uses the signal-to-noise ratio as the noise intensity information of the audio signal; or the electronic device extracts features such as zero-crossing rate, fundamental frequency variance, Mel-frequency cepstral coefficients, spectral entropy, etc. for all audio frames of the audio signal, provides these features to a machine learning model for noise information estimation, and uses the output of the machine learning model as the noise intensity information of the audio signal.

[0091] After obtaining the speech rate information and noise intensity information, the electronic device can perform weighted summation on the speech rate information and the noise intensity information, and use the weighted summation result as the complexity of the audio signal, or provide the speech rate information and the noise intensity information to a machine learning model for complexity prediction, such as a decision tree model or a neural network model, and use the output of the machine learning model as the complexity of the audio signal.

[0092] It can be understood that in the embodiment of the present application, the electronic device can perform speech rate analysis and noise intensity analysis on the audio signal to obtain the speech rate information and noise intensity information of the audio signal respectively, and at the same time combine the information of the two dimensions of speech rate information and noise intensity information to comprehensively analyze the overall processing difficulty of the audio signal, thereby obtaining a more accurate degree of complexity.

[0093] In other embodiments of the present application, obtaining the complexity of the audio signal in the above content can be achieved through the following processing: dividing the audio signal into multiple audio frames, extracting the energy value of each audio frame, and dividing the audio frames into speech frames and non-speech frames based on the energy value and energy threshold; counting the proportion of speech frames, and determining the complexity of the audio signal based on the proportion.

[0094] It should be noted that a higher proportion of speech frames indicates that speech in the audio signal is easier to recognize, thereby reducing the complexity of the audio signal. A lower proportion of speech frames indicates that speech in the audio signal is more difficult to recognize, possibly due to strong background noise interference, thereby increasing the complexity of the audio signal. Using the above processing method, the electronic device can also obtain the complexity of the audio signal.

[0095] It can be understood that in an embodiment of the present application, the electronic device can first determine the complexity of the audio signal, and based on its complexity, select a suitable first proportional parameter for the audio signal from multiple preset parameters. In this way, the obtained first proportional parameter can be more in line with the actual situation of the audio signal, thereby facilitating the subsequent acquisition of a more reasonable subsequence length based on the first proportional parameter.

[0096] It should also be noted that, in the embodiment of the present application, the first scale parameter may include a first sub-parameter and a second sub-parameter, that is, the first scale parameter may be a parameter set composed of multiple sub-parameters, that is, the first scale parameter may include multiple different sub-parameters, such as {0.1, 0.3}, or {0.2, 0.4}, etc. In this way, the electronic device subsequently needs to select a suitable sub-parameter from the first scale parameter to calculate the subsequence length with the total length of the audio signal.

[0097] After obtaining the first scale parameter, the electronic device calculates the subsequence length according to the subsequence length determination method corresponding to the length interval, combining the total length of the audio signal and the first scale parameter, and uses the calculated result as the subsequence length. If the length interval to which the total length belongs is the first length interval, the electronic device first selects the larger first subparameter from the first subparameter and the second subparameter as the subparameter for performing a scale operation on the total length. The electronic device then uses the first subparameter to adjust the total length of the audio signal, i.e., reducing the total length by the ratio of the first subparameter to obtain an adjusted length. The electronic device then compares the adjusted length with the first preset length, and determines the minimum of the adjusted length and the first preset length as the final subsequence length. If the length interval to which the total length belongs is the second length interval, the electronic device uses the smaller second subparameter of the first subparameter and the second subparameter as the subparameter for performing a scale operation on the total length. The electronic device then reduces the total length by the ratio of the second subparameter to obtain an adjusted length. The electronic device then compares the adjusted length with the second preset length, and determines the minimum of the adjusted length and the second subparameter as the final subsequence length. In this way, the value of the subsequence length can be limited between the second preset length and the first preset length, so as to avoid the extreme situation that the subsequence length is too small or too large.

[0098] It should be noted that the first sub-parameter and the second sub-parameter can be a single parameter or a parameter set. In this case, the subsequence length can also be a length set, with each length in the length set corresponding to a parameter in the parameter set. The electronic device will subsequently segment the speech feature sequence according to each length in the length set, so that the characteristics of the resulting multiple sub-feature sequences are different.

[0099] It can be understood that in the embodiment of the present application, the electronic device will first obtain a first proportional parameter, and select a suitable sub-parameter from the first proportional parameter to adjust the total length according to whether the length interval to which the total length belongs is the first length interval or the second length interval, so as to achieve adaptive segment length determination for audio signals of different lengths, and then use the first preset length and the second preset length to constrain the upper and lower limits of the adjusted length to ensure that the segment length is within a reasonable range.

[0100] In other embodiments of the present application, Figure 2 Step 103, i.e., determining the subsequence length of the audio feature sequence based on the total length of the audio feature sequence, can also be achieved by the following processing: reading the total length of the audio feature sequence through the sub-length prediction model to predict the subsequence length, and determining the prediction result of the large model as the subsequence length.

[0101] The sub-length prediction model can be obtained by fine-tuning a pre-trained large model to obtain the ability to predict reasonable sub-sequence lengths based on the total length of the audio feature sequence. The training data can be sample pairs consisting of the total length of a training sample and its most reasonable sub-sequence length.

[0102] Step 104: Divide the audio feature sequence into multiple first sub-feature sequences based on the sub-sequence length.

[0103] After obtaining the subsequence length, the electronic device can perform non-overlapping segmentation of the speech feature sequence according to the subsequence length to obtain multiple non-overlapping first subfeature sequences, or perform overlapping segmentation of the speech feature sequence to obtain multiple overlapping first subfeature sequences. The subsequence length represents the number of audio frame features contained in the first subfeature sequence, and the multiple first subfeature sequences are used as a basis for task processing of the audio signal, such as as a feature basis for generating a transcript of the audio signal, or as a feature basis for sentiment classification of the audio signal.

[0104] Step 105: Process the audio signal according to the multiple first sub-feature sequences.

[0105] Finally, the electronic device processes the audio signal based on the multiple first sub-feature sequences. For example, to generate a transcript of the audio signal, the electronic device may process the first sub-feature sequence using the first model to obtain a corresponding text segment, and then merge the multiple text segments corresponding to the first sub-feature sequences into a transcript of the audio signal. For another example, to perform emotion classification on the audio signal, the electronic device may perform emotion recognition on the multiple first sub-feature sequences using the second model to obtain a corresponding classification label.

[0106] It can be understood that, compared with the related art, since the audio feature sequences for audio signals of different lengths are segmented according to the same fixed rules, the length of the obtained sub-feature sequence does not match the length of the audio signal, which leads to the problem that the processing performance of the audio task processing is affected. In the embodiment of the present application, the electronic device will first segment the audio signal, and use the audio frame features extracted from the audio frame to construct the audio feature sequence of the speech signal, and then determine the personalized sub-sequence length for the audio feature sequence, so as to obtain a sub-sequence length that is adapted to audio signals of different lengths, and split the audio feature sequence based on the sub-sequence length to obtain a sub-feature sequence that matches the length of the audio signal, that is, the length of the obtained sub-feature sequence is more reasonable, and finally the audio signal is processed based on the sub-feature sequence with a more reasonable length, thereby improving the processing performance of the audio processing.

[0107] based on Figure 2 , see Figure 4 , Figure 4 This is a flow diagram of the audio processing method provided in the embodiment of the present application. Figure 3 In some embodiments of the present application, Figure 2 After step 104 in

[15] , that is, after dividing the audio feature sequence into a plurality of first sub-feature sequences based on the sub-sequence length, the method may further include the following processing:

[0108] Step 106: Determine context information of each first sub-feature sequence based on the first sub-feature sequences adjacent to each first sub-feature sequence.

[0109] It should be noted that the first subfeature sequences adjacent to each first subfeature sequence may include the first subfeature sequence preceding the first subfeature sequence, i.e., the forward-adjacent first subfeature sequence, or may include the first subfeature sequence following the first subfeature sequence, i.e., the backward-adjacent first subfeature sequence. The electronic device determines context information for each first subfeature sequence by combining at least one of the forward-adjacent first subfeature sequence and the backward-adjacent first subfeature sequence of each first subfeature sequence.

[0110] In an embodiment of the present application, the context information of the first sub-feature sequence refers to the temporal or semantic extension and association with the first sub-feature sequence, providing feature fragments that are temporally associated with the first sub-feature sequence, which is used to enhance the understanding and modeling of the first sub-feature sequence to resolve the ambiguity and uncertainty of the local first sub-feature sequence.

[0111] Figure 5 This is a flow diagram of the audio processing method provided in the embodiment of the present application. Figure 4 In some embodiments of the present application, Figure 4 Step 106 in the above, i.e., determining the context information of each first sub-feature sequence based on the first sub-feature sequence adjacent to each first sub-feature sequence, can be implemented by the following process:

[0112] Step 1061: Obtain a second ratio parameter, where the second ratio parameter is a parameter that controls the ratio between the length of the context information and the length of the subsequence.

[0113] When determining context information for each first sub-feature sequence, the electronic device first needs to determine a second scaling parameter for controlling the proportional relationship between the length of the context information and the length of the subsequence. In the embodiment of the present application, the second scaling parameter can be a preset fixed scaling parameter or a preset parameter obtained from a plurality of different preset parameters based on the complexity of the audio signal.

[0114] It should be noted that, when the adjacent first sub-feature sequences include only one of the forward-adjacent first sub-feature sequences and the backward-adjacent first sub-feature sequences, the second proportion parameter may be only a single parameter, and when the adjacent first sub-feature sequences include both the forward-adjacent first sub-feature sequences and the backward-adjacent sub-feature sequences, the second proportion parameter will also include the sub-parameters of the forward-adjacent first sub-feature sequence and the sub-parameters of the backward-adjacent first sub-feature sequence.

[0115] Step 1062: Determine the product of the second scale parameter and the subsequence length as the first context length.

[0116] After obtaining the second scale parameter, the electronic device performs a product operation on the second scale parameter and the subsequence length, and uses the resulting product as the first context length. If the second scale parameter is a single parameter, the resulting first context length is also a single length. If the second scale parameter includes multiple sub-parameters, the resulting first context length includes both the forward-adjacent context length and the backward-adjacent context length.

[0117] Step 1063: Extract a first sequence segment having a length equal to a first context length from adjacent first sub-feature sequences.

[0118] After obtaining the first context length, the electronic device extracts segments from at least one of the forward-adjacent first sub-feature sequence and the backward-adjacent first sub-feature sequence, from the first sub-feature sequence adjacent to each first sub-feature sequence, according to the first context length, to obtain first sequence segments. Thus, the obtained first sequence segments may include at least one of the forward-adjacent sequence segments and the backward-adjacent sequence segments.

[0119] It should be noted that when the electronic device extracts fragments from the forward adjacent first sub-feature sequence, it takes the last audio frame feature in the forward adjacent first sub-feature sequence as the starting point, and then extracts fragments in the reverse order of the time sequence to obtain sequence fragments; when the electronic device extracts fragments from the backward adjacent first sub-feature sequence, it takes the first audio frame feature in the backward adjacent first sub-feature sequence as the starting point, and then extracts fragments in the forward order of the time sequence to obtain sequence fragments.

[0120] Step 1064: Determine context information of each first sub-feature sequence based on the first sequence segments.

[0121] After obtaining the first sequence fragments corresponding to each first sub-feature sequence, the electronic device can directly use the obtained first sequence fragments as the context information of the first sub-feature sequence, or it can combine the first sequence fragments and extract the context again for each first sub-feature sequence to obtain the context information of each first sub-feature sequence.

[0122] In some embodiments of the present application, Figure 5 Step 1064, i.e., determining the context information of each first sub-feature sequence based on the first sequence fragment, can be achieved by the following processing: concatenating each first sub-feature sequence and the corresponding first sequence fragment to obtain a third sub-feature sequence corresponding to each first sub-feature sequence; determining the minimum value between the first context length and the third preset length as the second context length, where the third preset length is the maximum length of the context information; and intercepting a second sequence fragment of the second context length from the third sub-feature sequence adjacent to the third sub-feature sequence corresponding to each first sub-feature sequence, and determining the second sequence fragment as the context information of each first sub-feature sequence.

[0123] The electronic device first concatenates each first sub-feature sequence and its corresponding first sequence fragment into a sequence, which is recorded as the third sub-feature sequence. The electronic device then compares the first context length with a third preset length, i.e., the maximum length of the context information, and extracts the minimum of the first and third preset lengths as the second context length. The electronic device then obtains, for each first sub-feature sequence, a third sub-feature sequence adjacent to the corresponding third sub-feature sequence. It then re-segments the adjacent third sub-feature sequence according to the second context length to obtain a second sequence fragment, and uses the resulting second sequence fragment as the context information for the first sub-feature sequence.

[0124] It should be noted that each first sub-feature sequence is spliced ​​with the first sequence fragment in order to strengthen the local time dependence for the first sub-feature sequence, so as to extract the neighboring context for the first sub-feature sequence, and then extract the context in the third sub-feature sequence, so that the context information can cover a longer distance, thereby realizing the acquisition of long-distance context. The minimum value of the first context length and the third preset length is used as the second context length to ensure that the length of the context information is within a reasonable range, so as to avoid the computational redundancy caused by excessively long context information, that is, to eliminate the parts irrelevant to the first sub-feature sequence through the maximum last limit.

[0125] It can be understood that in an embodiment of the present application, the electronic device can first perform close-range context extraction on the first sub-feature sequence, and then perform long-range context extraction on the first sub-feature sequence, thereby providing more comprehensive context information for the first sub-feature sequence through hierarchical context extraction of close and long distances.

[0126] Step 107: Fuse the context information into the corresponding first sub-feature sequence to obtain a second sub-feature sequence.

[0127] The electronic device can use the attention mechanism to integrate the contextual information into the corresponding first sub-feature sequence. For example, based on the contextual information, the importance weight of each speech frame feature in the first sub-feature sequence is adjusted to obtain the second sub-feature sequence. The electronic device can also integrate the contextual information into the corresponding first sub-feature sequence through splicing, bitwise superposition, etc.

[0128] The attention mechanism is a technology that simulates the selective focus on important information in the human cognitive process. It can dynamically assign weights to the speech frame features in the first sub-feature sequence, so that subsequent task processing can focus on key speech frame features. The attention mechanism in the embodiments of the present application can refer to a multi-head attention mechanism or a cross-attention mechanism.

[0129] It can be understood that in an embodiment of the present application, context information can be determined for each first sub-feature sequence based on the first sub-feature sequence adjacent to each first sub-feature sequence, and the first sub-feature sequence can be context-enhanced through the context information to obtain a second sub-feature sequence that can carry both the features of the first sub-feature sequence itself and the features of the context information, so that the second sub-feature sequence can provide richer information for subsequent task processing, thereby helping to improve the processing performance of task processing.

[0130] The following describes the implementation of audio signal processing by taking the generation of a transcribed text of an audio signal as an example.

[0131] In some embodiments of the present application, Figure 4 After step 106, that is, fusing the context information into the corresponding first sub-feature sequence to obtain the second sub-feature sequence, the method may further include the following processing: generating a corresponding text segment based on each second sub-feature sequence through the first model; and merging the multiple text segments into a transcribed text of the audio signal.

[0132] When generating a transcribed text from an audio signal, the electronic device first invokes the first model and provides multiple second sub-feature sequences to the first model. The first model then generates text based on each second sub-feature sequence, generating corresponding text segments. The electronic device then concatenates all the resulting text segments into a complete text, which becomes the transcribed text of the audio signal.

[0133] It should be noted that the first model can be a deep learning model with text generation capabilities, such as a Transformer model, or other Transformer-based models with text generation capabilities. The first model can acquire this capability through training. In an embodiment of the present application, the electronic device can first be pre-trained using unlabeled training data containing multiple languages, and then fine-tuned using labeled training data from different fields and tasks to obtain the first model. In this way, the obtained first model can have multi-language adaptability and multi-field and multi-task adaptability.

[0134] It can be understood that in an embodiment of the present application, the electronic device can first combine the obtained second sub-feature sequence to generate a text fragment, thereby improving the accuracy of the text fragment, and then splice the text fragment into a transcribed text. In this way, the transcribed text is generated based on multiple second sub-feature sequences with a more reasonable length and carrying rich contextual information, and its accuracy will be improved.

[0135] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0136] The embodiment of the present application is implemented in a scenario where a server (referred to as an electronic device) converts a voice signal (referred to as an audio signal) into text, referred to as a transcribed text.

[0137] Figure 6 This is a schematic diagram of the processing process of text conversion of speech signals provided in an embodiment of the present application. First, the server performs data preprocessing 6-1 on the speech signal, then calculates the dynamic segment size 6-2, then performs audio segmentation 6-3, and then performs attention mechanism enhancement 6-4. In the prediction stage, after the attention mechanism enhancement 6-4 is completed, subsequent text conversion processing will be performed. In the training stage, after the attention mechanism enhancement 6-4, the model needs to be pre-trained in multiple languages ​​6-5, and then domain adaptation 6-6 is performed, followed by multi-task learning 6-7, and finally model training and optimization 6-8, and then the obtained model (called the first model) is evaluated and deployed 6-9. Below, each step is described in detail.

[0138] Data preprocessing: The server preprocesses the speech signal to extract acoustic features and perform alignment. This process includes: reading the speech signal from the speech database. The audio format is Pulse Code Modulation (PCM) or Waveform Audio File Format (WAV), with a sampling rate of 16kHz and monophonic; feature extraction, using the Convolution-augmented Former (Conformer) model to extract features from the speech signal. The Conformer model combines the convolution module and the transformer module to effectively capture the local and global features of the speech signal. The feature extraction process can be expressed as Equation (3):

[0139] H=Conformer(X) (3)

[0140] Wherein, X represents the speech signal, and H represents the extracted feature sequence (called audio feature sequence).

[0141] Calculating dynamic segment size: In an embodiment of the present application, the dynamic segment size is calculated through an adaptive segment processing mechanism (ASP), which aims to optimize the segmentation processing process according to the length and content complexity (called complexity level) of the speech signal, so that the optimal segment size (called subsequence length) can be automatically selected for long audio and short audio, thereby maximizing computing efficiency and resource utilization while maintaining the integrity of context information.

[0142] The core idea of ​​ASP is to dynamically adjust the segment size by monitoring the length and content characteristics of the speech signal (such as speaking rate and background noise) in real time. For long audio, a larger segment size is selected to reduce computing resource consumption, while for short audio, a smaller segment size is selected to ensure the integrity of contextual information.

[0143] Assume that the characteristic sequence of the speech signal is X, and its length is T (in frames (called audio frames)). The goal of ASP is to segment X into multiple segments {S1, S2, ..., S N}, each segment size is C i , and its calculation process can be expressed as formula (4):

[0144]

[0145] Among them, C max is the maximum segment size (called the first preset length), which is suitable for use in long audio segments, C minis the minimum segment size (called the second preset length), which is suitable for use in short audio segmentation; α and β are dynamic adjustment coefficients (called the first sub-parameter and the second sub-parameter), which are used to control the proportional relationship between the segment size and the length of the feature sequence (called the total length). α and β can take a single value or multiple different values, which correspond to the sequence number of the segment. For example, when i is 1, it is one value, and when i is 2, it is another value.

[0146] In order to ensure the consistency of context information between segments, ASP introduces a context cache mechanism. i The output of (called the first sub-feature sequence) depends not only on the data of the current segment, but also on the cache information C of the previous segment prev and the cache information of the next segment (C next ). The size of the context cache (called context information) can be calculated by formula (5):

[0147]

[0148] Among them, γ and δ are the scaling coefficients of the context buffer (called the second scaling parameter), which are used to control the size of the forward and backward contexts.

[0149] Audio segmentation: The server segments the speech signal based on the obtained segment size. The following is a pseudo code description of adaptive segmentation:

[0150] Input: Feature sequence X (length T), C max 、C min , α and β, γ and δ, output: {S1, S2, …, S N}, that is, the segmented feature sequence, and {C prev ,C next}, which is the context cache.

[0151] Processing flow:

[0152] Calculate the length T of the feature sequence X;

[0153] Initialize the segment list {S1, S2, ..., S N} and context cache {C prev ,C next};

[0154] If T>C max , then C i =min(C max ,α·T), otherwise C i =max(C min ,β·T);

[0155] The feature sequence X of the speech signal is divided into segments of size C. i Divide into multiple segments {S1, S2, ..., S N};

[0156] For each segment S i , calculate the forward context cache C prev =γ·C i , calculate the backward context cache C next =δ·C i ;

[0157] Appends the context buffer to the segment data and returns the segment data as the context buffer.

[0158] In this way, the server can dynamically adjust the segment size according to the length and content complexity of the voice signal, ensuring optimal performance in both long and short audio tasks, and through the context caching mechanism, ensure the information consistency between ASP segments to avoid the loss of context information due to segmentation.

[0159] Attention mechanism enhancement: The server continues to perform dynamic context enhancement on the segmented data with the context cache attached (called the third feature sequence). Dynamic Context Enhancement (DCE) is a method used to optimize contextual information in segmented speech recognition, ensuring that the model fully utilizes global information when processing each segment, thereby improving the accuracy and coherence of subsequent conversions.

[0160] The core idea of ​​DCE is to dynamically allocate scalable context information for each segment, while adjusting the scope of the context according to the length and content complexity of the segment. This method can not only enhance the information transfer between segmented data, but also avoid performance degradation caused by insufficient or redundant context information.

[0161] Assume that the characteristic sequence of the speech signal is divided into multiple segments {S1, S2, ..., S N}, each segment also has context information, each segment S i The length is C i DCE dynamically allocates context information for each segment, including forward context C prev,i and backward context C next,i The dynamic adjustment formula of the context range is shown in formula (6):

[0162]

[0163] Among them, γ and δ are context scale coefficients, which are used to control the proportional relationship between context range and segment size, Γ max and Δm ax is the maximum range limit of the context, used to avoid excessive context cache redundancy.

[0164] For each segment, DCE calculates the dynamic range of its context buffer and embeds it into the feature representation of the current segment. The embedding of the context buffer can be achieved through the attention mechanism, which can be expressed as Equation (7):

[0165] E i =Attention(S i ,C prev,i ,C next,i ) (7)

[0166] Among them, E i is the enhanced segment feature representation (called the second sub-feature sequence), and Attention is an attention mechanism for dynamic weighted context caching.

[0167] The following is a pseudo-code description of dynamic context enhancement:

[0168] Input: {S1,S2,…,S N}, γ and δ are context scaling coefficients, Γ max and Δ m ax;

[0169] Output: {E1,E2,…,E N}, represents the enhanced segmentation features;

[0170] Processing flow:

[0171] Initialize the enhanced segment features {E1,E2,…,E N};

[0172] For each segment S i , calculate the forward context range C prev,i =min(γ·C i ,Γ max ) and the backward context scope C next,i =min(δ·C i ,Δ max );

[0173] Extract forward context information C prev,i and backward context information C next,i , where if i=1 (i.e. the first segment), then C prev,i = 0, if i = N (i.e. the last segment), then C next,i =0;

[0174] Use attention mechanism to embed dynamic context cache into current segment;

[0175] Returns the enhanced segment features {E1,E2,…,E N}.

[0176] In this way, the context range can be dynamically adjusted according to the length and content complexity of the segment, ensuring that each segment can fully utilize global information; and the context cache is embedded into the feature segment through the attention mechanism, which significantly enhances the information transfer between segments, improves the performance of speech recognition, and avoids the waste of computing resources caused by excessive redundancy of the context cache.

[0177] Multi-language pre-training: The model used for feature extraction and text generation in the embodiment of this application is jointly trained on a dataset of multiple voices, so that the model can learn the common features and differential features of different voices. In this process, the server first needs to build a multi-language dataset D multi-lingual , contains speech data of multiple voices, and pre-training is performed on this dataset. The loss function during pre-training can be expressed as formula (8):

[0178]

[0179] Where L represents the speech set, A dataset representing language L.

[0180] Domain Adaptation: Automatically adjust model parameters based on the domain characteristics of the language signal (such as background noise, speaking style, etc.) to ensure the robustness of the model in different domains. For each domain D, a domain-specific parameter θ is introduced D , in the multi-domain dataset D multi-domain Fine-tune the loss function and optimize it. The loss function can be expressed as formula (9):

[0181]

[0182] in, Represents a collection of fields.

[0183] Multi-task learning: Simultaneously optimize the speech recognition task and other related tasks (such as speaker recognition, sentiment analysis, etc.). Its objective function can be expressed as formula (10):

[0184]

[0185] Where λ is the equilibrium parameter, is the loss function for the speech recognition task, is the loss function for the auxiliary task.

[0186] The following is a pseudo-code description of the process:

[0187] Input: Feature sequence X of speech signal, voice label L of speech signal, domain label D of speech signal, multilingual dataset D multi-lingual , multi-domain dataset D multi-domain ,λ is the balance parameter for multi-task learning;

[0188] Output: optimized model parameters;

[0189] Processing flow:

[0190] Multilingual pre-training: In the multilingual dataset D multi-lingual The pre-training process can be expressed as formula (11):

[0191]

[0192] Domain Adaptation: For each domain D, we introduce domain-specific parameters θ D , to train the model, the process can be expressed as formula (12):

[0193]

[0194] Multi-task learning: Optimize the speech recognition task and other related tasks simultaneously. The process can be expressed as formula (13):

[0195]

[0196] Among them, return θ final as the final model parameters.

[0197] In this way, through multilingual pre-training, the model can learn the common features of different languages ​​and improve its performance in multilingual tasks. Through domain adaptation, the model can automatically adjust parameters according to the domain characteristics of the language signal to ensure the robustness of the model in different domains. Through multi-task learning, the model can handle multiple tasks simultaneously, enhancing its versatility and scalability.

[0198] Evaluate and deploy the model: Apply the trained model to real-world scenarios for speech recognition. The trained model can be deployed to target devices, such as servers or embedded devices. These devices use the model to process input speech signals in real time, extract features, and input them into the model for recognition. The recognition results are then output as text for subsequent use.

[0199] It is understandable that in the embodiments of the present application, when user information, such as user voice and other related data, is involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent must be obtained, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.

[0200] The structure of the electronic device provided in the embodiments of the present application is described below.

[0201] See also Figure 7 , Figure 7 is a structural diagram of an electronic device provided in an embodiment of the present application, Figure 7 The electronic device 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 7 Various buses are labeled as bus system 440 .

[0202] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0203] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0204] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0205] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0206] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0207] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0208] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);

[0209] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0210] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.

[0211] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 7 An audio processing device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a feature extraction module 4551, a sequence generation module 4552, a length determination module 4553, a sequence processing module 4554, and a context processing module 4555. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0212] In other embodiments, the audio processing device provided in the embodiments of the present application can be implemented in hardware. As an example, the audio processing device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the audio processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0213] In some embodiments, the electronic device can implement the audio processing method provided in the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be microprogram-level commands, machine instructions or software instructions. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a voice transcription APP; it can also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to a browser environment to run. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0214] The following continues to describe the exemplary structure of the audio processing device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 7 As shown, the software modules stored in the audio processing device 455 of the memory 450 may include:

[0215] A feature extraction module 4551 is used to extract features from audio frames of an audio signal to obtain audio frame features;

[0216] A sequence generation module 4552 is configured to generate an audio feature sequence using the audio frame features;

[0217] a length determining module 4553, configured to determine a subsequence length of the audio feature sequence based on a total length of the audio feature sequence, wherein the total length represents the number of audio frame features included in the audio feature sequence;

[0218] The sequence processing module 4554 is configured to divide the audio feature sequence into a plurality of first sub-feature sequences based on the sub-sequence lengths, and process the audio signal according to the first sub-feature sequences.

[0219] In the above solution, the length determination module 4553 is further used to determine a first sub-length determination method that matches the audio feature sequence based on the total length of the audio feature sequence; and determine the sub-sequence length through the first sub-length determination method.

[0220] In the above scheme, the length determination module 4553 is further used to determine the length interval to which the total length belongs from a preset first length interval and a preset second length interval, where the left boundary of the first length interval is greater than the right boundary of the second length interval; and determine the second subsequence length determination method corresponding to the length interval to which the total length belongs from multiple second subsequence length determination methods as the first subsequence length determination method.

[0221] In the above scheme, the length determination module 4553 is also used to obtain a first proportional parameter, which includes: a first sub-parameter and a second sub-parameter, and the first sub-parameter is greater than the second sub-parameter; in response to the length interval to which the total length belongs being the first length interval, the total length is adjusted using the first sub-parameter, and the minimum value between the adjusted length and the first preset length is determined as the subsequence length, wherein the first preset length is the maximum value of the subsequence length; in response to the length interval to which the total length belongs being the second length interval, the total length is adjusted using the second sub-parameter, and the maximum value between the adjusted length and the second preset length is determined as the subsequence length, and the second preset length is the minimum value of the subsequence length.

[0222] In the above solution, the length determination module 4553 is further configured to obtain the complexity of the audio signal; and use a preset parameter corresponding to the complexity among a plurality of preset parameters as the first proportional parameter.

[0223] In the above scheme, the length determination module 4553 is also used to perform speech rate analysis on the audio signal to obtain the speech rate information of the audio signal; perform noise intensity analysis on the audio signal to obtain the noise intensity information of the audio signal; and calculate the complexity of the audio signal based on the speech rate information and the noise intensity information.

[0224] In the above scheme, the audio processing device 455 also includes: a context processing module 4555, which is used to determine the context information of each first sub-feature sequence based on the first sub-feature sequence adjacent to each first sub-feature sequence; and fuse the context information into the corresponding first sub-feature sequence to obtain a second sub-feature sequence.

[0225] In the above scheme, the context processing module 4555 is also used to obtain a second proportional parameter, which is a parameter that controls the proportional relationship between the length of the context information and the subsequence length; the product of the second proportional parameter and the subsequence length is determined as the first context length; a first sequence segment with a length of the first context length is cut off from the adjacent first sub-feature sequence; and the context information of each first sub-feature sequence is determined based on the first sequence segment.

[0226] In the above scheme, the context processing module 4555 is further used to splice each first sub-feature sequence and the corresponding first sequence fragment to obtain a third sub-feature sequence corresponding to each first sub-feature sequence; determine the minimum value between the first context length and the third preset length as the second context length, and the third preset length is the maximum length of the context information; from the third sub-feature sequence adjacent to the third sub-feature sequence of each first sub-feature sequence, cut off a second sequence fragment with a length of the second context, and determine the second sequence fragment as the context information of each first sub-feature sequence.

[0227] In the above solution, the audio processing device 455 further includes: a text generation module 4556, which is used to generate a corresponding text segment based on each second sub-feature sequence through a first model; and merge multiple text segments into a transcribed text of the audio signal.

[0228] The present invention provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the audio processing method described in the present invention.

[0229] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the audio processing method provided in the embodiment of the present application, for example, Figure 2 The audio processing method shown.

[0230] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0231] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0232] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0233] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0234] To sum up, through the embodiments of the present application, the audio signal is first segmented, and the audio frame features extracted from the audio frame are used to construct an audio feature sequence of the speech signal, and then the subsequence length that is adapted to the total length of the audio feature sequence is determined, so that a subsequence length that is adapted to audio signals of different lengths can be obtained, and the audio feature sequence is divided based on the subsequence length to obtain a sub-feature sequence that matches the length of the audio signal, that is, the length of the obtained sub-feature sequence is more reasonable, thereby improving the processing performance of audio processing.

[0235] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. An audio processing method, characterized in that: The method comprises: Perform feature extraction on the audio frame of the audio signal to obtain audio frame features; generating an audio feature sequence using the audio frame features; Determining a subsequence length of the audio feature sequence based on the total length of the audio feature sequence; Based on the subsequence length, the audio feature sequence is divided into a plurality of first sub-feature sequences, and the audio signal is processed according to the plurality of first sub-feature sequences.

2. The method according to claim 1, characterized in that The determining, based on the total length of the audio feature sequence, the subsequence length of the audio feature sequence includes: Determining a first sub-length determination method that matches the audio feature sequence based on the total length of the audio feature sequence; The subsequence length is determined by using the first subsequence length determination method.

3. The method according to claim 2, characterized in that The determining of the first sub-length determination method for the audio feature sequence based on the total length of the audio feature sequence includes: Determine, from a preset first length interval and a preset second length interval, a length interval to which the total length belongs, wherein a left boundary of the first length interval is greater than a right boundary of the second length interval; A second subsequence length determination method corresponding to the length interval to which the total length belongs, among the plurality of second subsequence length determination methods, is determined as the first subsequence length determination method.

4. The method according to claim 2, characterized in that The determining the subsequence length by using the first subsequence length determination method includes: Obtaining a first scale parameter, the first scale parameter including: a first sub-parameter and a second sub-parameter, the first sub-parameter being greater than the second sub-parameter; In response to the length interval to which the total length belongs being a first length interval, adjusting the total length using the first sub-parameter, and determining the minimum value between the adjusted length and a first preset length as the subsequence length, wherein the first preset length is the maximum value of the subsequence length; In response to the length interval to which the total length belongs being the second length interval, the total length is adjusted using the second sub-parameter, and the maximum value of the adjusted length and the second preset length is determined as the subsequence length, and the second preset length is the minimum value of the subsequence length.

5. The method according to claim 4, characterized in that The obtaining of the first ratio parameter includes: Obtaining the complexity of the audio signal; A preset parameter corresponding to the complexity level among a plurality of preset parameters is used as the first scale parameter.

6. The method according to claim 5, characterized in that The obtaining of the complexity of the audio signal includes: Performing speech rate analysis on the audio signal to obtain speech rate information of the audio signal; Performing noise intensity analysis on the audio signal to obtain noise intensity information of the audio signal; The complexity of the audio signal is calculated based on the speech rate information and the noise intensity information.

7. The method according to any one of claims 1 to 6, characterized in that After dividing the audio feature sequence into a plurality of first sub-feature sequences based on the sub-sequence length, the method further includes: determining context information of each of the first sub-feature sequences based on a first sub-feature sequence adjacent to each of the first sub-feature sequences; The context information is integrated into the corresponding first sub-feature sequence to obtain a second sub-feature sequence.

8. The method according to claim 7, characterized in that The determining, based on the first sub-feature sequence adjacent to each first sub-feature sequence, context information of each first sub-feature sequence includes: Obtaining a second ratio parameter, where the second ratio parameter is a parameter that controls a ratio between a length of the context information and a length of the subsequence; Determine the first context length as the product of the second scale parameter and the subsequence length; Extracting a first sequence segment having a length equal to the first context length from adjacent first sub-feature sequences; The context information of each of the first sub-feature sequences is determined based on the first sequence segments.

9. The method according to claim 8, characterized in that The determining, based on the first sequence segments, the context information of each first sub-feature sequence includes: splicing each of the first sub-feature sequences with the corresponding first sequence fragment to obtain a third sub-feature sequence corresponding to each of the first sub-feature sequences; determining a minimum value between the first context length and a third preset length as a second context length, wherein the third preset length is a maximum length of the context information; A second sequence segment having a length of the second context is cut off from the third sub-feature sequence adjacent to the third sub-feature sequence of each first sub-feature sequence, and the second sequence segment is determined as the context information of each first sub-feature sequence.

10. The method according to claim 7, characterized in that After fusing the context information into the corresponding first sub-feature sequence to obtain a second sub-feature sequence, the method further includes: Generate a corresponding text segment based on each of the second sub-feature sequences using the first model; The plurality of text segments are combined into a transcript of the audio signal.

11. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; A processor, configured to implement the method according to any one of claims 1 to 10 when executing computer-executable instructions or computer programs stored in the memory.

12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the method according to any one of claims 1 to 10 is implemented.

13. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the method according to any one of claims 1 to 10 is implemented.