A multi-thread based audio format conversion method and system

By using a multi-threaded audio format conversion method and leveraging a neural network learning model and a PESQ model to adjust output parameters, the problem of audio file quality degradation after conversion is solved, achieving efficient and high-quality audio format conversion.

CN117334205BActive Publication Date: 2026-05-08JIANGXIA INFORMATION TECH (HUIZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGXIA INFORMATION TECH (HUIZHOU) CO LTD
Filing Date
2023-09-19
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing technologies, audio files suffer from quality degradation after format conversion, mainly because selecting fixed output parameters cannot adapt to changes in audio specification parameters within the audio file.

Method used

A multi-threaded audio format conversion method is adopted. The audio segments are selectively extracted to form a conversion reference through a neural network learning model. The PESQ model is used for quality certification scoring. The output parameters are adjusted to meet the mathematical expectation, and learning rules are generated to adjust the audio format conversion process in real time.

Benefits of technology

It improves the quality and efficiency of audio format conversion, avoids multiple trials and adjustments, and achieves high-quality audio file format conversion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117334205B_ABST
    Figure CN117334205B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-threaded audio format conversion method and system, import audio file and select conversion reference body;Based on the initial audio format and target audio format of conversion reference body, audio file is classified, and neural network learning model is selected, and conversion reference body is segmented format conversion according to fixed output parameter;The segmented quality evaluation result of conversion reference body is obtained, the neural network learning model is adjusted based on quality evaluation result, until segmented quality evaluation result meets mathematical expectation;Start the audio format conversion of initial audio file;Audio file is segmented, and the audio segment of current segment is converted in format in sequence according to neural network learning real-time model, and the output parameter is determined in advance according to the audio specification parameter of next segment;The learning rule about output parameter is generated only before the format conversion of audio file in the application, and the output parameter is segmented and adjusted according to the learning rule, so as to improve the whole conversion efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of audio format conversion, and specifically to an audio format conversion method and system based on multithreading. Background Technology

[0002] Audio format conversion refers to the process of converting one audio format to another, usually to enable different devices or software to play, edit, or process audio files. The principle of audio format conversion is to decode audio files of different formats into an intermediate format, and then re-encode them into the target format.

[0003] Common audio formats have different characteristics, such as compression ratio, sound quality, and file size. Therefore, choosing the appropriate audio format is very important; different formats can be selected based on different needs.

[0004] Most existing audio format conversion methods select fixed output parameters for the audio file. However, since the audio specifications in the audio file are not constant and completely the same, but change according to different time periods, the audio file may suffer quality loss after conversion. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-threaded audio format conversion method and system to solve the technical problem in the prior art where selecting fixed output parameters for the audio file leads to quality degradation after audio format conversion.

[0006] To solve the above-mentioned technical problems, the present invention specifically provides the following technical solution:

[0007] A multi-threaded audio format conversion method includes the following steps:

[0008] Step 100: Import audio files and identify the initial audio format of the audio files; classify the different audio files based on the initial audio format and the target audio format of each audio file.

[0009] Step 200: Select a neural network learning model based on the initial audio format and target audio format of the conversion reference body, selectively extract the imported audio file to form a conversion reference body, the neural network learning model generates output parameters according to common audio specification parameters, and the conversion reference body performs audio format conversion sequentially according to the output parameters;

[0010] Step 300: Perform quality certification scoring on the conversion reference body that has been converted to the target audio format. Based on the scoring results and the audio specification parameters carried by the conversion reference body, adjust the neural network learning model, establish the correlation between the audio specification parameters and the neural network learning model, and generate a real-time neural network learning model that corresponds one-to-one with different groups of audio specification parameters, so as to adjust the output parameters in real time until the scoring results meet expectations.

[0011] Step 400: Start the audio format conversion of the initial audio file, obtain the audio specification parameters of each largest audio segment, and the neural network learning real-time model adjusts the output parameters in advance according to the audio specification parameters of the next segment, until each audio file is converted to the corresponding target format.

[0012] As a preferred embodiment of the present invention, in step 200, the neural network learning model is pre-saved, and the neural network learning model is initially selected based on the conversion relationship between the initial audio format and the target audio format. The neural network learning model adjusts the output parameters uniformly based on the general audio specification parameters corresponding to the initial audio format.

[0013] As a preferred embodiment of the present invention, in step 200, the step of selectively extracting the imported audio file to form a conversion reference body is as follows:

[0014] The imported audio file is parsed, and the audio file is sequentially cut into multiple audio segments to obtain the audio specification parameters of each audio segment;

[0015] Select audio segments with different combinations of the audio specification parameters to form a conversion reference body, wherein the conversion reference body is composed of audio segments of different durations in the audio file.

[0016] As a preferred embodiment of the present invention, the audio specification parameters include audio, sampling frequency, sampling bit depth, number of channels, and bit rate;

[0017] The output parameters include the number of output channels and the encoding format;

[0018] The audio formats include MP3, WAV, AAC, FLAC, and OGG.

[0019] As a preferred embodiment of the present invention, the specific implementation steps for performing quality certification scoring on the converted reference body that has been converted to the target format in step 300 are as follows:

[0020] The conversion reference body is re-decomposed into multiple audio segments according to the combination of audio specification parameters;

[0021] The neural network learning model sequentially adjusts the output parameters for each audio segment, while simultaneously acquiring the audio specification parameters for each audio segment;

[0022] Each audio segment that has been converted to the target audio format is scored for quality certification, resulting in multiple scores for the audio segments.

[0023] First, select the audio specification parameters of the audio segments with large differences in the scoring results to form the first dataset. Then, import the first dataset into the neural network learning model for multiple training sessions to adjust the output parameters until the scoring results of the audio segments with low scoring results after being converted to the target audio format meet the mathematical expectation.

[0024] Then, the audio specification parameters of the audio segments with small differences in the scoring results are selected to form a second dataset. The second dataset is then imported into the neural network learning model for testing. The stability of the adjusted output parameters is verified by using the scoring results after the audio segments are reconverted to the target audio format.

[0025] As a preferred embodiment of the present invention, an association relationship is established between the audio specification parameters of the audio segment and the real-time learning model of the neural network to form a learning rule, and the output parameter is adjusted based on the learning rule and the audio specification parameters of the audio segment until the scoring result of each audio segment in the conversion reference body meets the mathematical expectation.

[0026] The input value of the real-time neural network learning model is the audio specification parameter, and the output value of the real-time neural network learning model is the output parameter.

[0027] As a preferred embodiment of the present invention, the implementation model for quality certification scoring of the conversion reference body that has been converted to the target format is as follows: using PESQ to establish a simulated human ear auditory model to predict the listener's subjective score of the audio quality of the conversion reference body that has been converted to the target audio format.

[0028] The rating ranges from -0.5 to 4.5, with higher scores indicating higher audio quality. It is used to evaluate the audio quality after audio format conversion.

[0029] As a preferred embodiment of the present invention, in steps 100 to 300, the initial audio file pauses audio format conversion, and an audio specification conversion test is performed on the conversion reference body formed by the combination.

[0030] In step 400, the learning rules for adjusting the output parameters based on the different audio specification parameters obtained are used to initiate the audio format conversion of the initial audio file.

[0031] As a preferred embodiment of the present invention, the specific steps for initiating the audio format conversion of the initial audio file are as follows:

[0032] The audio specification parameters of the audio of a set duration are obtained according to the capture frequency, and the audio of the set duration is divided into different audio segments according to the standard of matching different real-time learning models of the neural network.

[0033] The neural network is used to learn a real-time model that adjusts the output parameters based on the audio specification parameters of the current audio segment;

[0034] Based on the audio specification parameters and learning rules of the next audio segment, the adjustment target of the real-time learning model of the neural network is determined in advance, so as to determine the adjustment mode of the output parameters corresponding to the next audio segment in advance, and to perform timely audio format conversion on the next audio segment.

[0035] To address the aforementioned technical problems, the present invention further provides the following technical solution: an audio format conversion system based on a multi-threaded audio format conversion method, comprising:

[0036] The initial audio classification unit is used to classify the initial audio file according to the initial audio format and the target conversion format of the received initial audio file;

[0037] An initial audio extraction unit is used to mark the audio specification parameters of the initial audio file and, based on the differences in the audio specification parameters of the initial audio file, selectively extracts a segment of audio from the initial audio file as a conversion reference.

[0038] The neural network learning model corresponds one-to-one with audio files of different categories, and the neural network learning model is selected according to the initial audio format and target conversion format corresponding to each conversion reference body. Each neural network learning model sets the general output parameters of the conversion reference body based on the initial audio format and target conversion format.

[0039] The quality certification unit performs quality certification scoring on the conversion reference body that has been converted to the target audio format;

[0040] The neural network learns a real-time model and adjusts the real-time output parameters of the conversion reference body according to the output results of the quality certification unit until the final output result of the conversion reference body after passing through the quality certification unit reaches the mathematical expectation. The audio specification parameters and output parameters in the neural network learns a real-time model to form a learning rule.

[0041] All audio conversion units perform full formal audio format conversion based on the initial audio file, determine a set of audio specification parameters corresponding to the output parameters according to the learning rules, and adjust the output parameters of each audio segment based on the audio specification parameters of the initial audio file and the learning rules.

[0042] Compared with the prior art, the present invention has the following advantages:

[0043] This invention selects a neural network learning model for the initial and target audio formats of an audio file. Then, based on the audio specifications of the complete audio file, it selects a conversion reference body composed of audio factors with distinct characteristics. Different audio factors within the conversion reference body are converted using the output parameters generated by the neural network learning model. The conversion results for different audio factors within the conversion reference body are quality-assessed, and the quality scores are used to train the neural network learning model until a real-time neural network learning model for different audio factors is obtained, generating learning rules. Finally, the entire audio file is converted sequentially. Different real-time neural network learning models are selected for different audio specifications, thus continuously converting the entire audio file segment by segment, improving the overall quality of audio format conversion. This eliminates the need for repeated experimentation and adjustment of output parameters across the entire audio file; the neural network learning model is only trained before the format conversion, allowing for high-quality format conversion directly across the entire audio file, thereby improving overall conversion efficiency. Attached Figure Description

[0044] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.

[0045] Figure 1 This is a flowchart illustrating the audio format conversion method provided in an embodiment of the present invention;

[0046] Figure 2 This is a structural block diagram of an audio format conversion system provided in an embodiment of the present invention.

[0047] The labels in the diagram represent the following:

[0048] 1-Initial audio classification unit; 2-Initial audio extraction unit; 3-Neural network learning model; 4-Quality certification unit; 5-Real-time neural network learning model; 6-Complete audio conversion unit. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] like Figure 1 As shown, this invention provides a multi-threaded audio format conversion method, including the following steps:

[0051] Step 100: Import audio files and identify their initial audio format. Based on the initial audio format and target audio format of each audio file, classify the different audio files.

[0052] Common audio formats include MP3, WAV, AAC, FLAC, and OGG. These formats have different characteristics, such as compression ratio, sound quality, and file size. Therefore, choosing the appropriate audio format is very important; different formats can be selected based on different needs.

[0053] Step 200: Select a neural network learning model based on the initial audio format and target audio format of the conversion reference body, selectively extract audio files from the imported audio files to form a conversion reference body, generate output parameters based on common audio specification parameters, and convert audio formats sequentially according to the output parameters of the conversion reference body.

[0054] In this embodiment, based on common audio formats, at least 20 audio format conversion requirements can be generated. The target audio format directly determines the selection of output parameters. Therefore, this embodiment classifies different audio files according to the initial audio format and the target audio format of the audio file, and selects a neural network learning model based on the conversion relationship between the initial audio format and the target audio format.

[0055] It should be noted that in step 200, the neural network learning model is a pre-saved model. The neural network learning model is initially selected based on the conversion relationship between the initial audio format and the target audio format. The neural network learning model adjusts the output parameters uniformly based on the general audio specification parameters corresponding to the initial audio format.

[0056] When using a neural network learning model to adjust output parameters based on audio specification parameters, the default is that the audio specifications of the audio files are universal and the same. Therefore, the output parameters can be adjusted uniformly based on the universal audio specification parameters.

[0057] Additional information: Audio specifications include audio value, sampling frequency, sampling bit depth, number of channels, and bit rate. Output parameters include the number of output channels and the encoding format.

[0058] Audio refers to sound waves with frequencies between 20Hz and 20kHz that can be heard by the human ear.

[0059] Sampling frequency refers to the number of sound samples acquired per second. Sound is actually an energy wave, and therefore also has the characteristics of frequency and amplitude. Frequency corresponds to the time axis, and amplitude corresponds to the level axis. Waves are infinitely smooth, and a string can be seen as composed of countless points. Since storage space is relatively limited, the points of the string must be sampled during the digital encoding process.

[0060] Sampling bit depth, also called sample size or quantization bit depth, is a parameter used to measure changes in sound fluctuations. It's essentially the resolution of a sound card, or the sound card's ability to process sound with high precision. The higher the value, the higher the resolution, and the more realistic the recorded and played-back sound.

[0061] The number of channels refers to the number of channels in the sound. Common types include mono and stereo (two-channel), and now it has evolved to quad surround (four-channel) and 5.1 channels.

[0062] Bit rate, also known as bit rate, refers to the amount of data played per second in music playback, and is expressed in bits.

[0063] The audio specifications in an audio file are not constant and completely identical, but change according to different time periods. Therefore, in order to adjust the output parameters for different audio specifications, this implementation selects different audio segments to form a conversion reference body based on the differences in audio specifications. According to the changes in audio specifications in the conversion reference body, the neural network learning model is adjusted to change, so as to adjust the output parameters in accordance with the changes in audio specifications.

[0064] Therefore, in step 200, the implementation steps for selectively extracting the imported audio file to form a conversion reference body are as follows:

[0065] The imported audio file is parsed and then cut into multiple audio segments, and the audio specification parameters of each audio segment are obtained.

[0066] Select audio segments with different combination forms and audio specification parameters to form a conversion reference body, wherein the conversion reference body is composed of audio segments of different durations in the audio file.

[0067] Because the audio specification parameters in the conversion reference are different and vary considerably, when using a general neural network learning model with fixed output parameters, the audio quality of different audio segments after conversion will differ. Therefore, it is necessary to adjust the output parameters of the neural network learning model according to the different audio specification parameters to ensure that the audio quality of each audio segment meets the requirements.

[0068] Therefore, as a preferred embodiment, the conversion reference body in this embodiment is not extracted sequentially from the audio file according to the time order. Instead, the audio file is first extracted sequentially to form multiple audio segments, and the audio segments with larger differences in audio specification parameters are selected as the conversion reference body. This makes it easier to train the neural network learning model, determine the correspondence between the audio specification parameters and the output parameter adjustment, adjust the neural network learning model for audio segments with different audio specification parameters, and generate a real-time neural network learning model that corresponds one-to-one with different groups of audio specification parameters. The specific implementation method is as shown in step 300.

[0069] Step 300: Perform quality certification scoring on the conversion reference body that has been converted to the target audio format. Based on the scoring results and the audio specification parameters carried by the conversion reference body, adjust the neural network learning model, establish the correlation between the audio specification parameters and the neural network learning model, and generate a real-time neural network learning model that corresponds one-to-one with different groups of audio specification parameters to adjust the output parameters in real time until the scoring results meet expectations.

[0070] In step 300, the specific steps for performing quality certification scoring on the converted reference body that has been converted to the target format are as follows:

[0071] (1) The conversion reference body is re-disassembled into multiple audio segments according to the combination of audio specification parameters.

[0072] (2) The neural network learning model adjusts the output parameters for each audio segment in turn, and obtains the audio specification parameters of each audio segment.

[0073] (3) Perform quality certification scoring on each audio segment that has been converted to the target audio format to obtain the scoring results of multiple audio segments.

[0074] The implementation model for quality certification scoring of the converted reference body that has been converted to the target format is as follows: a simulated human ear auditory model is established using PESQ to predict the listener's subjective score for the audio quality of the converted reference body that has been converted to the target audio format.

[0075] The rating ranges from -0.5 to 4.5, with higher scores indicating higher audio quality. It is used to evaluate the audio quality after audio format conversion.

[0076] (3) First, select the audio specification parameters of the audio segments with large differences in the scoring results to form the first dataset. Then, import the first dataset into the neural network learning model for multiple training sessions to adjust the output parameters until the scoring results of the audio segments with low scoring results after conversion to the target audio format meet the mathematical expectation.

[0077] (4) Select the audio specification parameters of the audio segments with small differences in the scoring results to form a second dataset. Import the second dataset into the neural network learning model for testing. Use the scoring results after the audio segments are converted back to the target audio format to verify the stability of the adjusted output parameters.

[0078] This involves establishing a correlation between the audio specification parameters of audio segments and the real-time learning model of the neural network, forming learning rules, and adjusting the output parameters based on the learning rules and the audio specification parameters of the audio segments until the scoring results of each audio segment in the transformation reference body meet the mathematical expectation.

[0079] The input value of the real-time neural network learning model is the audio specification parameter, and the output value of the real-time neural network learning model is the output parameter.

[0080] This implementation first selects a neural network learning model based on the initial audio format of the audio file and the target audio format to be converted. However, since the neural network learning model generates matching output parameters for general audio specification parameters (i.e., fixed audio rule parameters or audio specification parameters within a certain range of data), in actual applications, even for the same audio file, the audio specification parameters of different segments are not exactly the same along the duration. Therefore, the matching output parameters will damage the audio quality of that segment.

[0081] To address this issue, this implementation method uses a quality scoring approach to determine the quality of different audio factors in the conversion reference body after audio format conversion, which is used to train the neural network learning model. Based on whether the quality meets the mathematical expectation, the neural network learning model is continuously adjusted to continuously adjust the output parameters until the audio quality scores of all audio factors in the conversion reference body after conversion to the target format meet the mathematical expectation. At this point, the generated learning rules are more comprehensive and accurate.

[0082] In addition, as a preferred embodiment, in steps 100 to 300, the initial audio file pauses audio format conversion, and an audio specification conversion test is performed on the conversion reference body formed by the combination.

[0083] In step 400, the learning rules for adjusting the output parameters based on the different audio specification parameters obtained are used to start the audio format conversion of the initial audio file.

[0084] To avoid the problem of excessive datasets for adjusting the neural network learning model, which leads to long processing times and low audio format conversion efficiency, this implementation method does not adjust the neural network learning model based on the audio specification parameters of an entire audio file. Instead, it selects audio factor combinations corresponding to audio specification parameters with significant feature differences based on the characteristic differences of the audio specification parameters, forming a conversion reference body with more obvious audio specification parameters. The neural network learning model is trained using the audio specification parameters of the conversion reference body to determine the correspondence between the audio specification parameters and the output parameter adjustments. This results in a large training breadth and facilitates the determination of the learning rules for the real-time neural network learning model.

[0085] Step 400: Start the audio format conversion of the initial audio file, obtain the audio specification parameters of each largest audio segment, and the neural network learning real-time model adjusts the output parameters in advance according to the audio specification parameters of the next segment until each audio file is converted to the corresponding target format.

[0086] The specific steps to initiate audio format conversion of the initial audio file are as follows:

[0087] The audio specification parameters of the audio of a set duration are obtained according to the capture frequency, and the audio of the set duration is divided into different audio segments according to the standard of matching different neural network learning real-time models.

[0088] The neural network learns a real-time model that adjusts output parameters based on the audio specification parameters of the current audio segment;

[0089] Based on the audio specifications and learning rules of the next audio segment, the adjustment target of the real-time model of the neural network is determined in advance, so as to determine the adjustment mode of the output parameters corresponding to the next audio segment in advance, and to perform timely audio format conversion for the next audio segment.

[0090] In other words, in this embodiment, the audio specification parameters of the initial audio file are obtained in chronological order. Due to the limitations of computer computing power, the audio specification parameters of the initial audio file obtained in this instance can only be obtained for audio segments within a certain period of time. The output parameters are adjusted by using a neural network generated for the audio segments within the current time period to learn a real-time model and convert the audio segments into audio formats.

[0091] To ensure timely adjustment of output parameters for the next audio segment, when converting the audio format of the current audio segment, the output parameters of the next audio segment are determined based on the pre-acquired audio specification parameters of the next audio segment and the matching learning rules. After the audio segments in the current time period are converted, the output parameters are directly adjusted to the data that has been acquired. This achieves multi-threaded audio format conversion, which specifically includes three threads: format conversion of the current audio segment, conversion position monitoring, and output parameter adjustment for the next audio segment.

[0092] In addition, as an innovation of this invention, the learning rules obtained from training the neural network learning model are used to adjust the neural network learning model to obtain a real-time neural network learning model that matches different audio specification parameters. This model can directly convert audio segments with the same conversion requirements in the later stages according to step 400, thereby maximizing the efficiency of format conversion.

[0093] Based on the above-described multi-threaded audio format conversion method, this embodiment also provides an audio format conversion system, such as... Figure 2 As shown, it includes: initial audio classification unit 1, initial audio extraction unit 2, neural network learning model 3, quality certification unit 4, neural network learning real-time model 5, and full audio conversion unit 6.

[0094] The initial audio classification unit 1 is used to classify the initial audio file according to the initial audio format and the target conversion format of the received initial audio file;

[0095] The initial audio extraction unit 2 is used to mark the audio specification parameters of the initial audio file, and selectively extracts a segment of audio from the initial audio file as a conversion reference based on the differences in the audio specification parameters of the initial audio file.

[0096] The neural network learning model 3 corresponds one-to-one with audio files of different categories, and the neural network learning model 3 is selected according to the initial audio format and target conversion format corresponding to each conversion reference. Each neural network learning model 3 sets the general output parameters of the conversion reference based on the initial audio format and target conversion format.

[0097] Quality certification unit 4 performs quality certification scoring on the conversion reference that has been converted to the target audio format;

[0098] The real-time neural network learning model 5 adjusts the real-time output parameters of the conversion reference body according to the output results of the quality certification unit 4 until the final output result of the conversion reference body after passing through the quality certification unit 4 reaches the mathematical expectation, and the audio specification parameters and output parameters in the real-time neural network learning model 5 form learning rules.

[0099] All audio conversion units 6 perform full formal audio format conversion based on the initial audio file, determine the audio specification parameters corresponding to a set of output parameters according to the learning rules, and adjust the output parameters of each audio segment based on the audio specification parameters of the initial audio file and the learning rules.

[0100] This implementation selects a neural network learning model for the initial audio format and the target audio format of the audio file. Then, for the audio specifications of the complete audio file, it selects a conversion reference body with distinct audio factor combinations. Different audio factors in the conversion reference body are converted into audio formats using the output parameters generated by the neural network learning model. The conversion results of different audio factors in the conversion reference body are quality-evaluated, and the quality scores are used to train the neural network learning model until a real-time neural network learning model for different audio factors is obtained, generating learning rules. Finally, the entire audio file is converted into formats sequentially. Different real-time neural network learning models are selected for different audio specifications, thereby continuously converting the entire audio file into segments, improving the overall quality of audio format conversion. It eliminates the need for repeated experimentation and adjustment of output parameters on the entire audio file. By only performing neural network learning modeling before the audio file is converted into formats, high-quality format conversion can be performed directly on the entire audio file, thus improving the overall conversion efficiency.

[0101] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.

Claims

1. A multi-threaded audio format conversion method, characterized in that, Includes the following steps: Step 100: Import audio files and identify the initial audio format of the audio files; classify the different audio files based on the initial audio format and the target audio format of each audio file. Step 200: Select a neural network learning model based on the initial audio format and target audio format of the conversion reference body, selectively extract the imported audio file to form a conversion reference body, the neural network learning model generates output parameters according to common audio specification parameters, and the conversion reference body performs audio format conversion sequentially according to the output parameters; In step 200, the implementation steps for selectively extracting the imported audio file to form a conversion reference body are as follows: The imported audio file is parsed, and the audio file is sequentially cut into multiple audio segments to obtain the audio specification parameters of each audio segment; Select audio segments with different combination forms of the audio specification parameters to form a conversion reference body, wherein the conversion reference body is composed of audio segments of different durations in the audio file; Step 300: Perform quality certification scoring on the conversion reference body that has been converted to the target audio format. Based on the scoring results and the audio specification parameters carried by the conversion reference body, adjust the neural network learning model, establish the correlation between the audio specification parameters and the neural network learning model, and generate a real-time neural network learning model that corresponds one-to-one with different groups of audio specification parameters, so as to adjust the output parameters in real time until the scoring results meet expectations. Step 400: Start the audio format conversion of the initial audio file, obtain the audio specification parameters of each largest audio segment, and the neural network learning real-time model adjusts the output parameters in advance according to the audio specification parameters of the next segment's audio segment until each initial audio file is converted to the corresponding target format.

2. The audio format conversion method based on multithreading according to claim 1, characterized in that, In step 200, the neural network learning model is pre-saved. The neural network learning model is initially selected based on the conversion relationship between the initial audio format and the target audio format. The neural network learning model adjusts the output parameters uniformly based on the general audio specification parameters corresponding to the initial audio format.

3. The audio format conversion method based on multithreading according to claim 1, characterized in that, The audio specifications include audio frequency, sampling frequency, sampling bit depth, number of channels, and bit rate. The output parameters include the number of output channels and the encoding format; The audio formats include MP3, WAV, AAC, FLAC, and OGG.

4. The audio format conversion method based on multithreading according to claim 1, characterized in that, In step 300, the specific steps for performing quality certification scoring on the converted reference body that has been converted to the target format are as follows: The conversion reference body is re-decomposed into multiple audio segments according to the combination of audio specification parameters; The neural network learning model sequentially adjusts the output parameters for each audio segment, while simultaneously acquiring the audio specification parameters for each audio segment; Each audio segment that has been converted to the target audio format is scored for quality certification, resulting in multiple scores for the audio segments. First, select the audio specification parameters of the audio segments with large differences in the scoring results to form the first dataset. Then, import the first dataset into the neural network learning model for multiple training sessions to adjust the output parameters until the scoring results of the audio segments with low scoring results after being converted to the target audio format meet the mathematical expectation. Then, the audio specification parameters of the audio segments with small differences in the scoring results are selected to form a second dataset. The second dataset is then imported into the neural network learning model for testing. The stability of the adjusted output parameters is verified by using the scoring results after the audio segments are reconverted to the target audio format.

5. The audio format conversion method based on multithreading according to claim 4, characterized in that, Establish the correlation between the audio specification parameters of the audio segment and the real-time learning model of the neural network to form a learning rule, and adjust the output parameter based on the learning rule and the audio specification parameters of the audio segment until the scoring result of each audio segment in the transformation reference body meets the mathematical expectation. The input value of the real-time neural network learning model is the audio specification parameter, and the output value of the real-time neural network learning model is the output parameter.

6. The audio format conversion method based on multithreading according to claim 5, characterized in that, The implementation model for quality certification scoring of the conversion reference body that has been converted to the target format is as follows: using PESQ to establish a simulated human ear auditory model to predict the listener's subjective score of the audio quality of the conversion reference body that has been converted to the target audio format; The rating ranges from -0.5 to 4.5, with higher scores indicating higher audio quality. It is used to evaluate the audio quality after audio format conversion.

7. The audio format conversion method based on multithreading according to claim 5, characterized in that, In steps 100 to 300, the initial audio file is paused for audio format conversion, and an audio specification conversion experiment is performed on the combined conversion reference body; In step 400, the learning rules for adjusting the output parameters based on the different audio specification parameters obtained are used to initiate the audio format conversion of the initial audio file.

8. The audio format conversion method based on multithreading according to claim 7, characterized in that, The specific steps to initiate audio format conversion of the initial audio file are as follows: The audio specification parameters of the audio of a set duration are obtained according to the capture frequency, and the audio of the set duration is divided into different audio segments according to the standard of matching different real-time learning models of the neural network. The neural network is used to learn a real-time model that adjusts the output parameters based on the audio specification parameters of the current audio segment; Based on the audio specification parameters and learning rules of the next audio segment, the adjustment target of the real-time learning model of the neural network is determined in advance, so as to determine the adjustment mode of the output parameters corresponding to the next audio segment in advance, and to perform timely audio format conversion on the next audio segment.

Citation Information

Patent Citations

  • Variable bitrate control for distributed video encoding

    US20170078676A1

  • Multi-stage processing of audio signals to facilitate rendering of 3D audio via a plurality of playback devices

    US20220021997A1