Multi-track audio synthesis, processing method and system
Audio data is read through the SDIO protocol and ECC mechanism, combined with deep learning and neural network technology, intelligent management and personalized editing of multi-track audio data are realized, solving the problems of robustness and insufficient user experience of multi-track audio processing in the existing technology, and improving the efficiency and quality of audio creation.
Patent Information
- Application Number
- CN202411216502.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-08-30
AI Technical Summary
The existing multi-track audio processing technology has shortcomings in the robustness of data preprocessing, collaborative optimization between multiple tracks, and user feedback integration. It is especially challenging to achieve automation and refined control in terms of real-time and consistency of user experience.
Audio data is read from SD NAND solid-state memory through SDIO protocol, data checksum correction is used to analyze audio data in combination with deep learning algorithms, audio tracks are dynamically allocated and managed, and intelligent mixing functions and situation-aware editing method are used for automatic adjustment, and neural network model is used for synthesis and compression processing.
Improve the automation level of audio processing, enhance the personalized adjustment capabilities of audio, improve user experience, and optimize audio quality and file size to ensure high-quality audio output.
Smart Images

Figure CN119252224B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent audio production technology, and in particular to a multi-track audio synthesis and processing method and system. Background Art
[0002] In the field of audio engineering, with the rapid development of digital technology, multi-track audio synthesis has become an indispensable part of modern music production and sound design; traditional audio processing methods, such as linear time-invariant filters, FFT transforms, time-frequency analysis, etc., although they have met the needs of audio editing and mixing to a certain extent, have limitations in personalized editing, contextual awareness, and real-time processing capabilities in complex scenarios; especially in dynamic audio track management, intelligent mixing, and context-aware editing, existing technologies often rely on manual intervention, making it difficult to achieve automation and refined control, which limits the efficiency and creativity of audio creation.
[0003] In recent years, the integration of deep learning and neural network methods has brought revolutionary breakthroughs in audio processing, not only improving the accuracy of audio analysis and synthesis, but also showing great potential in intelligent mixing, adaptive parameter adjustment and personalized editing. However, most current solutions still need to be improved in terms of robustness of data preprocessing, collaborative optimization between multiple tracks, and user feedback integration, especially in terms of real-time performance and consistency of user experience. In view of this, there is an urgent need for an innovative multi-track audio synthesis, processing method and system to overcome the shortcomings of existing technologies in automation, intelligence and user interaction, and to achieve a more efficient and intelligent audio creation and editing process. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a multi-track audio synthesis and processing method and system to solve the problems of low data reading efficiency, high error rate and non-intelligent mixing and editing in multi-track audio processing.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, an embodiment of the present invention provides a multi-track audio synthesis and processing method and system, which includes reading audio data from an SD NAND solid-state memory through the SDIO protocol, sending the data into a circular buffer, and implementing an ECC mechanism for data verification and error correction; using a deep learning algorithm to analyze the audio data in the buffer, and dynamically allocating and managing multiple audio tracks; for the allocated and managed audio data, automatically adjusting it using an intelligent mixing function and a context-aware audio editing method; independently adjusting ADSR parameters based on the intelligently mixed and personalized edited audio data, and performing synthesis and compression processing; based on the synthesized and compressed audio data, fusing a neural network model to generate multi-track audio data; returning the synthesized audio and performing further processing.
[0008] As a preferred solution of the multi-track audio synthesis and processing method and system of the present invention, the audio data is read from the SD NAND solid-state memory via the SDIO protocol, and the data is sent to the ring buffer, and the ECC mechanism is implemented to perform data verification and error correction. The specific steps are as follows:
[0009] Read raw audio data from SD NAND solid-state storage via SDIO protocol;
[0010] Transfer audio data to a ring buffer in memory;
[0011] During the reading process of the ring buffer RB, the ECC mechanism is applied to perform data integrity check and error correction. The expression is:
[0012] RB=E(d)+l;
[0013] Where E is the ECC mechanism, d is the original audio data read from the SD NAND, and l is the data read delay value.
[0014] As a preferred embodiment of the multi-track audio synthesis and processing method and system of the present invention, the method of using a deep learning algorithm to analyze the audio data in the buffer and dynamically allocate and manage multiple audio tracks comprises the following specific steps:
[0015] Use convolutional neural networks to extract spectral features of audio data from the ring buffer;
[0016] The feature sequence output by the convolutional neural network is input into the recurrent neural network to capture the dependencies between audio packets;
[0017] The output of the recurrent neural network is further processed using a long short-term memory network to optimize the track selection probability. The distribution function is then combined to dynamically distribute audio data packets to the corresponding tracks. The expression is:
[0018]
[0019] The allocation function expression is:
[0020] f(a i ,j)=ω(a i ,t j );
[0021] Among them, t j is a track value, n j is the total number of audio packets on the jth track, i is the index of all audio packets to be allocated, p j is the probability value of the audio data packet assigned to the jth audio track being selected, f(a i , j) is the distribution function, a i is the value of the i-th audio segment, j is t j The identifier of ω(a i , t j ) is the evaluation of a i With t j Function of matching.
[0022] As a preferred embodiment of the multi-track audio synthesis and processing method and system of the present invention, the distributed and managed audio data is automatically adjusted using an intelligent mixing function and a context-aware audio editing method, specifically in the following steps:
[0023] Adjust audio tracks using context-aware editing to personalize edits based on context;
[0024] Introducing the context relevance value r, the smart mixing function is used to combine the edited audio track with the context relevance to generate the adjusted audio track. The expression is:
[0025] t′ j =m(c(t j )·r);
[0026] Smart mixing function, the expression is:
[0027] m(c(t j ).r)=c(t j )+ΔP(r);
[0028] The context-aware editing function c(t j ), the expression is:
[0029] c(t j )=g(t j θ c );
[0030] Where t′ jis the adjusted track value, m is the intelligent mixing function, ΔP(r) is the parameter adjustment vector value calculated according to the context r, θ c are model parameters.
[0031] As a preferred solution of the multi-track audio synthesis and processing method and system of the present invention, wherein: the ADSR parameters are independently adjusted according to the audio data that has been intelligently mixed and personalized edited, and synthesis compression processing is performed, the specific steps are:
[0032] Adjust each track t′ j ADSR parameters to adapt to different types of audio data;
[0033] Adaptive adjustment and calibration based on dynamically changing parameters of audio characteristics;
[0034] Compress the adjusted audio track. The expression is:
[0035] t″ j = cmp(ADSR(t j )+b);
[0036] ADSR parameter adjustment function, the expression is:
[0037] (A, D, S, R) t = adsr (z);
[0038] Where t″ j Represents the audio track value after ADSR adjustment and compression, cmp is the compression function, b is the parameter adaptive adjustment value, A is the time from the start of the note to the peak volume, D is the time from the volume reaching the peak to the maintenance level, S is the stable volume value of the note in the sustained phase, R is the time from the volume starting to drop after the user releases the key until it disappears completely, t represents time, and z is the audio feature value.
[0039] As a preferred embodiment of the multi-track audio synthesis and processing method and system of the present invention, wherein: the multi-track audio data is generated based on the synthesized and compressed audio data and integrated with the neural network model, and the specific steps are as follows:
[0040] Use neural network models to perform deep synthesis of compressed audio tracks;
[0041] During the audio synthesis process, a dynamic weight matrix is used to adjust the importance of each audio track;
[0042] Introducing time series, at each time point, audio data synthesis is performed based on the dynamic weight matrix W and the neural network model F. The expression is:
[0043]
[0044] Where s(t) is the synthesized audio data at time point t, n represents the total number of tracks, and w it is the time point t for track t″ j Dynamic weight value.
[0045] As a preferred embodiment of the multi-track audio synthesis and processing method and system of the present invention, the steps of returning the synthesized audio and further processing the audio are as follows:
[0046] The final effect processing of the synthesized audio data is performed using real-time audio processing methods. The expression is:
[0047] a f =ef(s(t),e);
[0048] Real-time audio effect processing function ef, the expression is:
[0049] ef(s(t),e)=eq(s(t),e eq );
[0050] Among them, a f is the final output audio signal value, e is an additional real-time effect parameter, eq represents the equalizer processing value, e eq is a set of parameters for the equalizer processing values;
[0051] Implement dynamic range control, the expression is:
[0052] as=li(af,th);
[0053] Among them, a s is the safe audio signal value after limiting processing, li is the limiting function, th is the threshold, and x is the input audio signal sample value;
[0054] The processed audio signal value is played through the audio output interface, and a feedback mechanism is established to monitor the playback quality and status. The expression is:
[0055] o t =ao(a s ,d);
[0056] Among them, t is the actual output audio value, ao is the audio output function, and d is the characteristic parameter value of the output device;
[0057] Using the actual output audio value, the audio data is finally optimized for sound quality through post-processing functions;
[0058] Collect user feedback and adjust the coefficient based on the feedback to control the impact of the feedback on the final output;
[0059] The final output is the combination of optimized audio data and user feedback, expressed as:
[0060] o t '=po(s(t))+k·u f (t);
[0061] The post-processing function po is expressed as:
[0062] po(s(t))=s(t)*h(t);
[0063] Among them, t ' is the final output audio data, u f is the user feedback value, k is the feedback adjustment coefficient value, and h(t) is a filter function.
[0064] In a second aspect, the present invention provides a multi-track audio synthesis and processing system, comprising a data reading module, an analysis and management module, a mixing and editing module, an adjustment and compression module, a synthesis module, and a processing module;
[0065] The data reading module is used to read audio data from the memory through the SDIO protocol and use ECC to perform error correction to ensure that the data enters the buffer accurately; the analysis and management module is used to use deep learning algorithms to analyze audio data, dynamically allocate and manage multiple audio tracks, and realize intelligent audio organization;
[0066] The mixing and editing module is used to automatically adjust audio tracks using intelligent mixing functions and context-aware editing to achieve the best listening experience; the adjustment and compression module is used to independently adjust the ADSR parameters of the audio, perform synthesis and compression processing, and optimize audio quality and file size; the synthesis module is used to use the processed audio data fused by the neural network model to generate high-quality multi-track audio output; the processing module is used to perform final optimization processing on the synthesized audio, prepare and output the finished audio file, and ensure that it meets the playback standards.
[0067] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the multi-track audio synthesis, processing method and system described in the first aspect of the present invention is implemented.
[0068] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the multi-track audio synthesis, processing method and system described in the first aspect of the present invention is implemented.
[0069] The beneficial effects of the present invention are as follows: the present invention ensures the complete transmission of audio data through efficient data reading and error correction, uses deep learning to realize intelligent audio track management, and improves the automation level of audio processing; context-aware editing and intelligent mixing functions give audio personalized adjustment capabilities, enhancing user experience; dynamic ADSR parameter adjustment and compression optimize audio quality and file size, while neural network deep synthesis greatly improves sound quality details; finally, the optimization processing combined with user feedback ensures high-quality audio output, meets professional standards and personalized needs, and the entire process significantly improves the efficiency and artistic expression of audio production. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0071] Figure 1 Flowchart of the multi-track audio synthesis and processing method and system in Example 1.
[0072] Figure 2 This is a flowchart of audio processing and optimization in Example 1. DETAILED DESCRIPTION
[0073] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0074] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0075] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0076] Example 1, with reference to Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides a multi-track audio synthesis and processing method and system, including the following steps:
[0077] S1. Read audio data from the SD NAND solid-state memory through the SDIO protocol, send the data to the ring buffer, and implement the ECC mechanism for data verification and error correction.
[0078] Furthermore, the original audio data is read from the SD NAND solid-state memory via the SDIO protocol;
[0079] Audio data includes:
[0080] Audio sample data, which is the basic component of audio signals, consists of a series of digital samples. Each sample represents the quantized value of the amplitude of the sound wave at a certain moment. These samples are usually converted from a continuous analog signal through an analog-to-digital converter ADC. Sample data can be mono or stereo and have different bit depths, such as 16 bits, 24 bits, and sampling rates such as 44.1kHz, 48kHz;
[0081] Metadata. In addition to the audio samples themselves, audio files may also contain metadata such as artist name, song title, album information, genre, track number, copyright information, etc. For MIDI files, metadata also includes note event timestamps, note duration, note pitch, velocity value, controller changes, etc.
[0082] Compression information: If the audio data is in a compressed format such as MP3 or AAC, the read audio data may contain compressed encoded information, which requires a decoder to decode into the original PCM pulse code modulation data;
[0083] Audio effect settings: Audio files or related configuration files may contain preset audio effect parameters, such as equalizer settings, reverb parameters, compression ratio, etc. These parameters will affect subsequent audio processing and mixing;
[0084] Audio markers and tags. Audio data may contain markers for navigation or editing, such as chapter markers, cut points, or keyframes. These can help with precise synchronization and editing in multi-track audio compositions.
[0085] Time code and synchronization information. For multi-track recordings or live performance recordings, the audio data may contain time code such as MTC-MIDI time code or other synchronization signals to ensure synchronization between different tracks or devices;
[0086] In MIDI files, in addition to note information, there are also control information such as Control Change, Program Change, and sequencer-specific events SysEx, which are crucial to the timbre and effects of playing instruments.
[0087] Transfer audio data to a ring buffer in memory to ensure the continuity of data flow;
[0088] During the reading process of the ring buffer RB, the ECC mechanism is applied to perform data integrity check and error correction to ensure data integrity and prevent data damage from affecting subsequent processing. The expression is:
[0089] RB=E(d)+l;
[0090] Data checksum error correction function, the expression is:
[0091] E(d)=d′;
[0092] Where E is the ECC mechanism, a data checksum and error correction function that ensures data integrity. d is the original audio data read from the SD NAND. l is the data read delay value, which represents the data read delay, that is, the total time from the data being read to the data arriving at the ring buffer after being checked and corrected by the ECC mechanism. This delay reflects the time required for the reading process and is usually determined by hardware performance. d′ represents the data after ECC checksum and error correction.
[0093] Data is transferred to a ring buffer to ensure data continuity and stability, avoiding interruptions or delays during data reading and processing that could affect smooth audio playback. The use of a ring buffer can smooth the data flow, maintaining a continuous supply of data even if the reading speed fluctuates.
[0094] The ECC mechanism is used to detect and correct errors that may occur during data transmission or storage. If data corruption occurs during audio data reading, ECC can promptly detect and attempt to repair these errors, thereby ensuring data integrity and avoiding audio distortion or playback failure caused by data corruption.
[0095] Preferably, the data read from the memory may be packed or compressed and needs to be unpacked or decompressed to the original PCM data format;
[0096] If the sampling rate of the read audio data does not meet the system requirements, resampling may be required;
[0097] If the data is mono and the system requires stereo, or vice versa, the number of channels needs to be converted;
[0098] Ensure that the data format is compatible with the system, such as converting from unsigned integer to floating point;
[0099] It should be noted that
[0100] Data read latency l refers to the time required for data to be read from the storage medium, undergo ECC checksum and error correction, and then be sent to the ring buffer. This latency mainly includes the following components:
[0101] Addressing delay: This is the time required for the memory to determine the location of the data, that is, the time from issuing a read instruction to the time the memory finds the address where the data is located;
[0102] Read time: The time required to actually read the data, which involves transferring the data from the memory cell to the memory controller;
[0103] Data transfer latency: The delay in transferring data from memory to the processor or cache, which may be limited by bus speed and bandwidth;
[0104] ECC checksum and error correction time: The time required to apply the ECC mechanism for data integrity and error correction; this includes calculating the check bits, detecting errors, and possible error correction operations;
[0105] Data processing time: Other processing may be required before the data enters the ring buffer, such as decompression and unpacking;
[0106] Buffer write time: the time it takes to write data into the ring buffer, which depends on the buffer's write speed;
[0107] This latency, l, is determined by factors such as hardware performance, memory type, data volume, and processing logic. Reducing l is crucial in high-speed audio processing systems because lower latency means smoother data flow, which is particularly important for real-time audio processing.
[0108] The ECC algorithm uses Reed-Solomon coding, a forward error correction (FEC) method widely used in storage systems and particularly well-suited for handling sudden errors. Reed-Solomon coding can tolerate a certain number of errors without affecting data recovery, making it an ideal choice for SD NAND solid-state memory read and write operations.
[0109] During encoding, the original data is divided into fixed-size data blocks. Each data block is regarded as the coefficient of a polynomial defined over a finite field. The encoder calculates the roots of the polynomial, which constitute the so-called check symbols. These check symbols are appended to the end of the original data block to form the encoded data block.
[0110] After receiving the data, the receiver recalculates the roots of the polynomial and compares them with the received checksum. If the received data contains errors, the decoder attempts to recover the original data by solving the roots of the polynomial. The decoding algorithm for Reed-Solomon codes can correct a certain number of random errors or burst errors, the specific number of which depends on the parameters set during encoding.
[0111] When writing data, the encoder adds redundant bits to the original data so that even if part of the data is damaged during the writing or reading process, the original data can be restored by the decoder; when reading data, the decoder checks the integrity of the data and attempts to repair it using check bits if an error is detected; the ECC algorithm adds redundant information to the data, enabling data transmission and storage systems to maintain data integrity and reliability in the face of various interferences and failures.
[0112] The data checksum and error correction function is in the process of reading data from the storage to the memory. E(d) ensures the integrity and correctness of the data. It uses the ECC (Error Correction Code) mechanism, which is a method of detecting and correcting errors in data transmission or storage;
[0113] ECC adds redundant information (check codes) to the original data. When the data is read, E(d) checks whether these check codes are consistent with the data. If not, it attempts to recover the original data.
[0114] Through the above steps, the system can efficiently and reliably read audio data from the SD NAND while ensuring the integrity and continuity of the data, laying a solid foundation for subsequent audio processing and playback.
[0115] S2. Use deep learning algorithms to analyze the audio data in the buffer and dynamically allocate and manage multiple audio tracks.
[0116] Furthermore, a convolutional neural network is used to extract the spectral features of the audio data from the ring buffer;
[0117] Convolutional neural networks can capture local frequency and time domain features, which are crucial for subsequent audio track classification;
[0118] The feature sequence output by the convolutional neural network is input into the recurrent neural network to capture the dependencies between audio data packets. The recurrent neural network can understand the contextual relationship of the sequence data.
[0119] The RNN outputs the track selection probability, which is a prediction based on the sequence history that determines the likelihood of the audio packet being assigned to which track;
[0120] The output of the recurrent neural network is further processed using a long short-term memory network because LSTM can better handle long-term dependencies, which is especially important for dynamically managing audio tracks. It optimizes the track selection probability and dynamically allocates audio data packets to the corresponding tracks in combination with the allocation function. The expression is:
[0121]
[0122] The allocation function expression is:
[0123] f(a i ,j)=ω(a i ,t j );
[0124] Among them, t j Is a track value, the jth track, used to receive and manage specific types of audio data, n j is the total number of audio packets on the jth track, i is the index of all audio packets to be allocated, p j is the probability value of the audio data packet assigned to the jth track being selected, which is used to determine which track the audio data packet is assigned to, f(a i ,j) is the allocation function, which allocates the audio data packet a i Assign to track t j , a i is the audio data packet, j is t j Identifier used to distinguish different tracks, ω(a i , t j ) is the evaluation of a i With t j A match function that returns a value where a higher value indicates a better match. This function can be rule-based or part of a machine learning model, such as a neural network, that predicts the best track assignment based on learned patterns.
[0125] Preferably, the allocation function f(a i ,j) is a core component that decides how to send audio data packets a i Assign to a specific track j ,Its goal is to determine the most appropriate audio track to receive the data packet based on the characteristics of the audio data packet and the properties of the audio track, in order to achieve optimization and efficiency of audio processing;
[0126] The main function of the allocation function is to evaluate the compatibility between each audio data packet and the audio track based on a series of criteria (such as audio type, frequency range, audio track capacity, etc.), and thus determine which audio track is most suitable to receive this data packet; in this way, a reasonable distribution of audio data can be ensured to avoid overloading some audio tracks while leaving other tracks idle, while ensuring the continuity and quality of the audio.
[0127] The allocation function determines how to allocate the audio segment a i Assign to a specific track j superior;
[0128] Use machine learning models, such as neural networks, to calculate a match based on the characteristics of the audio clip and the characteristics of the track to determine the allocation strategy;
[0129] Preferably, the matching function ω(a i , t j ) is used to quantize the audio segment a i and audio track t j compatibility or similarity between them;
[0130] Matching function ω(a i , t j ) may be calculated based on the following factors:
[0131] Spectral similarity: compares the spectral characteristics of an audio clip with those already in the track to determine whether they belong to the same category or style;
[0132] Temporal characteristics: Consider the temporal characteristics of the audio clip, such as intensity, rhythm, and similarity with other clips in the track;
[0133] Acoustic features: Use acoustic features such as MFCC (Mel Frequency Cepstral Coefficients), zero-crossing rate, energy, etc. for comparison;
[0134] Music theory: For example, if an audio clip matches the chord structure, melodic line, or rhythmic pattern of the track;
[0135] Contextual information: considers the position of an audio clip in the time sequence and its relationship to the preceding and following clips;
[0136] Output of the machine learning model: directly use the output of the trained neural network model as the measure of matching;
[0137] Use neural network to predict the matching degree, the expression is:
[0138] ω(a i ,t j )=NN match (a i ,t j);
[0139] Among them, NN match is a neural network trained to predict the audio segment a i and audio track t j The degree of match between
[0140] There is also a weighted summation of various features, the expression is:
[0141] ω(a i ,t j )=w1·sim freq (a i ,t j )+w2·sim time (a i ,t j )+w3·sim acoustic (a i ,t j );
[0142] Among them, sim freq Represents spectrum-based values, sim time Represents the time domain value, sim acoustic Represents the similarity measure of acoustic features, is the weight of the corresponding feature, which can be adjusted through experiments or automatically learned through optimization algorithms;
[0143] Preferably, before analyzing the audio data in the buffer using a deep learning algorithm, the following preprocessing may be required:
[0144] Standardize or normalize the audio data to suit the input requirements of the neural network;
[0145] Extract meaningful features from raw audio data, such as Mel-Frequency Cepstral Coefficients (MFCCs) and spectrograms, which will serve as input to deep learning models;
[0146] Split the audio data into small segments to facilitate model processing and subsequent track management;
[0147] It should be noted that convolutional neural networks (CNNs) capture local frequency and time domain features through convolutional layers. In audio processing, convolutional neural networks (CNNs) are often used to represent audio features such as spectrograms or Mel-frequency cepstral coefficients (MFCCs). The filters in the convolutional layers slide over the input feature maps and perform weighted summation operations on local areas, which can capture local patterns and structures, such as the presence of specific frequency components or transient changes in sound.
[0148] Convolutional Neural Network CNN workflow:
[0149] Design the filter width, height, and number of channels to match the input audio feature dimensions; for example, a 2D convolutional layer might use a smaller filter, such as 3x3, to scan the spectrogram to capture local frequency and time domain features;
[0150] Apply nonlinear activation functions such as ReLU after convolution operations to increase the expressive power of the model;
[0151] After the convolutional layer, the pooling layer can further reduce the spatial dimension while retaining the most important features, which helps reduce the amount of computation and prevent overfitting;
[0152] Once the convolutional neural network (CNN) extracts local features in the frequency and time domains, the recurrent neural network (RNN) can further process these feature sequences to capture dependencies and contextual information in the sequence.
[0153] The workflow is:
[0154] The feature sequence extracted by the convolutional neural network (CNN) is used as the input of the recurrent neural network (RNN), where each time step corresponds to a segment or feature vector of the audio data.
[0155] Recurrent neural networks (RNNs) combine information from previous time steps with information from the current time step through a hidden state transfer mechanism. This allows them to capture the historical dependencies of sequence data, which is crucial for understanding the continuity of musical melody, rhythm, or speech.
[0156] The output layer of the RNN can be another neural network layer to predict the probability of track selection;
[0157] Long Short-Term Memory (LSTM) is a special type of recurrent neural network (RNN). It uses gating mechanisms (input gate, forget gate, and output gate) to control the flow of information, thereby better handling long-term dependencies. In track management, LSTM can optimize track selection probabilities, ensuring that audio data packets are assigned to the most appropriate tracks.
[0158] The workflow is:
[0159] The gating mechanism in the Long Short-Term Memory (LSTM) unit allows the network to selectively remember or forget information, which is particularly important when processing long sequence data because it can avoid the gradient vanishing or gradient exploding problem;
[0160] The hidden states of the Long Short-Term Memory (LSTM) network transmitted in the sequence contain long-term dependency information, which enables the model to take into account more distant historical information when predicting the probability of track selection, thereby making more intelligent decisions;
[0161] Using three deep learning algorithms: convolutional neural network (CNN), recurrent neural network (RNN), and long short-term memory (LSTM), we ultimately generated a formula for track assignment.
[0162] The entire process uses a convolutional neural network (CNN) to extract features, a recurrent neural network (RNN) to capture dependencies, and a long short-term memory (LSTM) network to process long-term dependencies. Ultimately, dynamic allocation of audio data packets is achieved through track selection probabilities and allocation functions. This deep learning approach not only improves the efficiency of audio processing, but also enhances the intelligence and accuracy of track management, ensuring the correct processing and smooth playback of audio data.
[0163] S 3. For the allocated and managed audio data, use intelligent mixing functions and context-aware audio editing methods to automatically adjust.
[0164] Going a step further, use context-aware editing methods to adjust the audio track and make personalized edits based on the context;
[0165] The context-sensitive relevance value r is introduced to enable the audio editing and mixing process to respond to real-time environmental changes and user preferences, thereby achieving a more personalized and dynamic audio experience. It is a dynamic value determined by environmental sensors and user input;
[0166] Use the smart mix function to combine the edited audio track with the contextual relevance to generate the adjusted audio track. The expression is:
[0167] t′ j =m(c(t j )·r);
[0168] Smart mixing function, the expression is:
[0169] m(c(t j ).r)=c(t j )+ΔP(r);
[0170] The context-aware editing function c(t j ), the expression is:
[0171] c(t j )=g(t j θ c );
[0172] Where t′ j is the adjusted track value, m is the intelligent mixing function used to integrate the edited track and context relevance, used to adjust the track according to the context, r is the context relevance value, a dynamic value determined by environmental factors and user input, used to adjust the edit of the track, ΔP(r) is the parameter adjustment vector value calculated according to context r, and θc It is a model parameter, which comes from the sequence-to-sequence model in deep learning. This model can be based on the sequence-to-sequence (seq2seq) architecture, which is specially designed to process time series data, such as audio sequences;
[0173] Preferably, the intelligent mixing function combines the context-dependent r and the context-aware editing function c(t j ), adjust the audio track t j parameters to make them more suitable for the current situation;
[0174] It may be possible to use deep learning models to learn mixing rules for different situations and dynamically adjust the volume, equalization, and other parameters of the audio track.
[0175] Context-aware editing function c(t j ) is designed to adapt the soundtrack to real-time environmental changes and user preferences. j The purpose of this is to make the audio editing process more intelligent and dynamic, so as to provide a personalized audio experience;
[0176] The context-aware editing function analyzes data from environmental sensors and user input, such as sound levels, geographic location, time of day, and user preferences, and then adjusts the parameters of the audio track based on this information. For example, it might adjust the volume, pitch, timbre, and tempo to better suit the current context. This way, even when the environment changes, such as switching from outdoor exercise to a quiet library, the audio output still maintains an optimal listening experience.
[0177] Context-aware editing functions adjust the content and parameters of audio tracks according to the context, making editing more tailored to the user's listening experience;
[0178] Through the neural network model g(t j θ c ), according to the track characteristics and context parameters θ c to adjust the track's timbre, dynamic range, and more;
[0179] By using context-aware editing functions, the audio system can intelligently adapt to different scenarios, providing more personalized and dynamic audio content, ensuring that the audio output matches the user's environment and preferences.
[0180] Thus enhancing the user experience;
[0181] Preferably, between smart mixing and context-aware editing, the following steps may be involved:
[0182] Aggregate contextual information from different sensors and user inputs for subsequent analysis and application;
[0183] Dynamically update the parameters of the context-aware editing function based on the latest context information to reflect the latest status of the current context;
[0184] Based on preliminary situational analysis, pre-adjust track parameters such as volume, pitch, and timbre to prepare for subsequent intelligent mixing;
[0185] It should be noted that, first, the system collects information through environmental sensors, such as sound levels, geographical location, weather conditions, etc., while also taking into account user input, such as preferred music genres, volume preferences, etc.;
[0186] This data is fed into a context-aware editing function that analyzes the information and adjusts the soundtrack accordingly; for example, if the user is exercising outdoors, the system might increase the tempo and energy of the background music; if the user is reading in a quiet library, the system might lower the volume and select softer music.
[0187] Contextual relevance is a dynamic value that reflects the degree to which the current context influences the audio content. This value can be rule-based or predicted by a machine learning model, and it updates in real time based on changes in the environment and user interactions.
[0188] The adjusted tracks and contextual relevance are passed to a smart mixing function; this function can be a complex deep learning model that understands the best mixing strategies for different situations. For example, it might know how to improve vocal clarity in noisy environments, or how to maintain a balanced background music in quiet environments.
[0189] Finally, the smart mixing function combines the context-aware edited tracks and context-dependent features to generate an adjusted track that best suits the current context;
[0190] It should also be noted that context-aware editing refers to an audio processing method that can dynamically adjust audio content based on the characteristics of the current context. The context can include a variety of factors, such as: physical environment: such as noise level, room size, temperature, humidity, light intensity, etc.; geographical location: for example, in different locations such as parks, indoors, city centers or the seaside; time: the time of day, holidays or special dates; activities: what the user is doing, such as exercising, resting, working or gathering; user preferences: personal preferences, emotional state, health status, etc. The context-aware editing method collects this information and then applies algorithms to adjust audio parameters such as volume, pitch, timbre, rhythm, etc., so that the audio output is more in line with the current context and provides a better listening experience.
[0191] Contextual relevance refers to the degree to which a context influences the audio content. This is a dynamic value that reflects the strength of the connection between the context and the audio content. For example, if a user is running outdoors, high-energy music may have higher contextual relevance because it better matches the nature of the activity. Contextual relevance can be determined by a rules engine or machine learning model, and it updates in real time based on changes in the context and user behavior.
[0192] A smart mixing function is an algorithm or model that is responsible for combining multiple audio tracks or audio sources into a coordinated and consistent output. In context-aware audio systems, the smart mixing function not only focuses on the methodical mixing between audio tracks, but also considers the results of contextual relevance and context-aware editing to optimize the quality and adaptability of the final audio output.
[0193] Smart mixing functions may include:
[0194] Dynamic Range Control: Adjusts the loudness and contrast of audio tracks to suit varying ambient noise levels;
[0195] Spatialization: Changes the positioning of sounds based on the user's position or movement direction, creating a surround sound effect.
[0196] Tonal balance: ensuring that different instruments or sound types remain harmonious when mixed, without one part being too prominent or drowning out others;
[0197] Personalized adjustments: Fine-tune the mix based on user preferences, such as boosting bass, reducing treble, or adjusting vocal clarity;
[0198] The entire process reflects the integration of technology, environment, and human preferences, providing users with a personalized and intelligent audio experience. As technology develops, context-aware audio systems will become more and more mature, able to more accurately understand user needs and provide a more seamless and natural listening experience.
[0199] S 4. Independently adjust ADSR parameters based on the intelligently mixed and personalized edited audio data, and perform synthesis and compression processing.
[0200] Furthermore, adjust each track t′ j ADSR parameters to adapt to different types of audio data;
[0201] Adaptive adjustment and calibration based on dynamically changing parameters of audio characteristics;
[0202] Compress the adjusted audio track to prevent distortion and optimize the sound quality. The expression is:
[0203] t″ j = cmp(ADSR(t′ j)+b);
[0204] The cmp compression function is expressed as:
[0205] cmp(a)=a c ;
[0206] ADSR parameter adjustment function, the expression is:
[0207] (A,D,S,R) t = adsr(z);
[0208] Where t″ j Represents the audio track value after ADSR adjustment and compression, cmp is the compression function, a is the uncompressed audio data, a c It is compressed audio data, which is used to compress audio tracks to avoid distortion and optimize sound quality. b is the parameter adaptive adjustment value, which is used to fine-tune ADSR parameters to adapt to different audio data types. A is the abbreviation of Attack. The attack time is the time from the start of the note to the peak volume. In synthesized audio, this is usually the initial sound burst. D is the abbreviation of Decay. The decay time is the time from the peak volume to the hold level. This is the process of the note decreasing from maximum intensity to a stable level. S is the abbreviation of Sustain. The hold level is the stable volume value of the note in the sustained phase. This is the volume level of the note when the user continuously presses the key. R is the abbreviation of Release. The release time is the time from the user releasing the key until the volume starts to decrease until it completely disappears. This is the tail sound process at the end of the note. t represents time, which means that the ADSR parameters described by this expression change over time. That is, these parameters may be different at different time points in the audio. z is the audio characteristic value, which comes from audio signal processing and synthesizer theory and is used to guide the adjustment of ADSR parameters.
[0209] Preferably, the following steps may be required between ADSR parameter adjustment and compression processing:
[0210] Set initial ADSR parameters for each track, which may be based on the audio type and common playback environment;
[0211] Under the guidance of the dynamic weight matrix, ADSR parameters are adjusted according to the real-time analysis of audio content to adapt to the characteristics of audio data and context changes;
[0212] The compression function is to reduce the size of audio files for easy storage and transmission while maintaining the highest possible sound quality;
[0213] Use lossy or lossless encoding techniques, such as MP3, AAC, or FLAC, to reduce bitrate by removing information that is imperceptible to the human ear;
[0214] The ADSR parameter adjustment function dynamically adjusts the Attack, Decay, Sustain, and Release parameters of the audio track according to the audio feature c.
[0215] By analyzing the spectrum and dynamic characteristics of the audio signal, the ADSR curve is adjusted to match the natural dynamics of the audio material.
[0216] It should be noted that regarding the interpretation of c as an audio feature, c in (A, D, S, R) t =adsr(c) represents a series of characteristics of the audio signal, which can be:
[0217] Spectral characteristics: such as spectral density, spectral center, spectral bandwidth, etc. These characteristics reflect the characteristics of the audio signal in the frequency domain;
[0218] Time domain features: including zero crossing rate, zero crossing point, short-time energy, etc. These features describe the changes of audio signals in the time domain;
[0219] Cepstral features: such as Mel-frequency cepstral coefficients (MFCC), which are cepstral representations of the spectral characteristics of audio signals on the Mel-frequency scale and are commonly used in speech recognition and music information retrieval;
[0220] Statistical features: based on the statistical properties of the signal, such as mean, variance, skewness, kurtosis, etc.
[0221] Transient features: describe the characteristics of transient events in audio signals, such as the transient detection function (TDF);
[0222] Rhythm features: reflect the characteristics of the rhythm pattern in the audio signal, such as beat period and beat intensity;
[0223] These features are extracted from the original audio signal, usually through digital signal processing (DSP) technology; for example, the audio signal may first be converted to the frequency domain through Fourier transform, and then the spectral features are extracted from it; or the signal may be framed and features are calculated for each frame to obtain a time-varying feature sequence.
[0224] In the ADSR parameter adjustment function, the characteristics of c are used to dynamically adjust the attack, decay, sustain, and release parameters of the audio track. This is because different types of audio signals, such as percussion instruments, string instruments, and vocals, have different dynamic characteristics and require different ADSR parameters to accurately reproduce their natural timbre and dynamics. For example, percussion instruments may require fast attack and decay, while string instruments may require slower attack and longer sustain.
[0225] Therefore, c, as a set of audio features, is obtained by signal processing and feature extraction of the original audio signal, and is used to guide the adaptive adjustment of ADSR parameters to adapt to different types of audio data;
[0226] It should also be noted that
[0227] ADSR parameters are adjusted as follows:
[0228] Attack: Controls how quickly the sound rises from silence to full volume at the start of a note.
[0229] Decay: Specifies how quickly the volume drops from the peak of the attack phase to the sustained volume.
[0230] Sustain: Sets the volume level that a note maintains before being released.
[0231] Release: defines how quickly the volume drops to silence after the note is released;
[0232] The adjustment of ADSR parameters needs to take into account the characteristics of the audio. For example, for percussion, fast attack and decay may be needed to capture transients; for string instruments, slow attack and long sustain may be needed to simulate real performances.
[0233] The parameters are adaptively adjusted to:
[0234] Dynamically adjust ADSR parameters based on audio type and context to achieve optimal auditory effects. For example, in a noisy environment, a faster attack may be required to make the sound stand out more.
[0235] The compressor is used to adjust the dynamic range of the audio signal to avoid overload and distortion. The compression function cmp will reduce the signal amplitude that exceeds the threshold according to the preset ratio based on the input signal level, thereby optimizing the sound quality.
[0236] Finally, the adjusted and compressed audio track t″ j This will provide a more stable, high-quality listening experience, maintaining good sound quality and dynamic performance in both noisy and quiet environments; this approach is particularly suitable for live performances, studio production, and streaming audio applications, ensuring that audio content performs well under various conditions.
[0237] S 5. Based on the synthesized and compressed audio data, the neural network model is integrated to generate multi-track audio data.
[0238] Going a step further, a neural network model is used to perform deep synthesis on the compressed audio tracks to enhance the sound quality and details;
[0239] Considering the changing importance of each track during the synthesis process, a dynamic weight matrix is used;
[0240] Introducing time series, at each time point, audio data synthesis is performed based on the dynamic weight matrix W and the neural network model F. The expression is:
[0241]
[0242] Among them, s(t) is the synthesized audio data at time point t, n represents the total number of tracks, which is used to deeply synthesize compressed tracks to improve sound quality and details, and w it is the time point t for track t″ j Dynamic weight value, used for synthesis calculation, including all w it Elements are used to reflect the changes in the importance of each track in the synthesis process;
[0243] Preferably, before using the neural network model for deep synthesis, the following steps may be required:
[0244] Loading pre-trained neural network models to ensure the model is ready for deep synthesis;
[0245] Preprocess the audio tracks to be synthesized to ensure they meet the input requirements of the neural network model;
[0246] Based on the current context and audio content, a dynamic weight matrix is calculated for subsequent audio track synthesis.
[0247] It should be noted that the dynamic weight matrix W is used to reflect the importance of different tracks over time during the synthesis process; at each time point t, the weight w it Automatically adjusts based on the content and context of the tracks to ensure the most relevant or important tracks receive more attention in the mix;
[0248] Preferably, the neural network model F is used for deep synthesis, which can learn complex feature maps from the training data, enabling the model to improve audio quality, enhance details, and even recover information that may have been lost during compression. The model is usually based on architectures such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTMs), or generative adversarial networks (GANs).
[0249] At each time point t, the synthesized audio data s(t) is obtained by j The output of (t) and the corresponding weight w in the dynamic weight matrix it Multiply and then sum them up; the above formula performs a linear combination of all n tracks, with weights w itThis ensures that the contributions of different tracks are weighted according to their importance at a specific point in time;
[0250] The overall process can be simplified as follows:
[0251] First, for each track t″ j Applying a neural network model F to enhance its sound quality and details;
[0252] Then, using the weight w in the dynamic weight matrix W it , perform a weighted sum of these enhanced tracks at each time point t;
[0253] Ultimately, the output is a synthetic audio signal s(t), which combines the advantages of all tracks and reflects the relative importance and dynamic changes of the tracks at each time point;
[0254] This technology is particularly suitable for music production, film scoring, game sound effects and other fields. It can create a richer and more layered audio experience while also handling complex dynamic changes in audio. Through deep learning methods, the model and weight matrix can be continuously optimized to obtain higher quality synthesized audio.
[0255] S6. Return the synthesized audio for further processing.
[0256] The synthesized audio data is processed using real-time audio processing methods to achieve the final effect, ensuring low latency and high-quality audio output. The expression is:
[0257] a f =ef(s(t),e);
[0258] Real-time audio effect processing function, the expression is:
[0259] ef(s(t),e)=eq(s(t),e eq );
[0260] Among them, a f is the final output audio signal value, e f It is a real-time audio effect processing function. It receives two parameters, one is the synthesized audio data s(t), and the other is an additional real-time effect parameter e. e is an additional real-time effect parameter, such as equalizer settings, reverberation depth, compression ratio, etc. q is the equalizer processing value, e eq It is a parameter set of equalizer processing values. The parameters include frequency point, gain, and bandwidth. The frequency point is the specific frequency that needs to be enhanced or weakened. The gain is the amount of gain at the selected frequency point, indicating the degree to which the sound should be amplified or attenuated. The bandwidth is the width of the frequency range around the frequency point that affects the equalizer's effect.
[0261] Implement dynamic range control to ensure that the peak value of the audio signal does not exceed the safe range to avoid distortion. The expression is:
[0262] a s =li(a final ,th);
[0263] The limiting function is expressed as:
[0264]
[0265] Among them, a s Is the safe audio signal value after clipping, li is the clipping function, th is the threshold used to prevent the audio signal from exceeding the safe level, x is the input audio signal sample value, which can be positive or negative, indicating the instantaneous amplitude of the audio signal at a certain point, sign is used to determine the sign of a number, if is used to determine whether the absolute value of the audio signal sample exceeds a threshold, when |x|≤th means if the final output audio signal is less than or equal to the threshold, the clipping function will directly return x, which means that the signal is unaffected within this range because it has not reached the amplitude that may cause problems. In audio processing, this usually means that the signal amplitude is safe and does not require any adjustment or clipping. When |x|>th means if the final output audio signal is greater than the threshold, this means that the signal amplitude exceeds the range we consider safe or ideal. In this case, the clipping function clips or limits the signal to the size of th while keeping the positive and negative directions of the signal unchanged. Specifically, the function will return sign(x)·th, sign(x) determines the positive and negative sign of x, th is the preset threshold, and the signal amplitude will be clipped to this value;
[0266] The processed audio signal value is played through the audio output interface, taking into account the characteristic parameters d of the output device, such as the impedance of the speaker and the type of headphone jack, to ensure the best match. At the same time, a feedback mechanism is established to monitor the playback quality and status. The expression is:
[0267] o t =ao(a s ,d);
[0268] Audio output function, the expression is:
[0269] ao(a s ,d)=calibrate(a s ,d);
[0270] Among them, tis the actual output audio value, ao is the audio output function, d is the characteristic parameters of the output device, including speaker impedance, headphone jack type, etc., calibrate is used to calibrate the data, which is used to adjust the audio signal a s , making it more suitable for playback on devices with characteristic parameters d, which may include adjusting volume, frequency response, dynamic range, etc. to optimize output quality; 3
[0271] Using the actual output audio value, the audio data is finally optimized for sound quality through post-processing functions;
[0272] Collect user feedback and adjust the coefficient based on the feedback to control the impact of the feedback on the final output;
[0273] The final output is the combination of optimized audio data and user feedback, expressed as:
[0274] o t '=po(s(t))+k·u f (t);
[0275] Post-processing function, the expression is:
[0276] po(s(t))=s(t)*h(t);
[0277] Among them, t ' is the final output audio data, including post-processing and user feedback adjustment, po is the post-processing function, which optimizes the sound quality of the final audio data s(t), u f is the user feedback value, which is the user's evaluation of the audio effect collected at time point t. k is the feedback adjustment coefficient value, which controls the degree of influence of user feedback on the final output. h(t) is a filter function derived from digital signal processing theory.
[0278] Preferably, in audio processing, the filter function h(t) is a key component for modifying the spectral characteristics of an audio signal;
[0279] Filters can be used for a variety of purposes, including but not limited to equalization, noise reduction, dynamic range control, frequency separation, etc. Filter design is usually based on specific frequency response requirements to achieve the desired audio processing effect;
[0280] Filter functions can be of different types, such as low-pass filters, high-pass filters, band-pass filters, band-stop filters, notch filters, etc. Each type of filter has its own unique transfer function, which determines how it processes signals of different frequencies.
[0281] The filter function can be expressed in the time domain or frequency domain. In the frequency domain, the filter function H(f) is expressed as a function of frequency f, which defines the gain or attenuation of the signal at different frequency points. The expression is:
[0282]
[0283] Among them, f c is the cutoff frequency value, j is the imaginary unit value, and this expression describes the frequency response of a first-order Butterworth low-pass filter;
[0284] In the time domain, the filter function h(t) is expressed as a function of time t, which is the impulse response of the filter. For a linear time-invariant system, the convolution of the filter function h(t) with the input signal q(t) gives the output signal y(t), which is expressed as:
[0285] y(t)=q(t)*h(t);
[0286] The filter function h(t) or H(f) is used in audio processing to modify the spectral characteristics of the signal to achieve functions such as equalization, noise reduction, and dynamic range control. In post-processing functions, the filter function is applied to the audio signal through convolution operations to optimize the final output sound quality.
[0287] Preferably, before the final audio output, the following steps may be required:
[0288] Convert the synthesized audio data into a format suitable for the output device, such as PCM, MP3, etc.;
[0289] Adjust the audio signal to ensure the best match based on the characteristics of the output device, such as speaker impedance, headphone jack type, etc.
[0290] While outputting audio, a feedback mechanism is established to collect user feedback for subsequent system optimization;
[0291] The real-time audio effect processing function adds effects such as echo, reverberation, and equalization during audio playback to enhance the listening experience.
[0292] Use DSP (Digital Signal Processing) technology to process audio signals in real time and apply the required audio effects;
[0293] The limiting function prevents the audio signal from exceeding the dynamic range of the device and avoids clipping distortion;
[0294] Monitor the peak value of the audio signal and clip it when it exceeds the threshold th to prevent overload;
[0295] The audio output function converts the processed audio signal into a format that can be played by the physical device;
[0296] Taking into account the characteristic parameters d of the output device, such as sampling rate and bit depth, the digital signal is converted into an analog signal for playback by speakers or headphones;
[0297] The post-processing function is to perform the final optimization on the audio, improve the sound quality, and prepare for the final output;
[0298] Use Filter h(t) , such as equalizers and dynamic processors, to fine-tune the audio signal to ensure clear and pleasant sound;
[0299] Managing differences between different output devices, especially in frequency response, is a crucial aspect of audio engineering. Speakers and headphones often have significant frequency response differences due to differences in design and usage environments. Here are some ways to address and compensate for these differences:
[0300] Speakers are generally designed to provide balanced sound over a wide frequency range, but due to the limitations of physical size and driver units, the high and low frequency response may not be as flat as headphones;
[0301] Because headphones are closer to the ear canal, they tend to provide more accurate frequency response, especially in-ear headphones, but some frequencies may be overly boosted or attenuated;
[0302] Collect frequency response curves for different devices. This information can be found in the manufacturer's technical specifications, or measured using specialized measurement equipment such as an audio analyzer.
[0303] Applying an inverse frequency response curve (i.e., a compensation curve) to the audio signal to offset the device's own frequency response characteristics, usually accomplished through an equalizer (EQ) function in software;
[0304] Different devices have different dynamic ranges, and headphones may be more susceptible to damage from wide dynamic range audio. Using a compressor can reduce peak volume levels and ensure that the audio signal remains undistorted on all devices.
[0305] Many professional audio processing software and hardware devices allow you to create and save presets that are optimized for specific speaker or headphone types. Users can choose the preset that is closest to their device for the best sound quality.
[0306] For more advanced systems, built-in or external microphones can be used to analyze the acoustic characteristics of the room and automatically adjust the audio output to compensate for echo, reverberation, and other room effects. The same applies to headphones, which analyze the shape and position of the wearer's ears to optimize sound positioning and clarity.
[0307] Using virtualization technologies, such as virtual surround sound, it is possible to simulate the experience of speakers on headphones, and vice versa. This usually involves performing complex mathematical operations on the audio signal to create a more immersive listening environment;
[0308] Allows users to adjust audio settings such as volume, tone, bass boost, etc. according to personal preferences, which can compensate for small differences between different devices and provide a more personalized listening experience.
[0309] At the software level, these adjustments can be implemented using digital signal processing (DSP) algorithms, which can be integrated into audio playback software, operating systems, or dedicated audio processing devices. For example, equalizer algorithms can use FIR (Finite Impulse Response) filters to achieve frequency response correction, while dynamic range control and compression can be achieved using nonlinear algorithms.
[0310] At the hardware level, many modern audio devices have built-in DSP chips that can process audio signals directly without the need for external software intervention. High-end headphones and speaker systems may even have adaptive audio processing capabilities that can automatically identify the type of device connected and adjust accordingly.
[0311] In short, dealing with the differences between different output devices requires a comprehensive application of audio processing technologies, from frequency response correction to environmental awareness to personalization, to ensure the best audio experience on any device;
[0312] It should be noted that the entire process, starting from audio synthesis, through real-time effects processing, dynamic range control, audio output and feedback mechanism, and finally to post-processing and user feedback adjustment, forms a closed-loop system to ensure the quality of audio output and user satisfaction; this closed-loop system allows real-time adjustment to make the audio output more in line with user needs and preferences.
[0313] This embodiment also provides a multi-track audio synthesis and processing system, including: a data reading module, an analysis and management module, a mixing and editing module, an adjustment and compression module, a synthesis module, and a processing module; the data reading module is used to read audio data from the memory through the SDIO protocol and use ECC to correct errors to ensure that the data enters the buffer accurately; the analysis and management module is used to use a deep learning algorithm to analyze audio data, dynamically allocate and manage multiple audio tracks, and realize intelligent audio organization; the mixing and editing module is used to use intelligent mixing functions and context-aware editing to automatically adjust audio tracks to achieve the best listening experience; the adjustment and compression module is used to independently adjust the ADSR parameters of the audio, perform synthesis and compression processing, and optimize audio quality and file size; the synthesis module is used to use a neural network model to fuse the processed audio data to generate high-quality multi-track audio output; the processing module is used to perform final optimization processing on the synthesized audio, prepare and output the finished audio file, and ensure that it meets the playback standards.
[0314] This embodiment further provides a computer device suitable for use with a multi-track audio synthesis and processing method and system, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the multi-track audio synthesis and processing method and system proposed in the above embodiment.
[0315] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.
[0316] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-track audio synthesis and processing method and system proposed in the above embodiments; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0317] In summary, the present invention achieves efficient, intelligent, and high-quality multi-track audio processing through: SDIO high-speed data reading, ECC error correction, deep learning audio track management, context-aware intelligent mixing, dynamic ADSR adjustment and compression, and neural network deep synthesis and user feedback optimization technology, greatly improving the flexibility of audio creation and the listening experience.
[0318] Example 2, referring to Table 1, is the second embodiment of the present invention. To further verify the advancement of the present invention, experimental simulation data of a multi-track audio synthesis and processing method and system are provided.
[0319] In order to verify the innovation and advantages of the content of the present invention, a set of experiments were designed to compare the performance of the existing technology and the present invention in reading audio data, dynamic audio management, context-aware editing, ADSR parameter adjustment, multi-track synthesis and real-time audio processing. Two systems were used in the experiments: a traditional system (T) and the system of the present invention (I).
[0320] The experiments used the same set of audio files, including different types of audio data (music, dialogue, natural sounds, etc.); the audio files had different metadata, compression formats, audio tracks, and time codes; and both systems ran on similar hardware environments to eliminate the influence of external variables.
[0321] First, the device reads audio data from the SD NAND drive using the SDIO protocol, transfers the data to a ring buffer, and applies the ECC mechanism for data verification and error correction. We use Reed-Solomon coding as the ECC algorithm to ensure data integrity. We then use CNN, RNN, and LSTM deep learning algorithms to analyze the audio data in the buffer, dynamically allocate and manage audio tracks, and improve audio processing efficiency.
[0322] Secondly, we introduce context-aware editing methods to adjust audio tracks based on the user's environment and preferences. We use intelligent mixing functions to mix audio to ensure the personalization and dynamic nature of the audio content. We also adjust the ADSR parameters of the audio tracks to adapt to different audio types and perform compression processing to prevent distortion and optimize sound quality.
[0323] Finally, a neural network model is used to deeply synthesize the compressed audio track, taking into account the changes in the importance of the audio track and introducing time series for synthesis; then the final audio effect processing, dynamic range control, audio output, and user feedback are collected for optimization.
[0324] The details are shown in Table 1 below:
[0325] Table 1 Experimental record table
[0326]
[0327] By comparing the data in the above tables, the system (I) of the present invention is significantly superior to the traditional system (T) in all key indicators; specifically: the reading rate and ECC efficiency of the system of the present invention are improved by 25% in the reading rate, and the ECC correction rate is also improved by 5%, which proves the effectiveness of the SDIO protocol and ECC mechanism; the accuracy of audio data management and allocation in the system of the present invention is 13% higher in the audio track allocation, and the processing time is also significantly shortened, indicating that the deep learning algorithm is more efficient in dynamically managing audio tracks; the context-aware editing score and mixing quality score of the context-aware editing and intelligent mixing are 2.5 and 2.5 points higher respectively, confirming that The superiority of the context-aware editing method and the intelligent mixing function; the ADSR parameter adjustment and compression processing in the system of the present invention scored 1.5 points higher in terms of sound quality after adjusting the ADSR parameters and compression processing, showing the improvement of the parameter adaptive adjustment and compression algorithm; the synthesis detail score of multi-track synthesis and neural network deep synthesis was 2.5 points higher, and the application of neural network models in deep synthesis greatly improved the audio detail restoration and sound quality; finally, real-time audio processing and user feedback, the real-time processing delay was reduced by 40ms, and the user feedback adjustment score was 2 points higher, indicating the optimization of the real-time audio processing method and user feedback mechanism.
[0328] In summary, the system of the present invention demonstrates significant innovation and advantages in all aspects of audio data reading, management, editing, synthesis, and processing. In particular, it provides users with a higher-quality, more personalized audio experience in terms of contextual awareness, adaptive parameter adjustment, and deep neural network synthesis. This fully demonstrates that the technical solution of the present invention represents a significant advancement and creativity compared to existing technologies, satisfying the novelty and inventiveness requirements of patent applications.
[0329] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A multi-track audio synthesis and processing method, characterized by: include, Audio data is read from the SD NAND solid-state memory via the SDIO protocol and sent to a ring buffer, where the ECC mechanism is implemented for data verification and error correction. The original audio data is read from the SD NAND solid-state memory via the SDIO protocol. The audio data is transferred to the ring buffer in the memory. During the ring buffer RB reading process, the ECC mechanism is applied for data integrity checking and error correction. ECC uses Reed-Solomon encoding. During encoding, the original data is divided into fixed-size data blocks, each of which is a polynomial coefficient. The encoder calculates the roots of the polynomial to form check symbols, forming the encoded data blocks. After receiving the data, the roots of the polynomial are recalculated and compared with the received check symbols. If there are errors in the received data, the decoder recovers the original data by solving the roots of the polynomial. Analyze the audio data in the buffer using a deep learning algorithm and dynamically allocate and manage multiple audio tracks. Use a convolutional neural network to extract the spectral features of the audio data from the ring buffer. Input the feature sequence output by the convolutional neural network into a recurrent neural network to capture the dependencies between audio data packets. The output of the recurrent neural network is further processed using a long short-term memory (LSTM) network to optimize the track selection probability and dynamically assign audio data packets to corresponding tracks in combination with an allocation function. The feature sequence extracted by the convolutional neural network (CNN) serves as the input to the recurrent neural network (RNN), with each time step corresponding to a fragment of audio data. The RNN uses a hidden state transfer mechanism to combine information from previous time steps with information from the current time step to capture the historical dependencies of the sequence data. The long short-term memory (LSTM) network uses a gating mechanism to control the flow of information and optimize the track selection probability. The convolutional neural network (CNN), the recurrent neural network (RNN), and the long short-term memory (LSTM) network ultimately generate a formula for track assignment. Automatically adjust the allocated and managed audio data using intelligent mixing functions and context-aware audio editing methods; Independently adjust ADSR parameters based on intelligently mixed and personalized edited audio data, and perform synthesis and compression processing; Adjust the ADSR parameters of each audio track to adapt to different types of audio data; adaptively adjust the calibration parameters according to the dynamic changes of audio characteristics; compress the adjusted audio tracks; ADSR parameters are adjusted as follows: Attack controls the speed at which the sound rises from silence to maximum volume when a note starts; Decay specifies how quickly the volume drops from the peak of the attack phase to the sustained volume; Sustain sets the volume level maintained before the note is released; Release defines how quickly the volume drops to silence after a note is released; Based on the synthesized and compressed audio data, the neural network model is integrated to generate multi-track audio data; Returns the synthesized audio for further processing.
2. The multi-track audio synthesis and processing method according to claim 1, wherein: The ring buffer RB read expression is: RB=E(d)+l; Where E is the ECC mechanism, d is the original audio data read from the SD NAND, and l is the data read delay value.
3. The multi-track audio synthesis and processing method according to claim 2, wherein: The track value expression is: The allocation function expression is: f(a i ,j)=ω(a i ,t j ); Among them, t j is a track value, n j is the total number of audio packets on the jth track, i is the index of all audio packets to be allocated, p j is the probability value of the audio data packet assigned to the jth audio track being selected, f(a i , j) is the distribution function, a i is the value of the i-th audio segment, j is t j The identifier of ω(a i , t j ) is the evaluation of a i With t j Function of matching degree.
4. The multi-track audio synthesis and processing method according to claim 3, wherein: The allocated and managed audio data is automatically adjusted using the intelligent mixing function and the context-aware audio editing method, specifically in the following steps: Adjust audio tracks using context-aware editing to personalize edits based on context; Introducing the context relevance value r, the smart mixing function is used to combine the edited audio track with the context relevance to generate the adjusted audio track. The expression is: t′ j =m(c(t j )·r); Smart mixing function, the expression is: m(c(t j ).r)=c(t j )+ΔP(r); The context-aware editing function c(t j ), the expression is: c(t j )=g(t j ;θ c ); Where t′ j is the adjusted track value, m is the intelligent mixing function, ΔP(r) is the parameter adjustment vector value calculated according to the context r, θ c are model parameters.
5. The multi-track audio synthesis and processing method according to claim 4, wherein: The compressed audio track value expression is: t″ j =cmp(ADSR(t′ j )+b); ADSR parameter adjustment function, the expression is: (A,D,S,R) t =adsr(z); Where t″ j Represents the audio track value after ADSR adjustment and compression, cmp is the compression function, b is the parameter adaptive adjustment value, A is the time from the start of the note to the peak volume, D is the time from the volume reaching the peak to the maintenance level, S is the stable volume value of the note in the sustained phase, R is the time from the volume starting to drop after the user releases the key until it disappears completely, t represents time, and z is the audio feature value.
6. The multi-track audio synthesis and processing method according to claim 5, wherein: The method generates multi-track audio data based on the synthesized compressed audio data and integrates the neural network model. The specific steps are as follows: Use neural network models to perform deep synthesis of compressed audio tracks; During the audio synthesis process, a dynamic weight matrix is used to adjust the importance of each audio track; Introducing time series, at each time point, audio data synthesis is performed based on the dynamic weight matrix W and the neural network model F. The expression is: Where s(t) is the synthesized audio data at time point t, n represents the total number of tracks, and w it is the time point t for track t″ j Dynamic weight value.
7. The multi-track audio synthesis and processing method according to claim 6, wherein: The synthesized audio is returned and further processed, specifically in the following steps: The final effect processing of the synthesized audio data is performed using real-time audio processing methods. The expression is: a f =ef(s(t),e); Real-time audio effect processing function ef, the expression is: ef(s(t),e)=eq(s(t),e eq ); Among them, a f is the final output audio signal value, e is an additional real-time effect parameter, eq represents the equalizer processing value, e eq is a set of parameters for the equalizer processing values; Implement dynamic range control, the expression is: a s =li(a f ,th); Among them, a s is the safe audio signal value after limiting processing, li is the limiting function, th is the threshold, and x is the input audio signal sample value; The processed audio signal value is played through the audio output interface, and a feedback mechanism is established to monitor the playback quality and status. The expression is: o t =world(a s ,d); Among them, t is the actual output audio value, ao is the audio output function, and d is the characteristic parameter value of the output device; Using the actual output audio value, the audio data is finally optimized for sound quality through post-processing functions; Collect user feedback and adjust the coefficient based on the feedback to control the impact of the feedback on the final output; The final output is the combination of optimized audio data and user feedback, expressed as: o t '=po(s(t))+k·u f (t); The post-processing function po is expressed as: po(s(t))=s(t)*h(t); Among them, t ' is the final output audio data, u f is the user feedback value, k is the feedback adjustment coefficient value, and h(t) is a filter function.
8. A multi-track audio synthesis and processing system, based on the multi-track audio synthesis and processing method according to any one of claims 1 to 7, characterized in that: Including data reading module, analysis and management module, mixing and editing module, adjustment and compression module, synthesis module, and processing module; The data reading module is used to read audio data from the memory through the SDIO protocol and use ECC to perform error correction to ensure that the data enters the buffer accurately; The analysis and management module is used to analyze audio data using deep learning algorithms, dynamically allocate and manage multiple audio tracks, and achieve intelligent audio organization; The mixing and editing module is used to automatically adjust audio tracks using intelligent mixing functions and context-aware editing to achieve an optimal listening experience; The adjustment and compression module is used to independently adjust the ADSR parameters of the audio, perform synthesis and compression processing, and optimize the audio quality and file size; The synthesis module is used to generate high-quality multi-track audio output by fusing the processed audio data using a neural network model; The processing module is used to perform final optimization processing on the synthesized audio, prepare and output a finished audio file, and ensure that it meets the playback standards.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multi-track audio synthesis and processing method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-track audio synthesis and processing method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Audio processing method and device
CN116013332A
Tactile feedback method and system for music track matching vibration and related equipment
CN116185167A
Audio processing method and related equipment
CN116709162A