Method and System for Compressing and Transmitting Audio Data in AI Recording Scenarios

The AI-based audio data compression method addresses the challenge of suboptimal compression by using deep learning and adaptive bit rate adjustment to optimize encoding quality and efficiency in diverse recording scenarios.

CN120071946BActive Publication Date: 2025-07-15SHENZHEN BEIBO INTELLIGENT TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510537660.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-15
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing audio compression methods cannot adaptively adjust the compression parameters, resulting in poor compression efficiency and effectiveness on recording devices, and cannot achieve the best results on devices with different hardware performance, storage space and network bandwidth differences.

Method used

The deep learning model is used for audio content analysis, combined with adaptive bit rate adjustment strategy and hybrid encoding algorithm, the compression parameters are dynamically adjusted, and the lame library is used to encode and compress audio data based on the hardware performance, storage space and network bandwidth of the recording device.

Benefits of technology

Improve audio encoding effect and encoding efficiency, ensure optimal compression effect in different devices and network environments, and balance the relationship between sound quality and data volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071946B_ABST
    Figure CN120071946B_ABST
Patent Text Reader

Abstract

The method and system for compressing and transmitting audio data in the AI recording scenario provided by the present invention include: in the AI recording scenario of the recording device, obtaining pure audio data; performing audio content analysis on the pure audio data based on a deep learning model to obtain an audio content analysis result; according to the audio content analysis result, combining an adaptive bit rate adjustment strategy, and using a hybrid coding algorithm to encode the pure audio data to obtain preliminary encoded audio data; using the lame library to dynamically adjust compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compressing the preliminary encoded audio data to obtain compressed audio data; wherein, the compression parameters include a compression ratio; and transmitting the compressed audio data to a target device. In the present invention, the compression parameters are dynamically adjusted to overcome the defect that the current inability to adaptively adjust the compression parameters results in poor compression efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech processing, and particularly to a method and system for audio data compression and transmission in an AI recording scenario. Background Art

[0002] In the current field of audio processing, especially in mobile devices, audio data compression has always been an important issue. Existing audio compression methods, such as MP3, AAC, etc., although can effectively reduce the size of audio files, there is often a trade-off between sound quality and compression efficiency.

[0003] In addition, due to the differences in hardware performance, storage space, and network bandwidth of different devices, the standards and methods of audio compression cannot be generally applied to recording devices, and cannot adaptively adjust compression parameters, resulting in poor compression efficiency and compression effect. Therefore, a new method is needed to optimize the audio compression effect and compression efficiency on the recording platform. Summary of the Invention

[0004] The main object of the present invention is to provide a method and system for audio data compression and transmission in an AI recording scenario, aiming to overcome the defect of poor compression efficiency caused by the inability to adaptively adjust compression parameters currently.

[0005] To achieve the above object, the present invention provides a method for audio data compression and transmission in an AI recording scenario, including the following steps:

[0006] In the AI recording scenario of the recording device, obtain pure audio data;

[0007] Based on a deep learning model, perform audio content analysis on the pure audio data to obtain an audio content analysis result;

[0008] According to the audio content analysis result, combined with an adaptive bitrate adjustment strategy, use a hybrid coding algorithm to encode the pure audio data to obtain preliminary encoded audio data;

[0009] Use the lame library to dynamically adjust compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminary encoded audio data to obtain compressed audio data; wherein, the compression parameters include the compression ratio;

[0010] Transmit the compressed audio data to the target device.

[0011] Further, obtaining pure audio data includes:

[0012] After the recording device sets recording-related parameters, record, and use a noise reduction algorithm to reduce noise of the recorded audio data to obtain pure audio data.

[0013] Further, according to the audio content analysis result, in combination with the adaptive bitrate adjustment strategy, a hybrid coding algorithm is used to encode the pure audio data to obtain preliminary encoded audio data, including:

[0014] Using the information entropy calculation method, based on the audio content analysis result, analyze the information richness of the audio in different frequency bands and time windows to obtain an audio complexity quantization value;

[0015] Based on a pre-constructed complexity-bitrate mapping table, match the target bitrate mapped to the audio complexity quantization value as the adaptively adjusted bitrate;

[0016] From the lossless coding, lossy coding, and hybrid coding modes, dynamically match the most suitable coding mode and parameter combination according to the target bitrate and the audio complexity quantization value to obtain the optimal coding mode and parameter set;

[0017] Divide the pure audio data into multiple data blocks, and based on the optimal coding mode and parameter set, perform adaptive coding on each data block to obtain preliminary encoded audio data.

[0018] Further, according to the audio content analysis result, in combination with the adaptive bitrate adjustment strategy, a hybrid coding algorithm is used to encode the pure audio data to obtain preliminary encoded audio data, including:

[0019] Based on the audio content analysis result, divide the pure audio data into multiple time segments, and generate an audio feature vector for each time segment;

[0020] Using a reinforcement learning model, with each audio feature vector as the input, the adaptive bitrate adjustment strategy as the action space, and the comprehensive evaluation of the encoded audio quality and the encoded data volume as the reward function, dynamically select the optimal bitrate for each time segment;

[0021] Adopt a hybrid coding algorithm, and dynamically combine different coding sub-algorithms according to the spectral characteristics of each time segment;

[0022] Encode each time segment based on the selected optimal bitrate and the dynamically combined coding sub-algorithms, and merge the encoding results of each time segment to obtain preliminary encoded audio data.

[0023] Further, use the lame library to dynamically adjust the compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminary encoded audio data to obtain compressed audio data, including:

[0024] Establish a multi-dimensional state matrix for hardware performance, storage space, and network bandwidth, and store the optimal compression parameters in different states in the corresponding matrix elements;

[0025] Obtain the current hardware performance, storage space, and network bandwidth information of the recording device, and determine its position in the multi-dimensional state matrix;

[0026] When it is detected that the hardware performance, storage space, or network bandwidth changes, perform dynamic search in the multi-dimensional state matrix according to the change trend and amplitude to obtain multiple candidate compression parameters;

[0027] Assign different fuzzy membership degrees to each candidate compression parameter, calculate the comprehensive score of each candidate compression parameter according to the fuzzy inference rules, and select the one with the highest score as the final compression parameter;

[0028] Use the lame library to compress the preliminary encoded audio data according to the final compression parameter to obtain compressed audio data.

[0029] Furthermore, use the lame library to dynamically adjust the compression parameter according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminary encoded audio data to obtain compressed audio data, including:

[0030] Obtain the hardware performance data, storage space usage, and network bandwidth change data of the recording device within a preset time period before the current time, and construct a time series data set;

[0031] Input the time series data set into a long short-term memory network model for analysis to predict the hardware performance trend, remaining storage space, and network bandwidth fluctuation range of the recording device in the future for a period of time;

[0032] According to the prediction results, plan the compression parameters of the lame library in advance;

[0033] During the compression operation, obtain the real-time monitored hardware performance, storage space, and network bandwidth data, perform real-time correction on the pre-planned compression parameters, and use the lame library to compress the preliminary encoded audio data according to the corrected compression parameters to obtain compressed audio data.

[0034] Furthermore, after transmitting the compressed audio data to the target device, it includes:

[0035] The target device obtains the audio characteristic curve and audio attribute information of the compressed audio data;

[0036] Obtain the preset coding table in the database, and modify the preset coding table based on the audio characteristic curve to obtain a modified coding table;

[0037] Encode the audio attribute information based on the changed coding table to obtain an attribute code;

[0038] Generate a management key based on the attribute code to perform permission management on the compressed audio data.

[0039] Further, change the preset coding table based on the audio characteristic curve to obtain a changed coding table, including:

[0040] Move the audio characteristic curve into a preset coordinate system, and the starting point of the audio characteristic curve coincides with the origin of the coordinate system;

[0041] Extract the coding characters in the coding character column of the preset coding table and combine them into a character sequence in sequence; wherein, the preset coding table includes a coding character column and an original data column that are mapped one-to-one;

[0042] Tile the character sequence into the preset coordinate system according to a preset rule, and detect the positional relationship between each coding character and the audio characteristic curve; reorder each coding character according to the positional relationship;

[0043] Add each reordered coding character to the coding character column of the preset coding table one by one in sequence to obtain a changed coding table.

[0044] The present invention also provides a system for compressing and transmitting audio data in an AI recording scenario, including:

[0045] An acquisition module, configured to acquire pure audio data in the AI recording scenario of a recording device;

[0046] An analysis module, configured to perform audio content analysis on the pure audio data based on a deep learning model to obtain an audio content analysis result;

[0047] An encoding module, configured to encode the pure audio data by using a hybrid encoding algorithm according to the audio content analysis result in combination with an adaptive bitrate adjustment strategy to obtain preliminarily encoded audio data;

[0048] A compression module, configured to use the lame library to dynamically adjust compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminarily encoded audio data to obtain compressed audio data; wherein, the compression parameters include a compression ratio;

[0049] A transmission module, configured to transmit the compressed audio data to a target device.

[0050] The present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of any one of the above methods are implemented.

[0051] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0052] The method and system for compressing and transmitting audio data in the AI recording scenario provided by the present invention include: in the AI recording scenario of a recording device, obtaining pure audio data; performing audio content analysis on the pure audio data based on a deep learning model to obtain an audio content analysis result; according to the audio content analysis result, combining an adaptive bit rate adjustment strategy, and using a hybrid coding algorithm to encode the pure audio data to obtain preliminary encoded audio data; using the lame library to dynamically adjust compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compressing the preliminary encoded audio data to obtain compressed audio data; wherein the compression parameters include a compression ratio; and transmitting the compressed audio data to a target device. In the present invention, according to the audio content analysis result, combining an adaptive bit rate adjustment strategy, and using a hybrid coding algorithm to encode the pure audio data, the audio coding effect and coding efficiency are improved; at the same time, the compression parameters are dynamically adjusted according to the hardware performance, storage space, and network bandwidth of the recording device, overcoming the defect of poor compression efficiency caused by the current inability to adaptively adjust compression parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic diagram of the steps of the method for compressing and transmitting audio data in the AI recording scenario in an embodiment of the present invention;

[0054] Figure 2 is a block diagram of the system structure for compressing and transmitting audio data in the AI recording scenario in an embodiment of the present invention;

[0055] Figure 3 is a schematic block diagram of a computer device in an embodiment of the present invention.

[0056] The implementation, features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] Refer toFigure 1 , in an embodiment of the present invention, a method for compressing and transmitting audio data in an AI recording scenario is provided, including the following steps:

[0059] Step S1, in the AI recording scenario of the recording device, obtain pure audio data;

[0060] Step S2, based on a deep learning model, perform audio content analysis on the pure audio data to obtain an audio content analysis result;

[0061] Step S3, according to the audio content analysis result, combined with an adaptive bitrate adjustment strategy, use a hybrid coding algorithm to encode the pure audio data to obtain preliminarily encoded audio data;

[0062] Step S4, use the lame library to dynamically adjust compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminarily encoded audio data to obtain compressed audio data; wherein, the compression parameters include the compression ratio;

[0063] Step S5, transmit the compressed audio data to the target device.

[0064] In this embodiment, as described in step S1 above, in the AI recording scenario, there are usually various noise interferences, such as environmental noise, noise generated by the device itself, etc. In order to process audio data more accurately later, it is necessary to first obtain pure audio data. For example, by using hardware noise reduction technology, a special noise reduction microphone can be installed on the recording device, which can preliminarily filter the noise while collecting audio; software noise reduction algorithms can also be used to process the collected raw audio data. By analyzing the characteristics of the noise, it can be separated and removed from the audio to obtain pure audio data. Pure audio data is the basis for subsequent audio content analysis and encoding processing. High-quality pure audio can improve the accuracy and effectiveness of subsequent processing steps.

[0065] As described in step S2 above: Different audio contents have different characteristics. By performing content analysis on the pure audio data, the specific information of the audio can be understood, such as whether it contains speech, the semantic content of the speech, the emotional tendency of the audio, whether there are specific sound effects, etc., providing a basis for subsequent encoding and transmission. Construct a deep learning model, such as a recurrent neural network (RNN), a long short-term memory network (LSTM), or a convolutional neural network (CNN), etc. Take the pure audio data as input, and through the training and learning of the model, enable the model to recognize various features and patterns in the audio. For example, for speech audio, the model can perform speech recognition and convert the speech into text; for music audio, the model can analyze its melody, rhythm, instrument type, etc. The audio content analysis results can help the subsequent steps to adopt more appropriate encoding and transmission strategies according to the characteristics of different audio contents, so as to improve the efficiency and quality of audio processing.

[0066] As described in step S3 above, different audio contents have different requirements for audio quality, and there are also differences in network bandwidth and storage resources. Therefore, it is necessary to dynamically adjust the encoding bit rate according to the audio content analysis results and adopt appropriate encoding algorithms to minimize the data volume while ensuring the audio quality. First, judge factors such as the importance and complexity of the audio according to the audio content analysis results. For example, for the speech part containing important information, the bit rate can be appropriately increased to ensure clear audio quality; for relatively less important parts such as background music, the bit rate can be reduced to reduce the data volume. At the same time, combine the current network bandwidth and storage situation to dynamically adjust the bit rate.

[0067] Integrate the advantages of multiple encoding algorithms, such as lossless encoding algorithms (such as FLAC) and lossy encoding algorithms (such as MP3). Use lossless encoding for the audio parts that need to retain complete information, and use lossy encoding for the parts where the audio quality can be sacrificed appropriately in exchange for a smaller data volume. Through adaptive bit rate adjustment and hybrid encoding algorithms, it is possible to effectively reduce the data volume after encoding while ensuring the audio quality, and improve the efficiency of audio transmission and storage.

[0068] As described in step S4 above, the hardware performance of the recording device varies, resulting in differences in its processing and storage capabilities. At the same time, the network bandwidth also changes at any time. To ensure that audio data can be efficiently processed and transmitted on the device, it is necessary to dynamically adjust the compression parameters according to these factors. The lame library is a commonly used MP3 encoding library with high compression efficiency and sound quality performance. By real-time monitoring the hardware performance of the recording device (such as CPU usage, memory occupancy, etc.), the size of the storage space, and the current network bandwidth, the compression parameters, such as the compression ratio, are dynamically adjusted according to preset rules or algorithms. For example, when the device hardware performance is low or the network bandwidth is small, the compression ratio is increased to reduce the data volume; when the device performance is good and the network bandwidth is sufficient, the compression ratio is appropriately reduced to improve the sound quality. Dynamically adjusting the compression parameters can enable the audio data to achieve the best compression effect in different device and network environments, and balance the relationship between sound quality and data volume.

[0069] As described in step S5 above, after the audio data is compressed, it needs to be transmitted to the target device for subsequent playback, storage, or further processing. According to the target device and network environment, select the appropriate transmission protocol and method. For example, if it is transmitted within a local area network, the TCP / IP protocol can be used for stable transmission; if it is transmitted through the Internet, protocols such as HTTP and FTP can be adopted. At the same time, to ensure the reliability and security of the transmission, data encryption, error checking, etc. may also be required. The compressed audio data is successfully transmitted to the target device to achieve the sharing and use of the audio data.

[0070] In one embodiment, obtaining pure audio data includes:

[0071] After the recording device sets the recording-related parameters, it records, and a noise reduction algorithm is used to reduce the noise of the recorded audio data to obtain pure audio data.

[0072] In one embodiment, according to the audio content analysis result, combined with the adaptive bitrate adjustment strategy, a hybrid coding algorithm is used to encode the pure audio data to obtain preliminary encoded audio data, including:

[0073] Using the information entropy calculation method, based on the audio content analysis result, analyze the information richness of the audio in different frequency bands and time windows to obtain the audio complexity quantization value;

[0074] Based on the pre-constructed complexity-bitrate mapping table, match the target bitrate mapped to the audio complexity quantization value as the adaptively adjusted bitrate;

[0075] From lossless coding, lossy coding, and hybrid coding modes, dynamically match the most suitable coding mode and parameter combination according to the target bit rate and the audio complexity quantization value to obtain the optimal coding mode and parameter set;

[0076] Divide the pure audio data into multiple data blocks, and perform adaptive coding on each data block based on the optimal coding mode and parameter set to obtain the preliminary coded audio data.

[0077] In this embodiment, the information distribution of different audio contents is uneven in different frequency bands and time. For example, the high-frequency and low-frequency parts in music, and the pauses and speaking parts in speech have different degrees of information richness. To perform more accurate coding, it is necessary to quantify the complexity of the audio. Information entropy is an index to measure information uncertainty and richness. In audio processing, based on the audio content analysis results, it is divided according to different frequency bands (such as low frequency, medium frequency, high frequency) and time windows (the audio is divided into several small segments by time). For the audio data within each frequency band and time window, calculate its information entropy. The larger the information entropy, the richer the audio information in this area and the higher the complexity; conversely, the lower the complexity. In this way, the information richness of the audio in different frequency bands and time windows is converted into a specific quantization value, that is, the audio complexity quantization value. The above quantization value provides a key basis for the subsequent adaptive bit rate adjustment and coding mode selection, enabling the coding process to be optimized according to the actual complexity of the audio.

[0078] Audios with different complexities require different bit rates to ensure appropriate sound quality and coding efficiency. If the bit rate is set too high, it will result in too large a data volume, wasting storage space and transmission bandwidth; if the bit rate is set too low, the audio quality will decline. Therefore, it is necessary to establish the correspondence between complexity and bit rate. In this embodiment, a complexity-bit rate mapping table is pre-constructed, which is obtained through experiments and analyses on a large number of audio samples with different complexities. The table records the optimal bit rates corresponding to different audio complexity quantization values. When the complexity quantization value of the current audio is obtained, the target bit rate matching it is searched in the mapping table and used as the bit rate for adaptive adjustment. In the above way, the bit rate can be dynamically adjusted according to the actual complexity of the audio, effectively controlling the data volume while ensuring the audio quality and improving the coding efficiency.

[0079] In this embodiment, the above lossless encoding can completely retain the original information of the audio, but usually generates a large amount of data; lossy encoding sacrifices a certain audio quality in exchange for a smaller amount of data; hybrid encoding combines the advantages of both. Different target bitrates and audio complexities require different encoding modes and parameter combinations to achieve the best encoding effect. According to the target bitrate and the quantization value of audio complexity, a set of matching rules is established. For example, when the target bitrate is high and the audio complexity is low, the lossless encoding mode may be selected to ensure high quality; when the target bitrate is low and the audio complexity is high, the lossy encoding mode may be selected to reduce the amount of data; when the situation is more complex, the hybrid encoding mode may be selected, and different encoding methods are used for different parts of the audio. At the same time, for each encoding mode, specific parameters need to be determined, such as the sampling rate of encoding, the specific parameters of the encoding algorithm, etc. Through dynamic matching, the most suitable encoding mode and parameter combination, that is, the optimal encoding mode and parameter set, are found. In this embodiment, the most appropriate encoding mode and parameters can be selected according to the actual situation of the audio and the target bitrate, achieving the best balance between audio quality and data volume.

[0080] Audio data usually has a certain degree of continuity and variability. Dividing it into multiple data blocks can make the encoding more flexible. Different encoding strategies are adopted according to the characteristics of each data block. The pure audio data is divided into multiple data blocks of equal or unequal sizes according to certain rules. Each data block can be regarded as a relatively independent audio segment. Then, according to the optimal encoding mode and parameter set obtained previously, each data block is encoded separately. During the encoding process, each data block can be adaptively adjusted according to its own characteristics and the overall encoding requirements to achieve the best encoding effect. The above-mentioned block encoding method improves the flexibility and adaptability of encoding, can better handle the changes in audio data, further optimizes the encoding result, and obtains the preliminarily encoded audio data.

[0081] In one embodiment, according to the audio content analysis result, combined with the adaptive bitrate adjustment strategy, a hybrid encoding algorithm is used to encode the pure audio data to obtain the preliminarily encoded audio data, including:

[0082] Based on the audio content analysis result, the pure audio data is divided into multiple time segments, and an audio feature vector is generated for each time segment;

[0083] Using a reinforcement learning model, with each audio feature vector as the input, the adaptive bitrate adjustment strategy as the action space, and the comprehensive evaluation of the encoded audio quality and encoded data volume as the reward function, the optimal bitrate for each time segment is dynamically selected;

[0084] Adopt a hybrid coding algorithm, and dynamically combine different coding sub-algorithms according to the spectral characteristics of each time segment;

[0085] Based on the selected optimal bitrate and the dynamically combined coding sub-algorithms, each time segment is encoded, and the encoding results of each time segment are merged to obtain the preliminary encoded audio data.

[0086] In this embodiment, there are significant differences in the characteristics and content of audio in different time periods. For example, a voice audio may contain pauses, different intonations, etc.; a music audio may have rhythm changes, instrument performance switches, etc. Dividing the audio by time and extracting feature vectors helps to analyze and process the audio more meticulously. According to the results of audio content analysis, such as silent segments in the audio, semantic pauses in speech, rhythm change points in music, etc., the pure audio data is segmented into multiple time segments. For each time segment, a series of audio features are extracted, such as frequency, energy, pitch, timbre, etc., and these features are combined into a vector, namely the audio feature vector. These features can reflect the characteristics of the audio within this time segment. By dividing the audio into time segments and generating feature vectors, it provides a basis for subsequent precise adjustment of coding parameters for each segment, and can more flexibly adapt to changes in audio content.

[0087] In audio coding, it is necessary to minimize the amount of encoded data while ensuring the audio quality. Different audio contents require different bitrates to achieve this goal. The reinforcement learning model can continuously learn and optimize strategies based on environmental feedback, and is suitable for dynamically selecting the optimal bitrate.

[0088] Specifically, the audio feature vector of each time segment is input into the reinforcement learning model. The different bitrate options included in the adaptive bitrate adjustment strategy constitute the action space, and the model can select an action (i.e., a bitrate) from it. The reward function comprehensively considers the quality of the encoded audio (measured by indicators such as signal-to-noise ratio, perceptual audio quality assessment, etc.) and the amount of encoded data. If the selected bitrate results in high audio quality and small data volume, the model will obtain a higher reward; otherwise, it will obtain a lower reward. By continuously trying different actions and learning based on the rewards, the model will gradually find the optimal bitrate for each time segment. Using the reinforcement learning model to dynamically select the optimal bitrate can achieve more precise bitrate adjustment according to the actual characteristics of the audio and the comprehensive requirements for quality and data volume, improving the coding efficiency and quality.

[0089] Different encoding sub - algorithms are applicable to audio with different spectral characteristics. For example, some encoding algorithms perform better in processing low - frequency signals, while others are more suitable for high - frequency signals. The hybrid encoding algorithm combines the advantages of multiple encoding sub - algorithms and can better adapt to the spectral changes of audio. Specifically, perform spectral analysis on the audio of each time segment to understand its spectral characteristics, such as the main frequency components, frequency distribution range, etc. According to these spectral characteristics, select appropriate algorithms from a predefined set of encoding sub - algorithms for combination. For example, for a time segment with more low - frequency components, select an encoding sub - algorithm that is good at processing low - frequency; for a segment with rich high - frequency components, combine an algorithm suitable for high - frequency processing. Through dynamic combination, form an optimal encoding scheme for each time segment. Dynamically combining encoding sub - algorithms according to the spectral characteristics of the audio can give full play to the advantages of different encoding algorithms, improve the pertinence and effect of encoding, and further optimize the balance between audio quality and data volume.

[0090] After determining the optimal bit - rate and suitable encoding sub - algorithms for each time segment, it is necessary to perform actual encoding operations on each segment and merge the encoding results into a complete audio data. According to the selected optimal bit - rate and the dynamically combined encoding sub - algorithms, encode the audio data of each time segment. After encoding, merge the encoding results of each time segment in the time order of the original audio to form a continuous audio data stream, that is, the preliminary encoded audio data. By performing targeted encoding on each time segment and merging the results, finally obtain the preliminary encoded audio data that meets the requirements. While ensuring the audio quality, effectively control the data volume, providing a good basis for subsequent compression and transmission.

[0091] In one embodiment, use the lame library to dynamically adjust the compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminary encoded audio data to obtain compressed audio data, including:

[0092] Establish a multi - dimensional state matrix of hardware performance, storage space, and network bandwidth, and store the optimal compression parameters in different states in the corresponding matrix elements;

[0093] Obtain the current hardware performance, storage space, and network bandwidth information of the recording device, and determine its position in the multi - dimensional state matrix;

[0094] When it is detected that the hardware performance, storage space, or network bandwidth changes, perform dynamic search in the multi - dimensional state matrix according to the trend and amplitude of the change to obtain multiple candidate compression parameters;

[0095] Assign different fuzzy membership degrees to each candidate compression parameter. According to the fuzzy inference rules, calculate the comprehensive score of each candidate compression parameter, and select the one with the highest score as the final compression parameter;

[0096] Using the lame library, compress the preliminary encoded audio data according to the final compression parameter to obtain compressed audio data.

[0097] In this embodiment, the hardware performance, storage space, and network bandwidth of the recording device are dynamically changing, and different state combinations have different requirements for audio compression parameters. To quickly find the optimal compression parameter suitable for the current device state, a comprehensive mapping relationship needs to be established. Specifically, the hardware performance, storage space, and network bandwidth are respectively divided into multiple levels or intervals to construct a three-dimensional multi-dimensional state matrix. For example, the hardware performance can be divided into three levels: high, medium, and low according to indicators such as CPU usage rate and memory occupancy; the storage space can be divided into multiple intervals according to the percentage of the remaining capacity; the network bandwidth can be divided into different levels according to the upload and download speeds. For each state combination (i.e., an element in the matrix), through a large number of experiments and tests, determine the corresponding optimal compression parameters (such as compression ratio, sampling rate, bit rate, etc.), and store these parameters in the corresponding positions of the matrix. This multi-dimensional state matrix provides a basis for quickly finding suitable compression parameters according to the current device state, avoiding complex calculations and evaluations each time.

[0098] To select a suitable compression parameter for the current recording device, it is necessary to understand its current hardware performance, storage space, and network bandwidth status. Specifically, use the interfaces or tools provided by the system to obtain the hardware performance indicators (such as CPU usage rate, remaining memory, etc.), the remaining capacity of the storage space, and the current network bandwidth speed of the recording device in real time. Then, according to the pre-set level or interval division criteria, determine the corresponding levels or intervals of this information, so as to determine the specific position of the device's current state in the multi-dimensional state matrix. Clearly defining the position of the device's current state in the matrix provides an accurate positioning for subsequent searching and adjusting of compression parameters.

[0099] During the recording and compression process, the hardware performance, storage space, and network bandwidth of the device may change. For example, the device runs other resource-consuming programs simultaneously, resulting in a decrease in hardware performance, or the network environment is unstable, causing bandwidth fluctuations. To adapt to these changes, it is necessary to adjust the compression parameters in a timely manner. Specifically, continuously monitor the changes in hardware performance, storage space, and network bandwidth. When any of these indicators is detected to change, analyze the trend (such as rising or falling) and amplitude (such as the percentage of change) of the change. Based on this information, conduct a dynamic search in the multi-dimensional state matrix. For example, if the hardware performance decreases significantly, it may search in the area of lower hardware performance in the matrix; if the network bandwidth suddenly increases, it may search in the area of higher bandwidth. Through this search method, multiple candidate compression parameters that may be suitable for the current change state are found. Dynamically searching for candidate compression parameters ensures that appropriate parameter options can be found in a timely manner when the device state changes, providing multiple possibilities for subsequent selection of the final compression parameters.

[0100] Due to the certain uncertainty in the changes of hardware performance, storage space, and network bandwidth, and the relatively complex impact of different compression parameters on audio quality and compression effect, using fuzzy logic can handle this uncertainty more flexibly. Assign a fuzzy membership degree to each candidate compression parameter, which represents the degree of matching of this parameter with the current device state and compression requirements. The determination of the fuzzy membership degree can be based on experience, experimental data, or expert knowledge. Then, according to the predefined fuzzy inference rules, comprehensively consider multiple factors such as hardware performance, storage space, network bandwidth, and audio quality, and calculate the comprehensive score of each candidate compression parameter. Finally, select the candidate compression parameter with the highest comprehensive score as the final compression parameter. Selecting the final compression parameter through fuzzy inference can more comprehensively and flexibly consider various uncertain factors, improve the accuracy and rationality of the selection, and ensure good compression effects under different device states.

[0101] After determining the final compression parameters, it is necessary to use appropriate tools to compress the initially encoded audio data to reduce the data volume for easy storage and transmission. Specifically, the above-mentioned lame library is a powerful audio coding library with high compression efficiency and good sound quality performance. Pass the finally determined compression parameters (such as compression ratio, sampling rate, etc.) to the corresponding interface of the lame library, and call the library function to compress the initially encoded audio data. After compression, smaller-sized compressed audio data is obtained. Using the lame library and the final compression parameters to compress the audio data achieves the purpose of effectively reducing the data volume while ensuring a certain audio quality, providing convenience for subsequent audio transmission and storage.

[0102] In one embodiment, using the lame library, the compression parameters are dynamically adjusted according to the hardware performance, storage space, and network bandwidth of the recording device, and the preliminary encoded audio data is compressed to obtain compressed audio data, including:

[0103] Obtain the hardware performance data, storage space usage, and network bandwidth change data of the recording device within a preset time period before the current time, and construct a time series data set;

[0104] Input the time series data set into a long short-term memory network model for analysis to predict the hardware performance trend, remaining storage space, and network bandwidth fluctuation range of the recording device within a future period of time;

[0105] According to the predicted results, plan the compression parameters of the lame library in advance;

[0106] During the compression operation, obtain the real-time monitored hardware performance, storage space, and network bandwidth data, make real-time corrections to the pre-planned compression parameters, and use the lame library to compress the preliminary encoded audio data according to the corrected compression parameters to obtain compressed audio data.

[0107] In this embodiment, the hardware performance, storage space, and network bandwidth of the recording device are not constant, but change dynamically over time. In order to accurately predict the future state of the device, it is necessary to collect relevant data over a period of time to capture its change pattern. Specifically, with the help of the system's built-in monitoring tools, API interfaces, or specialized hardware monitoring software, within a preset time period before the current time (such as the past 1 hour, 6 hours, etc., and the specific duration can be set according to the actual situation), regularly collect hardware performance data (such as CPU usage rate, memory occupancy rate, etc.), storage space usage (remaining storage space size), and network bandwidth change data (upload and download bandwidth). Combine these data arranged in chronological order to form a time series data set. The time series data set is the basis for subsequent prediction analysis. It contains information on the device state changing over time, can reflect the usage pattern and trend of the device, and provides a basis for accurately predicting the future state.

[0108] Long short-term memory network (LSTM) is a special recurrent neural network that can effectively handle long-term dependencies in sequence data and is suitable for predicting data with time correlation. By learning from historical data, LSTM can capture the law of device state changes and thus predict the situation in the future. Specifically, first, the constructed time series data set is preprocessed, such as normalization, to ensure the consistency of the data scale and improve the training effect of the model. Then, the preprocessed data set is input into the pre-trained LSTM model. After a lot of training, the model has learned the patterns and laws of hardware performance, storage space, and network bandwidth changes. Based on the input historical data, the model will output the hardware performance trend (such as whether the CP utilization rate is rising or falling), the remaining storage space (the approximate range of the remaining storage space), and the network bandwidth fluctuation range (the maximum and minimum bandwidth values) of the recording device in the future (for example, the next 30 minutes, 1 hour, etc., the specific duration is related to business needs). Through the prediction of the LSTM model, the future state of the device can be understood in advance, providing a basis for planning compression parameters in advance, and avoiding poor compression effect or errors due to sudden changes in device status during the compression process.

[0109] Different hardware performance, storage space and network bandwidth conditions have different requirements for audio compression parameters. For example, when the hardware performance is low or the network bandwidth is narrow, a higher compression ratio needs to be selected to reduce the amount of data; when the hardware performance is good and the network bandwidth is sufficient, a lower compression ratio can be selected to ensure audio quality. Therefore, planning the compression parameters in advance according to the predicted future state of the device can make the compression process more efficient and reasonable. Specifically, according to the hardware performance trend, storage space remaining situation and network bandwidth fluctuation range predicted by the LSTM model in the future, combined with pre-set rules or empirical data, determine the appropriate compression parameters. For example, if it is predicted that the network bandwidth will become narrower in the future, the compression ratio of lame will be appropriately increased; if it is predicted that the hardware performance will improve, the compression ratio can be appropriately reduced to improve the audio quality. These determined compression parameters are used as the result of advance planning to prepare for subsequent compression operations. Planning the compression parameters in advance can make the compression process better adapt to the future state of the device, reduce the amount of data as much as possible while ensuring audio quality, and improve compression efficiency and transmission stability.

[0110] Although the compression parameters are planned in advance through prediction, there may be certain deviations between the actual situation and the prediction results. To ensure the accuracy and adaptability of the compression process, it is necessary to monitor the status of the device in real time during the actual compression operation and correct the parameters planned in advance. Specifically, when starting to use the lame library to compress the preliminarily encoded audio data, the current hardware performance, storage space, and network bandwidth data of the recording device are obtained in real time. These real-time data are compared with the previously predicted results. If a large difference is found, the compression parameters planned in advance are corrected according to the real-time data and the preset adjustment rules. For example, if the real-time monitored network bandwidth is wider than predicted, the compression ratio can be appropriately reduced to improve the audio quality. Finally, the lame library is used to compress the preliminarily encoded audio data according to the corrected compression parameters to obtain the final compressed audio data. Correcting the compression parameters in real time can make the compression process better adapt to the actual status of the device, improve the compression effect and audio quality, and ensure optimal compression and transmission in various situations.

[0111] In one embodiment, after transmitting the compressed audio data to the target device, it includes:

[0112] The target device obtains the audio characteristic curve and audio attribute information of the compressed audio data;

[0113] Obtain the preset coding table in the database, and change the preset coding table based on the audio characteristic curve to obtain a changed coding table;

[0114] Encode the audio attribute information based on the changed coding table to obtain an attribute code;

[0115] Generate a management key based on the attribute code to perform permission management on the compressed audio data.

[0116] In this embodiment, when further processing and managing the compressed audio data, it is necessary to understand its audio characteristics and related attributes. The audio characteristic curve can reflect the characteristic changes of the audio in dimensions such as frequency and time, while the audio attribute information includes basic information such as the duration, sampling rate, number of channels, and encoding format of the audio. This information is crucial for subsequent operations such as encoding and permission management. Specifically, after receiving the compressed audio data, the target device first parses it. For the audio characteristic curve, specific audio processing algorithms, such as spectral analysis and time-domain analysis, are required to extract the characteristic values of the audio at different frequencies and time points from the compressed audio data, and then draw the audio characteristic curve. The audio attribute information is usually included in the file header or metadata of the compressed audio data, and the target device can directly read this information from the corresponding location. Obtaining the audio characteristic curve and audio attribute information provides the basic data for subsequent encoding and permission management, enabling subsequent operations to be accurately processed according to the actual characteristics and attributes of the audio.

[0117] The preset encoding table is a predefined set of rules for encoding audio attribute information. However, different audios have different characteristics, and using a fixed encoding table may not fully reflect the uniqueness of the audio. Therefore, it is necessary to adjust the preset encoding table according to the audio characteristic curve to generate an encoding table that is more suitable for the current audio. Specifically, the target device obtains the preset encoding table from a local or remote database. This encoding table usually contains the mapping relationships between the possible values of various audio attributes and the encoding values. Then, analyze the audio characteristic curve, for example, judge the unique characteristics and change rules of the audio according to the characteristics such as the shape, peak value, and valley value of the curve. Based on these analysis results, modify, add, or delete the mapping relationships in the preset encoding table. For example, if the audio characteristic curve shows obvious peak values in a specific frequency range, the encoding rules for the audio attributes related to this frequency may be adjusted to obtain the changed encoding table. By changing the preset encoding table according to the audio characteristic curve, the encoding table can better adapt to the actual characteristics of the audio, improve the accuracy and pertinence of encoding, and lay the foundation for generating more effective attribute encodings in the future.

[0118] Encoding the audio attribute information is to convert it into a form that is convenient for storage, transmission, and management. Using the changed encoding table for encoding can ensure that the encoding result can reflect the actual characteristics and attributes of the audio. Specifically, the previously obtained audio attribute information (such as duration, sampling rate, number of channels, etc.) is converted according to the mapping rules in the changed encoding table. For example, if the changed encoding table stipulates that the encoding value corresponding to the audio duration within a certain range is a specific combination of numbers or characters, then the actual audio duration is converted into the corresponding encoding value according to this rule. Encoding all the audio attribute information in turn, a set of attribute encodings containing all attribute encodings is finally obtained. The attribute encoding represents the audio attribute information in a standardized form, which is convenient for subsequent unified management and processing, and also provides the necessary input for generating the management key.

[0119] Specifically, in order to protect the security and copyright of the compressed audio data, it is necessary to perform permission management on it. The management key is the key to implementing permission management. Different management keys can control the access, use, and dissemination permissions of different users or devices to the compressed audio data. Specifically, a specific key generation algorithm is used, and the attribute encoding is used as an input parameter. The above key generation algorithm can be a symmetric encryption algorithm, an asymmetric encryption algorithm, or other custom key generation algorithms. The algorithm generates a unique management key according to the content of the attribute encoding and its own rules. The management key is associated with the compressed audio data. In subsequent access and use processes, only users or devices holding the correct management key can perform corresponding operations on the compressed audio data. For example, different permission levels, such as read-only, modifiable, and disseminateable, can be set, and these permissions are assigned according to different management keys. By generating the management key and performing permission management, the security and copyright of the compressed audio data can be effectively protected, unauthorized access and use can be prevented, and the audio data can be ensured to be used and disseminated within a legal range.

[0120] In one embodiment, changing the preset encoding table based on the audio characteristic curve to obtain a changed encoding table includes:

[0121] Moving the audio characteristic curve to a preset coordinate system, and the starting point of the audio characteristic curve coincides with the origin of the coordinate system;

[0122] Extracting the encoding characters in the encoding character column of the preset encoding table and sequentially combining them into a character sequence; wherein, the preset encoding table includes an encoding character column and an original data column that are mapped one by one;

[0123] Tiling the character sequence onto the preset coordinate system according to a preset rule, detecting the positional relationship between each encoding character and the audio characteristic curve; reordering each encoding character according to the positional relationship;

[0124] Add each of the re - sorted coded characters to the coded character column of the preset coding table one by one in sequence to obtain a modified coding table.

[0125] In this embodiment, to facilitate the subsequent analysis of the positional relationship between the audio characteristic curve and the coded characters, it is necessary to place the audio characteristic curve in a unified and quantifiable space for processing. The preset coordinate system provides such a standardized space, and making the starting point of the curve coincide with the origin can simplify the subsequent position detection and analysis process. First, determine the specifications of the preset coordinate system, including the meaning of the coordinate axes (for example, the horizontal axis may represent time or frequency, and the vertical axis may represent the amplitude of the audio, etc.), the coordinate range, and the scale, etc. Then, according to the original coordinates of the starting point of the audio characteristic curve, by means of coordinate translation, move the entire audio characteristic curve into the preset coordinate system so that its starting point coincides with the origin (0, 0) of the coordinate system. Placing the audio characteristic curve in the preset coordinate system and making its starting point coincide with the origin provides a unified spatial basis for accurately detecting the positional relationship between the coded characters and the audio characteristic curve, which helps to improve the accuracy and consistency of position detection.

[0126] The preset coding table is a table that maps the original data to the coded characters one by one. In order to re - sort the coded characters according to the audio characteristic curve subsequently, it is necessary to first extract the coded characters from the coded character column and combine them into a continuous character sequence. Traverse the coded character column of the preset coding table and extract each coded character in each cell in sequence according to the order in the table. Connect the extracted coded characters in the extraction order to form a character sequence. For example, if the coded characters in the coded character column are 'A', 'B', 'C' in sequence, then the combined character sequence is 'ABC'. Combining the coded characters into a character sequence facilitates the subsequent tiling operation in the preset coordinate system, so as to be able to detect the positional relationship with the audio characteristic curve, laying a foundation for re - sorting the coded characters.

[0127] By tiling the character sequence in a preset coordinate system and detecting its positional relationship with the audio characteristic curve, the characteristics of the audio characteristic curve can be used to reorder the encoded characters, so that the changed encoding table can better reflect the characteristics of the audio. The above preset rule can be to place each encoded character in the character sequence successively at a certain interval on the horizontal or vertical axis of the coordinate system. For example, an encoded character can be placed at a certain unit length interval on the horizontal axis. For each encoded character tiled in the coordinate system, its relative positional relationship with the audio characteristic curve is detected. The positional relationship can include situations where the encoded character is above the curve, below the curve, intersecting the curve, etc. The positional relationship can be determined by comparing the ordinate of the coordinate point where the encoded character is located with the ordinate of the audio characteristic curve at this abscissa. According to the detected positional relationship, corresponding sorting rules are formulated. For example, it can be stipulated that the encoded characters located above the curve are arranged in the front, and the encoded characters located below the curve are arranged in the back; or they are sorted according to the distance from the curve. The encoded characters are reordered according to these rules. Reordering the encoded characters according to the audio characteristic curve enables the changed encoding table to be associated with the characteristics of the audio, improving the adaptability and pertinence of the encoding table to audio data.

[0128] After completing the reordering of the encoded characters, it is necessary to refill the sorted encoded characters into the encoded character column of the preset encoding table to generate the changed encoding table. In the order of reordering, each encoded character is successively added to the encoded character column of the preset encoding table to replace the original order of the encoded characters. The original data column remains unchanged because changing the encoding table mainly adjusts the order of the encoded characters without changing the mapping relationship between the original data and the encoded characters. Adding the reordered encoded characters to the encoded character column finally results in the changed encoding table. This changed encoding table can better reflect the characteristics of the audio and provides more unique encoding rules for subsequent encoding of audio attribute information based on this encoding table.

[0129] In one embodiment, changing the preset encoding table based on the audio characteristic curve to obtain a changed encoding table includes:

[0130] Moving the audio characteristic curve into a preset coordinate system, and the starting point of the audio characteristic curve coincides with the origin of the coordinate system;

[0131] Obtaining the digital encoded characters in the encoded character column of the preset encoding table; wherein, the preset encoding table includes an encoded character column and an original data column that are mapped one by one;

[0132] Performing a summation calculation on each digital encoded character, and substituting the calculation result as the abscissa value into the audio characteristic curve in the preset coordinate system to calculate the corresponding ordinate value;

[0133] Decompose the ordinate value to obtain a plurality of different decomposed numbers; insert each of the decomposed numbers into the head of the encoding character column of the preset encoding table in sequence, and shift the subsequent encoding characters in the encoding character column backward until the encoding character column is filled completely, thereby obtaining the changed encoding table.

[0134] In this embodiment, the preset coordinate system provides a unified mathematical space for analyzing the audio characteristic curve and subsequent operations. Coinciding the starting point of the audio characteristic curve with the origin of the coordinate system is to eliminate the influence of the difference in the starting position of the curve on subsequent calculations, so that all audio characteristic curves are processed under the same benchmark, facilitating quantitative analysis and operations. Specifically, determine the meanings, scales, and ranges of the coordinate axes of the preset coordinate system. According to the original coordinate information of the audio characteristic curve, through translation transformation, adjust the starting point of the curve to the origin (0, 0) of the coordinate system. The specific calculation method of translation is to obtain the original coordinates (x0, y0) of the curve starting point, and then perform coordinate transformation on each point (x, y) on the curve. The new coordinates are (x - x0, y - y0). Unifying the data benchmark simplifies the process of calculating the relationship between the audio characteristic curve and other data (such as the calculation results related to encoding characters) subsequently, and improves the accuracy and consistency of operations.

[0135] The preset encoding table is used to establish the mapping relationship between the original data and the encoding characters. Extracting the digital encoding characters from the encoding character column is for subsequent specific calculations and processing based on these numbers to realize the change of the encoding table and make it associated with the audio characteristic curve. Specifically, traverse the encoding character column of the preset encoding table and screen out the encoding characters belonging to the digital type. This may involve judging the type of each character and using character encoding rules (such as ASCII code) to identify digital characters (usually 0 - 9 correspond to specific encoding ranges). Extract the key data for subsequent calculations. These digital encoding characters will serve as the basis for calculations, and through specific operations, they will be associated with the audio characteristic curve, thereby adjusting the encoding table.

[0136] By summing up the numerically encoded characters, a numerical value representing the comprehensive characteristics of these encoded characters is obtained. Substituting this value as the abscissa into the audio characteristic curve is to obtain the corresponding ordinate value from the audio characteristic curve. This ordinate value will be used for subsequent adjustment of the coding table to establish the connection between the encoded characters and the audio characteristics. Accumulatively sum up all the numerically encoded characters extracted to obtain a total value. Take this total value as the abscissa value x and substitute it into the audio characteristic curve equation placed in the preset coordinate system (if the audio characteristic curve can be represented by a mathematical function) or by looking up the curve data points (if the curve is stored in the form of discrete data points) to obtain the corresponding ordinate value y. Associate the information of the encoded characters with the audio characteristic curve and use the characteristics of the audio characteristic curve to generate a new numerical value (ordinate value), which will be used as the basis for adjusting the coding table so that the change of the coding table can reflect the characteristics of the audio.

[0137] Decompose the ordinate value to obtain multiple decomposed numbers. Insert these numbers at the head of the encoded character column to change the structure of the coding table so that it matches the characteristics of the audio characteristic curve and generate a changed coding table that better conforms to the characteristics of the audio data. Use a specific decomposition algorithm to decompose the calculated ordinate value. For example, if the ordinate value is an integer, it can be decomposed digit by digit into the numbers on each digit. Insert the decomposed numbers into the head of the encoded character column of the preset coding table in sequence. Each time a number is inserted, shift all the encoded characters after this position in the encoded character column backward by one position to ensure the integrity and order of the encoded character column until all the decomposed numbers are inserted. By inserting the decomposed numbers, the structure of the coding table is changed, making the coding table not only unique and more secure, but also able to reflect the characteristics of the audio characteristic curve.

[0138] In one embodiment, changing the preset coding table based on the audio characteristic curve to obtain a changed coding table includes:

[0139] Perform multi-scale wavelet transform on the audio characteristic curve to obtain sub-band signals of different scales and frequencies;

[0140] Calculate the fractal dimension of the audio characteristic curve; where the fractal dimension describes the complexity and self-similarity of the curve;

[0141] Fuse the sub-band signals of different scales and frequencies with the fractal dimension to form a comprehensive feature vector;

[0142] Layer the coding rules of the preset coding table according to their functions and applicable scopes, analyze the correlation relationships between the coding rules of each layer, and construct a coding rule correlation graph;

[0143] Calculate the similarity between the comprehensive feature vector and the feature templates corresponding to the coding rules at each layer in the preset coding table;

[0144] According to the similarity calculation results, screen out the coding rules whose similarity reaches the threshold; for the screened coding rules, perform adaptive parameter adjustment according to the specific characteristics of the audio characteristic curve;

[0145] Fuse each of the screened coding rules to obtain a new coding rule; integrate the new coding rule into the preset coding table to generate a modified coding table;

[0146] In this embodiment, the audio characteristic curve contains rich audio information, but this information is manifested differently at different scales and frequencies. Multiscale wavelet transform can decompose the audio characteristic curve into subband signals of different scales and frequencies, so that the characteristics of the audio at different levels can be analyzed more meticulously, providing more comprehensive information for subsequent feature fusion and coding rule adjustment. Use an appropriate wavelet basis function (such as Daubechies wavelet, Haar wavelet, etc.) to perform multiscale wavelet transform on the audio characteristic curve. This transform will decompose the curve into a series of subband signals of different scales and frequencies, and each subband signal represents the characteristics of the audio within a specific scale and frequency range. For example, the subband signals of small scales usually correspond to the high-frequency detail information of the audio, while the subband signals of large scales correspond to the low-frequency overall trend of the audio. Through multiscale wavelet transform, the complex information of the audio characteristic curve is decomposed into multiple subband signals with clear physical meanings, facilitating subsequent in-depth analysis and processing of audio characteristics.

[0147] The fractal dimension is an important index used to describe the complexity and self-similarity of a curve. The complexity and self-similarity of the audio characteristic curve reflect the internal structure and characteristics of the audio. Calculating the fractal dimension can quantify the characteristics of the audio characteristic curve from another perspective, providing unique information for subsequent feature fusion. Methods such as the box-counting method and the correlation dimension method can be used to calculate the fractal dimension of the audio characteristic curve. Taking the box-counting method as an example, divide the plane where the curve is located into boxes of different sizes, count the number of boxes required to cover the curve, and then calculate the fractal dimension according to the relationship between the box size and the number of boxes. The fractal dimension provides a quantitative index of complexity and self-similarity for the audio characteristic curve. Combined with the subband signals obtained by multiscale wavelet transform, it can describe the characteristics of the audio more comprehensively.

[0148] Sub-band signals of different scales and frequencies, as well as the fractal dimension, describe the characteristics of the audio characteristic curve from different perspectives. By fusing them into a comprehensive feature vector, these scattered feature information can be integrated together to form a feature representation that can comprehensively represent the audio characteristics, facilitating subsequent matching and comparison with the coding rules. The eigenvalue of the sub-band signal of different scales and frequencies and the fractal dimension can be arranged in a certain order to form a vector. For example, the energy values of each sub-band signal can be arranged in sequence first, and then the fractal dimension is added to the end of the vector. The comprehensive feature vector integrates the feature information obtained from multi-scale wavelet transform and fractal dimension calculation, providing a unified and comprehensive feature representation for subsequent screening and adjustment of coding rules.

[0149] The coding rules in the preset coding table may have different functions and application scopes. Through layering, these rules can be organized more clearly, facilitating their management and analysis. Constructing a coding rule association graph can visually display the dependency relationships and mutual influences among the coding rules at each layer, providing a basis for subsequent screening and adjustment of coding rules. The coding rules in the preset coding table are layered according to their functions (such as audio noise reduction, audio compression, etc.) and application scopes (such as for specific types of audio, audio in a specific frequency range, etc.). For example, the general coding rules can be placed in the first layer, and the coding rules for specific audio characteristics can be placed in the second layer. Then, analyze the association relationships among the coding rules at each layer. For example, some rules may depend on the output results of other rules. Use the method of graph theory to construct a coding rule association graph, where the nodes represent the coding rules and the edges represent the association relationships among the rules. The layering and construction of the association graph make the coding rules in the preset coding table more structured and visual, facilitating understanding and operation, and providing a clear framework for subsequent screening and adjustment of coding rules according to audio characteristics.

[0150] Each layer of coding rules in the preset coding table usually has a corresponding feature template, which describes the range of audio characteristics applicable to this coding rule. By calculating the similarity between the comprehensive feature vector and these feature templates, the matching degree of each coding rule with the current audio characteristic curve can be judged, providing a basis for screening suitable coding rules. Methods such as Euclidean distance and cosine similarity can be used to calculate the similarity between the comprehensive feature vector and the feature template. Taking cosine similarity as an example, calculate the dot product of the comprehensive feature vector and the feature template vector, and then divide it by the product of their modulus lengths. The resulting value is the cosine similarity. The closer the value is to 1, the higher the similarity. By calculating the similarity, the matching degree of each coding rule with the current audio characteristic curve can be quantified, providing an objective basis for subsequent screening of the most suitable coding rule.

[0151] Screening out coding rules with a similarity reaching the threshold can ensure that the selected coding rules have a high degree of matching with the current audio characteristic curve, enabling better audio encoding. Adaptive parameter adjustment of the screened coding rules according to the specific characteristics of the audio characteristic curve can further optimize the coding rules to make them more in line with the actual situation of the current audio. Set a similarity threshold, compare the calculated similarity with this threshold, and screen out coding rules with a similarity reaching or exceeding the threshold. For the screened coding rules, adjust their parameters according to the specific characteristics of the audio characteristic curve (such as the sub-band signal characteristics obtained by multi-scale wavelet transform, fractal dimension, etc.). For example, if a certain coding rule is used for audio compression and the energy of the audio is high in a certain frequency band, the compression parameters for this frequency band in the coding rule can be adjusted. Screening and adaptive parameter adjustment can make the coding rules better adapt to the characteristics of the audio characteristic curve, improving the accuracy and efficiency of encoding.

[0152] The screened coding rules may each have advantages for different aspects of the audio. Fusing them can combine the advantages of these rules to form a more optimized new coding rule. Integrating the new coding rule into the preset coding table to generate a changed coding table can make the coding table better adapt to the characteristics of the current audio and improve the encoding and processing capabilities for audio data. According to the correlation relationship between the coding rules and the characteristics of the audio characteristic curve, fuse the screened coding rules. For example, the coding methods for different frequency ranges in different rules can be combined to form a new coding rule. Then, add the new coding rule to the preset coding table to replace or supplement the original coding rule to generate a changed coding table. Through fusion and integration, the generated changed coding table can more accurately reflect the characteristics of the audio characteristic curve, providing more effective rules for subsequent audio encoding and processing.

[0153] Refer to Figure 2 , in another embodiment of the present invention, a system for audio data compression and transmission in an AI recording scenario is further provided, including:

[0154] An acquisition module, configured to acquire pure audio data in the AI recording scenario of the recording device;

[0155] An analysis module, configured to perform audio content analysis on the pure audio data based on a deep learning model to obtain an audio content analysis result;

[0156] An encoding module, configured to encode the pure audio data according to the audio content analysis result, in combination with an adaptive bit rate adjustment strategy, using a hybrid encoding algorithm to obtain preliminarily encoded audio data;

[0157] A compression module, which is used to utilize the lame library to dynamically adjust compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminary encoded audio data to obtain compressed audio data; wherein, the compression parameters include the compression ratio;

[0158] A transmission module, which is used to transmit the compressed audio data to the target device.

[0159] In this embodiment, for the specific implementation of each module in the above system embodiment, please refer to that described in the above method embodiment, and details will not be repeated here.

[0160] Refer to Figure 3 , in an embodiment of the present invention, a computer device is further provided. This computer device can be a server, and its internal structure can be as Figure 3 shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of this computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the above method.

[0161] Those skilled in the art can understand that Figure 3 the structure shown in

[0162] is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0163] In summary, the present invention provides a method and system for compressing and transmitting audio data in an AI recording scenario in an embodiment of the present invention, including: in an AI recording scenario of a recording device, obtaining pure audio data; performing audio content analysis on the pure audio data based on a deep learning model to obtain an audio content analysis result; according to the audio content analysis result, combining an adaptive bitrate adjustment strategy, and encoding the pure audio data using a hybrid coding algorithm to obtain preliminary encoded audio data; using the lame library to dynamically adjust compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compressing the preliminary encoded audio data to obtain compressed audio data; wherein the compression parameters include a compression ratio; and transmitting the compressed audio data to a target device. In the present invention, according to the audio content analysis result, combining an adaptive bitrate adjustment strategy, and encoding the pure audio data using a hybrid coding algorithm improves the audio coding effect and coding efficiency; at the same time, dynamically adjusting compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device overcomes the defect of poor compression efficiency caused by the current inability to adaptively adjust compression parameters.

[0164] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0165] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, device, article or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, device, article or method including such an element.

[0166] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A method for compressing and transmitting audio data in an AI recording scenario, characterized in that, It includes the following steps: Under the AI recording scenario of the recording device, obtain pure audio data; Based on a deep learning model, perform audio content analysis on the pure audio data to obtain an audio content analysis result; According to the audio content analysis result, in combination with an adaptive bitrate adjustment strategy, use a hybrid coding algorithm to encode the pure audio data to obtain preliminary encoded audio data; Utilize the lame library to dynamically adjust compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminary encoded audio data to obtain compressed audio data; wherein, the compression parameters include the compression ratio; Transmit the compressed audio data to the target device; Utilize the lame library to dynamically adjust compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminary encoded audio data to obtain compressed audio data, including: Establish a multi-dimensional state matrix of hardware performance, storage space, and network bandwidth, and store the optimal compression parameters in different states in the corresponding matrix elements; Obtain the current hardware performance, storage space, and network bandwidth information of the recording device, and determine its position in the multi-dimensional state matrix; When it is detected that the hardware performance, storage space, or network bandwidth changes, perform dynamic search in the multi-dimensional state matrix according to the change trend and amplitude to obtain multiple candidate compression parameters; Assign different fuzzy membership degrees to each candidate compression parameter, calculate the comprehensive score of each candidate compression parameter according to the fuzzy inference rules, and select the one with the highest score as the final compression parameter; Utilize the lame library to compress the preliminary encoded audio data according to the final compression parameter to obtain compressed audio data.

2. The method for compressing and transmitting audio data in the AI recording scenario according to claim 1, wherein Obtain pure audio data, including: After the recording device sets the recording-related parameters, record, and use a noise reduction algorithm to perform noise reduction on the recorded audio data to obtain pure audio data.

3. The method for compressing and transmitting audio data in the AI recording scenario according to claim 1, wherein, According to the audio content analysis result, in combination with an adaptive bitrate adjustment strategy, use a hybrid coding algorithm to encode the pure audio data to obtain preliminary encoded audio data, including: Apply the information entropy calculation method to analyze the information richness of the audio in different frequency bands and time windows based on the audio content analysis result to obtain an audio complexity quantization value; Based on a pre-constructed complexity-bitrate mapping table, match the target bitrate mapped to the audio complexity quantization value as the adaptively adjusted bitrate; From the lossless coding, lossy coding, and hybrid coding modes, dynamically match the most suitable coding mode and parameter combination according to the target bitrate and audio complexity quantization value to obtain the optimal coding mode and parameter set; Divide the pure audio data into multiple data blocks, and perform adaptive encoding on each data block based on the optimal coding mode and parameter set to obtain preliminary encoded audio data.

4. The method for compressing and transmitting audio data in the AI recording scenario according to claim 1, characterized in that According to the audio content analysis result, in combination with an adaptive bitrate adjustment strategy, use a hybrid coding algorithm to encode the pure audio data to obtain preliminary encoded audio data, including: Based on the analysis results of the audio content, the pure audio data is divided into multiple time segments, and an audio feature vector is generated for each time segment; Using a reinforcement learning model, with each audio feature vector as the input, the adaptive bitrate adjustment strategy as the action space, and the comprehensive evaluation of the encoded audio quality and the encoded data volume as the reward function, dynamically select the optimal bitrate for each time segment; Adopt a hybrid coding algorithm, and dynamically combine different coding sub-algorithms according to the spectral characteristics of each time segment; Based on the selected optimal bitrate and the dynamically combined coding sub-algorithms, encode each time segment, and merge the encoding results of each time segment to obtain the preliminary encoded audio data.

5. The method for compressing and transmitting audio data in the AI recording scenario according to claim 1, wherein Using the lame library, dynamically adjust the compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device, and compress the preliminary encoded audio data to obtain compressed audio data, including: Obtain the hardware performance data, storage space usage, and network bandwidth change data of the recording device within a preset time period before the current time, and construct a time series data set; Input the time series data set into a long short-term memory network model for analysis to predict the hardware performance trend, remaining storage space, and network bandwidth fluctuation range of the recording device in the next period of time; According to the predicted results, plan the compression parameters of the lame library in advance; During the compression operation, obtain the real-time monitored hardware performance, storage space, and network bandwidth data, make real-time corrections to the pre-planned compression parameters, and use the lame library to compress the preliminary encoded audio data according to the corrected compression parameters to obtain compressed audio data.

6. The method for compressing and transmitting audio data in the AI recording scenario according to claim 1, wherein After transmitting the compressed audio data to the target device, including: The target device obtains the audio characteristic curve and audio attribute information of the compressed audio data; Obtain the preset coding table in the database, and modify the preset coding table based on the audio characteristic curve to obtain a modified coding table; Encode the audio attribute information based on the modified coding table to obtain an attribute code; Generate a management key based on the attribute code to perform permission management on the compressed audio data.

7. The method for compressed transmission of audio data in the AI recording scenario according to claim 6, wherein Modifying the preset coding table based on the audio characteristic curve to obtain a modified coding table includes: Move the audio characteristic curve to a preset coordinate system, and the starting point of the audio characteristic curve coincides with the origin of the coordinate system; Extract the coding characters in the coding character column of the preset coding table and combine them into a character sequence in sequence; wherein, the preset coding table includes a coding character column and an original data column in one-to-one correspondence; Tile the character sequence onto the preset coordinate system according to a preset rule, and detect the position relationship between each coding character and the audio characteristic curve; reorder each coding character according to the position relationship; Add each reordered coding character to the coding character column of the preset coding table in sequence to obtain a modified coding table.

8. A system for compressing and transmitting audio data in an AI recording scenario, characterized in that, Including: An acquisition module for acquiring pure audio data in the AI recording scenario of the recording device; An analysis module for performing audio content analysis on the pure audio data based on a deep learning model to obtain an audio content analysis result; An encoding module for encoding the pure audio data by using a hybrid encoding algorithm according to the audio content analysis result in combination with an adaptive bitrate adjustment strategy to obtain preliminarily encoded audio data; A compression module for dynamically adjusting compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device by using the lame library, and compressing the preliminarily encoded audio data to obtain compressed audio data; wherein the compression parameters include a compression ratio; A transmission module for transmitting the compressed audio data to a target device; Dynamically adjusting compression parameters according to the hardware performance, storage space, and network bandwidth of the recording device by using the lame library, and compressing the preliminarily encoded audio data to obtain compressed audio data, including: Establishing a multi-dimensional state matrix of hardware performance, storage space, and network bandwidth, and storing the optimal compression parameters in different states in the corresponding matrix elements; Obtaining the current hardware performance, storage space, and network bandwidth information of the recording device, and determining its position in the multi-dimensional state matrix; When it is detected that the hardware performance, storage space, or network bandwidth changes, dynamically searching in the multi-dimensional state matrix according to the trend and amplitude of the change to obtain multiple candidate compression parameters; Assigning different fuzzy membership degrees to each candidate compression parameter, calculating the comprehensive score of each candidate compression parameter according to the fuzzy inference rules, and selecting the one with the highest score as the final compression parameter; Using the lame library to compress the preliminarily encoded audio data according to the final compression parameter to obtain compressed audio data.

Citation Information

Patent Citations

  • Transmission method and device and storage medium

    CN110601872A

  • Bluetooth earphone low-delay transmission method

    CN117440440A

  • Wireless audio equipment low-power-consumption communication method and device, SoC chip and storage medium

    CN119649835A