A method and device for processing audio frame loss and a Bluetooth headset
By extracting and comprehensively judging audio data in multiple aspects, the problems of low accuracy of audio frame drop detection and poor frame filling effect in the prior art are solved, and more efficient frame drop processing and audio quality maintenance are achieved.
Patent Information
- Application Number
- CN202211496286.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-11-25
AI Technical Summary
When processing audio frame drops, the detection accuracy is low and the frame filling effect is poor, resulting in incoherent audio waveforms and distortion in auditory perception.
By extracting frame features of audio data, comprehensively considering energy characteristics, time domain characteristics, frequency domain characteristics, music theory characteristics and perception characteristics, we judge whether audio frame drops occur, and differentiated processing is carried out according to the degree of frame drops, and priority is given to reducing the quality of audio data or performing complementary frame processing.
It improves the accuracy of audio frame drop detection, improves the efficiency of frame filling processing, and reduces the breakage of audio waveforms and distortion in auditory perception.
Smart Images

Figure CN115881143B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of audio equipment, and in particular relates to a method and device for processing audio frame loss, and a Bluetooth headset. Background Art
[0002] With the rapid development of mobile Internet, users' demand for music is increasing. Wireless Bluetooth headsets are popular among users because they avoid the drag of headphone wires. However, the wireless Bluetooth communication environment is complex and changeable, which leads to problems such as bit errors and frame loss in the audio data obtained by wireless Bluetooth headsets, low voice transmission quality, and poor audio communication service quality. Frame loss will reduce the quality of audio decoding, and the audio waveform cannot remain coherent, which will further cause audio distortion in auditory perception.
[0003] When facing audio frame loss, the prior art often only judges whether the audio frame loss occurs based on one audio parameter. For example, it judges whether the audio frame loss occurs based on whether the frequency domain component is complete. However, the manifestations of audio frame loss are often multifaceted. For example, when some audio frames are actively discarded due to insufficient hardware conditions during the caching process, the frequency domain components are still complete. Obviously, it is not possible to accurately judge the occurrence of frame loss based on the frequency domain parameters alone. Moreover, when it is necessary to conceal the audio frame loss, whether it is caused by communication quality problems or hardware configuration problems, the frame filling process is often only performed mechanically without actually paying attention to the cause of the frame loss, resulting in inefficient concealment of the frame loss. In the process of frame filling, the frame filling is often only performed mechanically by copying and pasting the attachment frame, resulting in poor frame filling effect, the audio waveform cannot remain coherent, and even after the frame filling, the audio will still be distorted in auditory perception. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides a method and device for processing audio frame loss and a Bluetooth headset.
[0005] First aspect
[0006] The present invention provides a method for processing audio frame loss, which is applied to a Bluetooth headset, wherein a Bluetooth communication connection is established between the Bluetooth headset and an external device, and the method for processing audio frame loss includes:
[0007] S101: Acquire audio data from an external device;
[0008] S102: extracting features from the audio data frame by frame, and obtaining energy feature values, time domain feature values, frequency domain feature values, music theory feature values, and perception feature values corresponding to the frame data of each frame;
[0009] S103: Calculate the energy coherence parameter value between every two frames of frame data according to the energy characteristic value; if the energy coherence parameter value is less than a first preset value, set the energy coherence result to 1 to indicate that the energy characteristic is coherent; otherwise, set the energy coherence result to 0 to indicate that the energy characteristic is incoherent; in this way, calculate the time domain coherence result, the frequency domain coherence result, the music theory coherence result and the perceptual coherence result;
[0010] S104: Calculate a coherence result between two frames of data, where the coherence result is the sum of an energy coherence result, a time domain coherence result, a frequency domain coherence result, a music theory coherence result, and a perception coherence result. If the coherence result is less than a second preset value, determine that there is frame loss between the two frames of data, and set the frame loss result between the two frames of data to 1, otherwise, set it to 0;
[0011] S105: taking a preset number of frame data as a group, calculating a group frame loss result, wherein the group frame loss result is the sum of the frame loss results in the group;
[0012] S106: when the group frame loss result is greater than a third preset value and less than a fourth preset value, reducing the quality of the audio data in a preset order, wherein the preset order is premium quality, lossless quality, high quality, and standard quality;
[0013] S107: when the group frame loss result is greater than or equal to a fourth preset value, determining the number of frames N that need to be supplemented between the first frame data and the second frame data according to the first frame data and the second frame data between which frame loss exists;
[0014] S108: searching whether there is similar frame data that meets the conditions before the first frame data, wherein the similarity between the energy characteristic value, the time domain characteristic value, the frequency domain characteristic value, the music theory characteristic value and the perception characteristic value of the similar frame data and the first frame data is within a preset range;
[0015] S109: if there is similar frame data that meets the conditions, copy N frames of data following the similar frame data and insert them after the first frame of data;
[0016] S110: when there is no similar frame data meeting the condition, extracting energy feature values of each frame data to form an energy feature curve, extracting time domain feature values of each frame data to form a time domain feature curve, and extracting frequency domain feature values of each frame data to form a frequency domain feature curve;
[0017] S111: In the energy characteristic curve, the time domain characteristic curve, and the frequency domain characteristic curve, N horizontal coordinates are inserted between the horizontal coordinate corresponding to the first frame data and the horizontal coordinate corresponding to the second frame data, and the curve is fitted by using the least square method according to the known values in the curve to obtain the target energy characteristic value, the target time domain characteristic value, and the target frequency domain characteristic value corresponding to the inserted N horizontal coordinates;
[0018] S112: copying the first frame data, and transforming the first frame data so that the transformed frame data reaches a target energy characteristic value, a target time domain characteristic value, and a target frequency domain characteristic value;
[0019] S113: inserting the transformed frame data between the first frame data and the second frame data.
[0020] Second aspect
[0021] The present invention provides an audio frame loss processing device, which is applied to a Bluetooth headset, wherein a Bluetooth communication connection is established between the Bluetooth headset and an external device, and the audio frame loss processing device comprises:
[0022] An acquisition module, used for acquiring audio data from an external device;
[0023] A first extraction module is used to extract features from the audio data frame by frame, and obtain energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perception eigenvalues corresponding to the frame data of each frame;
[0024] A first calculation module is used to calculate the energy coherence parameter value between every two frames of frame data according to the energy characteristic value, and when the energy coherence parameter value is less than a first preset value, the energy coherence result is set to 1 to indicate that the energy characteristic is coherent, otherwise the energy coherence result is set to 0 to indicate that the energy characteristic is incoherent; in this way, the time domain coherence result, the frequency domain coherence result, the music theory coherence result and the perceptual coherence result are calculated;
[0025] A second calculation module is used to calculate a coherence result between two frames of data, where the coherence result is the sum of an energy coherence result, a time domain coherence result, a frequency domain coherence result, a music theory coherence result, and a perception coherence result. When the coherence result is less than a second preset value, it is determined that there is frame loss between the two frames of data, and the frame loss result between the two frames of data is set to 1, otherwise it is set to 0;
[0026] A third calculation module is used to calculate a group frame loss result by taking a preset number of frame data as a group, wherein the group frame loss result is the sum of the frame loss results in the group;
[0027] a reducing module, configured to reduce the quality of the audio data in a preset order when the group frame loss result is greater than a third preset value and less than a fourth preset value, wherein the preset order is premium quality, lossless quality, high quality, and standard quality;
[0028] A determination module, configured to determine the number of frames N that need to be supplemented between the first frame data and the second frame data according to the first frame data and the second frame data between which frame loss occurs when the group frame loss result is greater than or equal to a fourth preset value;
[0029] A retrieval module is used to retrieve whether there is similar frame data that meets the conditions before the first frame data, wherein the similarity between the energy characteristic value, the time domain characteristic value, the frequency domain characteristic value, the music theory characteristic value and the perception characteristic value of the similar frame data and the first frame data is within a preset range;
[0030] A first copying module is used to copy N frames of data following the similar frame data when there is similar frame data meeting the conditions, and insert the N frames of data after the similar frame data, and insert the N frames of data after the first frame data;
[0031] The second extraction module is used to extract the energy characteristic value of each frame data to form an energy characteristic curve, extract the time domain characteristic value of each frame data to form a time domain characteristic curve, and extract the frequency domain characteristic value of each frame data to form a frequency domain characteristic curve when there is no similar frame data that meets the conditions;
[0032] A first insertion module is used to insert N horizontal coordinates between the horizontal coordinate corresponding to the first frame data and the horizontal coordinate corresponding to the second frame data in the energy characteristic curve, the time domain characteristic curve and the frequency domain characteristic curve, and to perform curve fitting using the least square method according to the known values in the curve to obtain the target energy characteristic value, the target time domain characteristic value and the target frequency domain characteristic value corresponding to the inserted N horizontal coordinates;
[0033] A second copying module is used to copy the first frame data and transform the first frame data so that the transformed frame data reaches a target energy characteristic value, a target time domain characteristic value and a target frequency domain characteristic value;
[0034] The second inserting module is used to insert the transformed frame data between the first frame data and the second frame data.
[0035] The third aspect
[0036] The present invention provides a Bluetooth headset, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the audio frame loss processing method as in the first aspect when executing the computer program.
[0037] Compared with the prior art, the present invention has at least the following beneficial effects:
[0038] 1. In the present invention, the external manifestations of the device audio frame loss are fully considered, the energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perception eigenvalues of the audio are integrated together, and the energy coherence results, time domain coherence results, frequency domain coherence results, music theory coherence results and perception coherence results are comprehensively judged whether the audio data has frame loss, which greatly improves the accuracy of frame loss detection.
[0039] 2. In the present invention, differentiated processing is performed according to the degree of frame loss. If the degree of frame loss is low, the audio data transmission quality can be reduced first, and priority can be given to checking whether the frame loss is caused by communication quality problems. If the degree of frame loss is high, frame supplementation processing is performed. Compared with frame supplementation, reducing the audio data transmission quality is more efficient in concealing frame loss, avoiding mechanical frame supplementation processing, which results in low processing efficiency.
[0040] 3. In the present invention, if the audio is of a musical nature, there will often be some repeated tunes. In the process of frame supplementation, it is possible to preferentially search whether there is audio data with very similar energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perceptual eigenvalues in the audio before the frame loss position, and preferentially use similar data for frame supplementation. If no similar audio exists, the audio contour is extracted for frame supplementation, and the audio waveform is kept consistent as much as possible to repair the distortion of the audio in auditory perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The preferred implementation modes will be described below in a clear and understandable manner with reference to the accompanying drawings to further illustrate the above-mentioned characteristics, technical features, advantages and implementation methods of the present invention.
[0042] Figure 1 It is a flowchart of a method for processing audio frame loss provided by the present invention;
[0043] Figure 2 It is a flowchart of a method for detecting audio frame loss provided by the present invention;
[0044] Figure 3 It is a flowchart of an audio feature extraction method provided by the present invention;
[0045] Figure 4 It is a flow chart of a method for calculating energy coherence results provided by the present invention;
[0046] Figure 5 It is a flow chart of a method for calculating the number of supplemented frames provided by the present invention;
[0047] Figure 6 It is a structural schematic diagram of an audio frame loss processing device provided by the present invention;
[0048] Figure 7 The present invention provides a schematic diagram of the hardware structure of a Bluetooth headset. DETAILED DESCRIPTION
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the specific implementation methods of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings and other implementation methods can be obtained based on these drawings without creative work.
[0050] In order to simplify the drawings, only the parts related to the invention are schematically shown in each figure, and they do not represent the actual structure of the product. In addition, in order to simplify the drawings and facilitate understanding, in some figures, only one of the parts with the same structure or function is schematically drawn or marked. In this article, "one" not only means "only one", but also means "more than one".
[0051] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0052] In this document, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0053] In addition, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0054] Example 1
[0055] In one embodiment, the reference specification Figure 1 , a flowchart of a method for processing audio frame loss provided by the present invention. Figure 2 , is a flow chart of an audio frame loss detection method provided by the present invention.
[0056] The present invention provides a method for processing audio frame loss, which is applied to a Bluetooth headset, wherein a Bluetooth communication connection is established between the Bluetooth headset and an external device.
[0057] Among them, the external devices can be mobile phones, laptops, wearable devices, etc.
[0058] Methods for handling audio frame loss include:
[0059] S101: Acquire audio data from an external device.
[0060] S102: Extract features of the audio data frame by frame to obtain energy feature values, time domain feature values, frequency domain feature values, music theory feature values, and perception feature values corresponding to the frame data of each frame.
[0061] It should be noted that each audio has energy characteristics, time domain characteristics, frequency domain characteristics, music theory characteristics and perception characteristics.
[0062] In one possible implementation, the energy characteristic value is the amplitude;
[0063] The time domain feature value is the autocorrelation value, where the autocorrelation value refers to the similarity between the signal and its version shifted along the time axis, which can be used to calculate the fundamental frequency of a single tone;
[0064] The frequency domain eigenvalue is the spectrum centroid value, where the spectrum centroid value refers to the concentration point of the signal energy in the spectrum, which can be used to describe the brightness of the signal timbre. It can be understood that the brighter the sound energy is concentrated in the high-frequency part, the larger the value of the spectrum centroid;
[0065] The music theory characteristic value is the detuning value, where the detuning value refers to the degree of deviation between the overtone frequency of an audio signal and an integer multiple of its fundamental frequency; the fundamental frequency, referred to as fundamental frequency, the sound can be decomposed into the superposition of several sine waves of different frequencies. The wave with the lowest frequency is the fundamental frequency, and the other high frequencies are overtones. The higher the frequency, the less energy is allocated.
[0066] The perceptual characteristic value is the loudness value, where the loudness value refers to the subjective feeling of the signal strength felt by the human ear, which can also be understood as the volume.
[0067] It should be noted that once frame loss occurs in audio data, audio parameters such as the audio amplitude, autocorrelation value, spectral centroid value, detuning value and loudness value at the frame loss point will often change suddenly. By comprehensively judging whether frame loss occurs in audio data through these different features, the accuracy of audio frame loss detection can be greatly improved.
[0068] Furthermore, in addition to amplitude, autocorrelation value, spectral centroid value, detuning value and loudness value, there are many specific audio parameters of energy characteristics, time domain characteristics, frequency domain characteristics, music theory characteristics and perceptual characteristics. Technical personnel in this field can select corresponding audio parameters according to actual conditions.
[0069] In a possible implementation manner, refer to the attached specification Figure 3 , is a flow chart of an audio feature extraction method provided by the present invention. S102 specifically includes:
[0070] S1021: adding windows and framing to the audio data, and extracting energy eigenvalues and time domain eigenvalues by frame;
[0071] S1022: Perform short-time Fourier transform on the audio data to obtain a short-time spectrum, and extract frequency domain eigenvalues of the short-time spectrum frame by frame;
[0072] S1023: extracting the fundamental frequency from the short-time spectrum, and calculating the music theory characteristic value according to the overtone frequency and the fundamental frequency of the frame data;
[0073] S1024: Extract perceptual feature values for the audio data frame by frame through the auditory perception model.
[0074] In actual application, the energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perceptual eigenvalues of audio data can be extracted in stages and orderly in the above manner, which can improve processing efficiency while taking into account the accuracy of frame loss detection.
[0075] S103: Calculate the energy coherence parameter value between every two frames of frame data according to the energy characteristic value. When the energy coherence parameter value is less than a first preset value, set the energy coherence result to 1 to represent coherence in terms of energy characteristics. Otherwise, set the energy coherence result to 0 to represent incoherence in terms of energy characteristics.
[0076] In a possible implementation manner, refer to the attached specification Figure 4 , is a flow chart of a method for calculating energy coherence results provided by the present invention. S103 specifically includes:
[0077] S1031: Calculate the average value of the energy characteristic values of the frame data from the first frame to the Mth frame;
[0078] S1032: using the energy characteristic value of the frame data of the M+1th frame minus the average value of the energy characteristic value of the previous M frames of data as the energy coherence parameter value between the Mth frame data and the M+1th frame data;
[0079] It should be noted that the energy characteristic value of the frame data of the M+1th frame minus the average value of the energy characteristic values of the previous M frames is used as the energy coherence parameter value, which can be understood as a jump mutation value of the energy characteristic value of the frame data of the M+1th frame relative to the energy characteristic values of all previous data.
[0080] S1033: When the energy coherence parameter value is less than the first preset value, the energy coherence result is set to 1; otherwise, the energy coherence result is set to 0.
[0081] It should be noted that if the energy coherence parameter value is less than the first preset value, it means that the energy characteristic value of the frame data of the M+1th frame has not undergone a sudden jump, and the corresponding possibility of frame loss is small, and the energy characteristic value can be considered to be coherent, which is recorded as 1. On the contrary, if the energy coherence parameter value is greater than the first preset value, it means that the energy characteristic value of the frame data of the M+1th frame has undergone a sudden jump, and the corresponding possibility of frame loss is large, and the energy characteristic value can be considered to be incoherent, which is recorded as 0.
[0082] Among them, those skilled in the art can adjust the size of the first preset value according to actual conditions, and the present invention does not limit the specific numerical value of the first preset value.
[0083] According to this method, the time domain coherence results, frequency domain coherence results, music theory coherence results and perceptual coherence results are calculated.
[0084] Among them, each of the energy coherence results, time domain coherence results, frequency domain coherence results, music theory coherence results and perception coherence results is valuable for judging whether the audio has frame loss. The more 1s there are in the energy coherence results, time domain coherence results, frequency domain coherence results, music theory coherence results and perception coherence results of the audio data, the smaller the probability of frame loss of the audio data; conversely, the more 0s there are in the energy coherence results, time domain coherence results, frequency domain coherence results, music theory coherence results and perception coherence results of the audio data, the greater the probability of frame loss of the audio data.
[0085] S104: Calculate the coherence result between two frames of data, where the coherence result is the sum of the energy coherence result, the time domain coherence result, the frequency domain coherence result, the music theory coherence result and the perceptual coherence result. When the coherence result is less than a second preset value, determine that there is frame loss between the two frames of data, and set the frame loss result between the two frames of data to 1, otherwise set it to 0.
[0086] For example, the energy coherence result between A and A+1 frame data is 0, the time domain coherence result is 0, the frequency domain coherence result is 1, the music theory coherence result is 0 and the perceptual coherence result is 0; at this time, the coherence result between A and A+1 frame data is 0+0+1+0+0=1.
[0087] The energy coherence result between B and B+I frame data is 1, the time domain coherence result is 0, the frequency domain coherence result is 1, the music theory coherence result is 1 and the perceptual coherence result is 0; at this time, the coherence result of the B frame data is 1+0+1+1+0=3.
[0088] The energy coherence result between C and C+1 frame data is 1, the time domain coherence result is 1, the frequency domain coherence result is 1, the music theory coherence result is 1 and the perceptual coherence result is 1; at this time, the coherence result of A frame data is 1+1+1+1+1=5.
[0089] If the second preset value is set to 3, then it can be determined that there is frame loss between A and A+1 frame data, there is no frame loss between B and B+1 frame data, and there is no frame loss between C and C+1 frame data.
[0090] If the size of the second preset value is set to 4, at this time, it can be determined that there is frame loss between A and A+1 frame data, there is frame loss between B and B+1 frame data, and there is no frame loss between C and C+1 frame data.
[0091] Accordingly, the scale of frame loss judgment can be adjusted by setting a specific size of the second preset value to balance accuracy and efficiency.
[0092] In a possible implementation, in the process of calculating the coherence result between two frames of data, the energy coherence result, the time domain coherence result, the frequency domain coherence result, the music theory coherence result and the perception coherence result may be summed by weighted summation.
[0093] Specifically, let the value of the energy coherence result be y 1 , the value of the time domain coherence result is y 2 , the value of the frequency domain coherence result is y 3 , the value of the music theory coherence result is y 4 , the value of the perceptual coherence result is y 5 , the weight of the energy coherence result is ζ, the weight of the time domain coherence result is β, the weight of the frequency domain coherence result is γ, the weight of the music theory coherence result is δ, and the weight of the perceptual coherence result is ε;
[0094] Then the coherence result z between two frames of data can be calculated by the following formula:
[0095] z=ζ·y 1 +β·y 2 +γ·y 3 +δ·y 4 +ε·y 5
[0096] Furthermore, in order to balance accuracy and efficiency, not only the specific size of the second preset value can be set, but also the weight of each coherent result can be adjusted to adjust the scale of frame loss judgment.
[0097] S105: Taking a preset number of frame data as a group, calculating a group frame loss result, wherein the group frame loss result is the sum of the frame loss results in the group.
[0098] It should be noted that by observing the frame data in groups, it is possible to evaluate whether the frame loss of the audio data is serious from a more macro and overall perspective.
[0099] Among them, those skilled in the art can adjust the size of the preset number according to actual conditions, and the present invention does not limit the specific value of the preset number.
[0100] It can be understood that the larger the group frame loss result is, the more serious the frame loss of audio data in the group is; conversely, the smaller the group frame loss result is, the less the frame loss of audio data in the group is.
[0101] S106: When the group frame loss result is greater than the third preset value and less than the fourth preset value, reduce the quality of the audio data in a preset order, wherein the preset order is premium quality, lossless quality, high quality and standard quality.
[0102] It can be understood that if the group frame loss result is less than the third preset value, it means that the frame loss result of the audio data is not serious and no processing is required.
[0103] If the group frame loss result is greater than the third preset value and less than the fourth preset value, it can be understood that there is frame loss, but the degree of frame loss is low. You can first reduce the audio data transmission quality to prioritize whether the frame loss is caused by communication quality problems. For communication quality problems, you can reduce the quality of audio data for processing. Compared with frame supplementation, reducing the quality of audio data transmission is more efficient in concealing frame loss, avoiding mechanical frame supplementation processing, resulting in low processing efficiency.
[0104] In a possible implementation, when reducing the quality of audio data in a preset order, only one level is allowed to be reduced, for example, from premium quality to lossless quality, from lossless quality to high quality, but not from premium quality to high quality. If the frame loss problem is still not solved after reducing the quality by one level, it is not appropriate to continue to sacrifice the user experience and continue to reduce the quality, and frame filling should be performed to cover up the frame loss.
[0105] S107: When the group frame loss result is greater than or equal to a fourth preset value, determine the number of frames N that need to be supplemented between the first frame data and the second frame data according to the first frame data and the second frame data between which frame loss occurs.
[0106] If the group frame loss result is greater than or equal to the fourth preset value, it can be understood that frame loss exists and the degree of frame loss is very high, which has exceeded the frame loss result that may be caused by poor communication quality. At this time, frame filling is directly performed to conceal the frame loss.
[0107] Among them, the present invention does not limit the specific values of the third preset value and the fourth preset value. If the fourth preset value is set too high, the consequence of misjudging the communication quality problem will be borne at this time, because time is wasted to reduce the quality, but the frame loss problem still cannot be solved. If the fourth preset value is set too low, frame loss will be concealed more by filling frames, and filling frames obviously takes more time than reducing audio quality. Therefore, those skilled in the art can adjust the size of the third preset value and the fourth preset value according to actual conditions to take into account both communication quality issues and hardware equipment issues.
[0108] It should be noted that, in the process of frame supplementation, how many frames to supplement for the most favorable frame loss concealment has always been a thorny issue. In the present invention, a corresponding number of frames are adaptively inserted according to the data characteristics of the first frame data and the second frame data.
[0109] In a possible implementation manner, refer to the attached specification Figure 5 , is a flow chart of a method for calculating the number of supplementary frames provided by the present invention. S107 specifically includes:
[0110] S1071: extracting perceptual feature values of frame data of each frame to form a perceptual feature matrix S1;
[0111] It should be noted that the perceptual eigenvalue can better reflect the human ear's perception of audio data than the energy eigenvalue, time domain eigenvalue, frequency domain eigenvalue, and music theory eigenvalue. Therefore, the perceptual eigenvalue is selected as the basis for calculating the number of interpolated frames.
[0112] S1072: In the perceptual feature matrix S1, the difference between the perceptual feature values of every two frames of data is calculated to form a difference matrix S2;
[0113] It should be noted that this step calculates the difference between the perceptual feature values of every two frames of frame data, which is used to represent the degree of jump mutation of the perceptual feature.
[0114] S1073: remove the elements whose values are greater than the fifth preset value in the difference matrix S2 to form a matrix S3;
[0115] It should be noted that removing elements greater than the fifth preset value is used to remove some obviously unreasonable data, and the corresponding audio data may have problems such as popping sounds. If frame infilling is performed based on the data with popping sound problems, the concealing effect of the frame infilling will be weakened.
[0116] Among them, those skilled in the art can adjust the size of the fifth preset value according to actual conditions, and the present invention does not limit the specific numerical value of the first preset value.
[0117] S1074: Calculate the average value of the elements in the matrix S3 as the frame filling parameter value α;
[0118] It should be noted that the average value at this time can be considered as the average value of the perceptual characteristics of normal data in the audio data, which is used to express the tone of the entire audio data.
[0119] S1075: Let the perceptual feature value of the first frame data be x 1 , the perceptual feature value of the second frame data is x 2 , then the number of frames N that need to be added between the first frame data and the second frame data can be calculated by the following formula:
[0120]
[0121]
[0122] in, Indicates rounding up.
[0123] It should be noted that the number of supplementary frames calculated by perceptual feature values is more consistent with the user's auditory perception and can better achieve the effect of concealing frame loss.
[0124] S108: Check whether there is similar frame data meeting the conditions before the first frame data.
[0125] Among them, the similarity degree of energy characteristic value, time domain characteristic value, frequency domain characteristic value, music theory characteristic value and perception characteristic value between the similar frame data and the first frame data is within a preset range.
[0126] Furthermore, the preset ranges required for the similarity of energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perceptual eigenvalues may be different. For example, the range of eigenvalues that can better reflect the human ear's perception of audio data may be set to a stricter point, while the range of time domain eigenvalues that are less perceived by the human ear may be set to a looser point, so as to better retrieve similar audio data.
[0127] S109: When similar frame data that meets the conditions exists, copy N frames of data following the similar frame data and insert them after the first frame of data.
[0128] It should be noted that if the audio is of a musical nature, there will often be some repeated tunes. In the process of frame supplementation, it is possible to first search for audio data with very similar energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues, and perceptual eigenvalues in the audio before the frame loss position, and give priority to using similar data for frame supplementation. At this point, on the one hand, compared with extracting the audio contour for frame supplementation, the amount of calculation is reduced, which can improve the efficiency of frame supplementation. On the other hand, frame supplementation based on similar audio data is more conducive to maintaining the waveform coherence of the audio and repairing the distortion of the audio in auditory perception.
[0129] S110: When there is no similar frame data that meets the conditions, extract the energy characteristic value of each frame data to form an energy characteristic curve, extract the time domain characteristic value of each frame data to form a time domain characteristic curve, and extract the frequency domain characteristic value of each frame data to form a frequency domain characteristic curve.
[0130] It should be noted that since energy features, time domain features and frequency domain features can be obtained directly or simply through processing of audio data, and the music theory feature values need to extract the fundamental frequency to calculate, and the perceptual features need to be obtained with the help of the auditory perception model, it will increase the complexity for subsequent calculations. Therefore, energy features, time domain features and frequency domain features are selected for frame complementation.
[0131] Optionally, energy features, time domain features, frequency domain features, music theory features and perception features may all be used for frame supplementation. This may achieve the best frame supplementation concealment effect, but may increase the calculation and processing time of frame supplementation.
[0132] S111: In the energy characteristic curve, the time domain characteristic curve and the frequency domain characteristic curve, N horizontal coordinates are inserted between the horizontal coordinates corresponding to the first frame data and the horizontal coordinates corresponding to the second frame data, and the curves are fitted using the least squares method according to the known values in the curves to obtain the target energy characteristic values, target time domain characteristic values and target frequency domain characteristic values corresponding to the inserted N horizontal coordinates.
[0133] It should be noted that, by selecting points corresponding to the horizontal coordinates in the fitted characteristic curve, and then counting their vertical coordinates, the target energy eigenvalue, the target time domain eigenvalue, and the target frequency domain eigenvalue can be obtained.
[0134] S112: copying the first frame data, and transforming the first frame data so that the transformed frame data reaches a target energy eigenvalue, a target time domain eigenvalue, and a target frequency domain eigenvalue.
[0135] It should be noted that the first frame data is adjacent data and has a high degree of similarity to the lost frame data. Therefore, the first frame data is selected as a basis and transformed on this basis to obtain the supplementary frame data, which can improve the efficiency of the supplementary frame.
[0136] S113: inserting the transformed frame data between the first frame data and the second frame data.
[0137] In a possible implementation manner, between S101 and S102, the following is further included:
[0138] S114: Identify silent segments in the audio data, and delete the silent segments.
[0139] It should be noted that deleting silent segments can be understood as a preprocessing of audio data before frame loss detection to prevent silent segments from interfering with the accuracy of subsequent frame loss detection. It can also reduce the amount of data processing and improve processing efficiency.
[0140] Specifically, the voice activity detection algorithm detects voice segments and silent segments in the audio data and removes the silent segments.
[0141] In a possible implementation manner, after S113, the method further includes:
[0142] S115: Perform cross-fading processing on the first frame data, the inserted N frames of data, and the second frame data.
[0143] The audio after inserting N frames of data is mixed and reconstructed. Cross-fading can make the reconstructed audio smoother. At the same time, cross-fading can complete the audio repair at a very low computational cost.
[0144] Compared with the prior art, the present invention has at least the following beneficial effects:
[0145] 1. In the present invention, the external manifestations of the device audio frame loss are fully considered, the energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perception eigenvalues of the audio are integrated together, and the energy coherence results, time domain coherence results, frequency domain coherence results, music theory coherence results and perception coherence results are comprehensively judged whether the audio data has frame loss, which greatly improves the accuracy of frame loss detection.
[0146] 2. In the present invention, differentiated processing is performed according to the degree of frame loss. If the degree of frame loss is low, the audio data transmission quality can be reduced first, and priority can be given to checking whether the frame loss is caused by communication quality problems. If the degree of frame loss is high, frame supplementation processing is performed. Compared with frame supplementation, reducing the audio data transmission quality is more efficient in concealing frame loss, avoiding mechanical frame supplementation processing, which results in low processing efficiency.
[0147] 3. In the present invention, if the audio is of a musical nature, there will often be some repeated tunes. In the process of frame supplementation, it is possible to preferentially search whether there is audio data with very similar energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perceptual eigenvalues in the audio before the frame loss position, and preferentially use similar data for frame supplementation. If no similar audio exists, the audio contour is extracted for frame supplementation, and the audio waveform is kept consistent as much as possible to repair the distortion of the audio in auditory perception.
[0148] Example 2
[0149] In one embodiment, the reference specification Figure 6 , a structural schematic diagram of an audio frame loss processing device provided by the present invention.
[0150] The present invention provides an audio frame loss processing device 20, which is applied to a Bluetooth headset. A Bluetooth communication connection is established between the Bluetooth headset and an external device. The audio frame loss processing device 20 includes:
[0151] An acquisition module 201 is used to acquire audio data from an external device;
[0152] The first extraction module 202 is used to extract features from the audio data frame by frame, and obtain energy feature values, time domain feature values, frequency domain feature values, music theory feature values and perception feature values corresponding to the frame data of each frame;
[0153] The first calculation module 203 is used to calculate the energy coherence parameter value between every two frames of frame data according to the energy characteristic value. When the energy coherence parameter value is less than a first preset value, the energy coherence result is set to 1 to indicate that the energy characteristic is coherent; otherwise, the energy coherence result is set to 0 to indicate that the energy characteristic is incoherent; in this way, the time domain coherence result, the frequency domain coherence result, the music theory coherence result and the perceptual coherence result are calculated;
[0154] A second calculation module 204 is used to calculate a coherence result between two frames of data, where the coherence result is the sum of an energy coherence result, a time domain coherence result, a frequency domain coherence result, a music theory coherence result, and a perception coherence result. If the coherence result is less than a second preset value, it is determined that there is frame loss between the two frames of data, and the frame loss result between the two frames of data is set to 1, otherwise it is set to 0;
[0155] The third calculation module 205 is used to calculate a group frame loss result by taking a preset number of frame data as a group, wherein the group frame loss result is the sum of the frame loss results in the group;
[0156] A reducing module 206, configured to reduce the quality of the audio data in a preset order when the group frame loss result is greater than a third preset value and less than a fourth preset value, wherein the preset order is premium quality, lossless quality, high quality, and standard quality;
[0157] A determination module 207 is used to determine the number of frames N that need to be supplemented between the first frame data and the second frame data according to the first frame data and the second frame data between which frame loss occurs when the group frame loss result is greater than or equal to a fourth preset value;
[0158] A search module 208 is used to search whether there is similar frame data that meets the conditions before the first frame data, wherein the similarity between the energy characteristic value, the time domain characteristic value, the frequency domain characteristic value, the music theory characteristic value and the perception characteristic value of the similar frame data and the first frame data is within a preset range;
[0159] The first copying module 209 is used to copy N frames of data following the similar frame data when there is similar frame data meeting the conditions, and insert the N frames of data after the similar frame data, and insert the N frames of data after the first frame data;
[0160] The second extraction module 210 is used to extract the energy characteristic value of each frame data to form an energy characteristic curve, extract the time domain characteristic value of each frame data to form a time domain characteristic curve, and extract the frequency domain characteristic value of each frame data to form a frequency domain characteristic curve when there is no similar frame data that meets the conditions;
[0161] A first insertion module 211 is used to insert N horizontal coordinates between the horizontal coordinate corresponding to the first frame data and the horizontal coordinate corresponding to the second frame data in the energy characteristic curve, the time domain characteristic curve and the frequency domain characteristic curve, and to perform curve fitting using the least square method according to known values in the curve to obtain target energy characteristic values, target time domain characteristic values and target frequency domain characteristic values corresponding to the inserted N horizontal coordinates;
[0162] The second copying module 212 is used to copy the first frame data and transform the first frame data so that the transformed frame data reaches the target energy characteristic value, the target time domain characteristic value and the target frequency domain characteristic value;
[0163] The second inserting module 213 is used to insert the transformed frame data between the first frame data and the second frame data.
[0164] In a possible implementation, the first extraction module 202 specifically includes:
[0165] The energy time domain feature extraction submodule is used to perform windowing and framing on the audio data, and extract energy feature values and time domain feature values by frame;
[0166] The frequency domain feature extraction submodule is used to perform short-time Fourier transform on the audio data to obtain a short-time spectrum, and extract frequency domain feature values on a frame-by-frame basis from the short-time spectrum;
[0167] The music theory feature extraction submodule is used to extract the fundamental frequency from the short-time spectrum and calculate the music theory feature value according to the overtone frequency and the fundamental frequency of the frame data;
[0168] The perceptual feature extraction submodule is used to extract perceptual feature values for audio data frame by frame through an auditory perception model.
[0169] In a possible implementation, the energy eigenvalue is the amplitude, the time domain eigenvalue is the autocorrelation value, the frequency domain eigenvalue is the spectrum centroid value, the music theory eigenvalue is the detuning value, and the perception eigenvalue is the loudness value.
[0170] In a possible implementation, the first calculation module 203 specifically includes:
[0171] A first average value calculation submodule, used to calculate the average value of the energy characteristic values of the frame data from the first frame to the Mth frame;
[0172] A subtraction submodule, used to use the energy characteristic value of the frame data of the M+1th frame to subtract the average value of the energy characteristic value of the previous M frames of data, as a value for calculating the energy coherence parameter between the Mth frame data and the M+1th frame data;
[0173] The energy consistency result submodule is used to set the energy consistency result to 1 when the energy consistency parameter value is less than a first preset value, and to set the energy consistency result to 0 otherwise.
[0174] In a possible implementation, the determination module 207 specifically includes:
[0175] An extraction submodule, used for extracting perceptual feature values of frame data of each frame to form a perceptual feature matrix S1;
[0176] A calculation submodule, used for calculating the difference between the perceptual feature values of every two frames of data in the perceptual feature matrix S1 to form a difference matrix S2;
[0177] A removal submodule, used for removing elements whose values are greater than a fifth preset value in the difference matrix S2, to form a matrix S3;
[0178] A second average value calculation submodule is used to calculate the average value of the elements in the matrix S3 as the frame filling parameter value α;
[0179] The frame number determination submodule is used to set the perceptual feature value of the first frame data to be x1 and the perceptual feature value of the second frame data to be x 2 , then the number of frames N that need to be added between the first frame data and the second frame data can be calculated by the following formula:
[0180]
[0181]
[0182] in, Indicates rounding up.
[0183] In a possible implementation manner, the audio frame loss processing device 20 further includes:
[0184] The deleting module 214 is used to identify the silent segments in the audio data and perform the following operations on the silent segments:
[0185] The gradient processing module 215 is used to perform cross gradient processing on the first frame data, the inserted N frames of data and the second frame of data.
[0186] The audio frame loss processing device 20 provided by the present invention can implement each process implemented in the above method embodiment, and will not be described again here to avoid repetition.
[0187] The virtual device provided by the present invention may be a device, or a component, an integrated circuit, or a chip in a terminal.
[0188] Compared with the prior art, the present invention has at least the following beneficial effects:
[0189] 1. In the present invention, the external manifestations of the device audio frame loss are fully considered, the energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perception eigenvalues of the audio are integrated together, and the energy coherence results, time domain coherence results, frequency domain coherence results, music theory coherence results and perception coherence results are comprehensively judged whether the audio data has frame loss, which greatly improves the accuracy of frame loss detection.
[0190] 2. In the present invention, differentiated processing is performed according to the degree of frame loss. If the degree of frame loss is low, the audio data transmission quality can be reduced first, and priority can be given to checking whether the frame loss is caused by communication quality problems. If the degree of frame loss is high, frame supplementation processing is performed. Compared with frame supplementation, reducing the audio data transmission quality is more efficient in concealing frame loss, avoiding mechanical frame supplementation processing, which results in low processing efficiency.
[0191] 3. In the present invention, if the audio is of a musical nature, there will often be some repeated tunes. In the process of frame supplementation, it is possible to preferentially search whether there is audio data with very similar energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perceptual eigenvalues in the audio before the frame loss position, and preferentially use similar data for frame supplementation. If no similar audio exists, the audio contour is extracted for frame supplementation, and the audio waveform is kept consistent as much as possible to repair the distortion of the audio in auditory perception.
[0192] Example 3
[0193] In one embodiment, the reference specification Figure 7 An exemplary embodiment of the present invention is a Bluetooth headset, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for processing audio frame loss in Example 1 when executing the computer program.
[0194] The Bluetooth headset may include a processor 301 and a memory 302 storing computer program instructions.
[0195] Specifically, the processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiment of the present invention.
[0196] Among them, the memory 302 may include a large capacity memory for data or instructions. For example, but not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 302 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 302 may be inside or outside the data processing device. In a specific embodiment, the memory 302 is a non-volatile memory. In a specific embodiment, the memory 302 includes a read-only memory (ROM) and a random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (Programmable Read-Only Memory, PROM for short), an erasable PROM (Erasable Programmable Read-Only Memory, EPROM for short), an electrically erasable PROM (Electrically Erasable Programmable Read-Only Memory, EEPROM for short), an electrically alterable ROM (Electrically Alterable Read-Only Memory, EAROM for short) or a flash memory (FLASH) or a combination of two or more of these. Under appropriate circumstances, the RAM can be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM can be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0197] The memory 302 may be used to store or cache various data files that need to be processed and / or used for communication, as well as possible computer program instructions executed by the processor 301 .
[0198] The processor 301 implements any one of the audio frame loss processing methods in Embodiment 1 by reading and executing computer program instructions stored in the memory 302 .
[0199] In some embodiments, the Bluetooth headset may further include a communication interface 303 and a bus 300. Figure 7 As shown, the processor 301, the memory 302, and the communication interface 303 are connected via a bus 300 and communicate with each other.
[0200] The communication interface 303 is used to implement communication between the modules, devices, units and / or equipment in the embodiment of the present invention. The communication interface 303 can also implement data communication with other components such as: external devices, image / data acquisition equipment, databases, external storage, and image / data processing workstations.
[0201] The bus 300 includes hardware, software or both, and couples the components of the Bluetooth headset to each other. The bus 300 includes but is not limited to at least one of the following: a data bus, an address bus, a control bus, an expansion bus, and a local bus. By way of example and not limitation, bus 300 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of the above. Bus 300 may include one or more buses, where appropriate. Although embodiments of the present invention describe and illustrate a particular bus, the present invention contemplates any suitable bus or interconnect.
[0202] Compared with the prior art, the present invention has at least the following beneficial effects:
[0203] 1. In the present invention, the external manifestations of the device audio frame loss are fully considered, the energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perception eigenvalues of the audio are integrated together, and the energy coherence results, time domain coherence results, frequency domain coherence results, music theory coherence results and perception coherence results are comprehensively judged whether the audio data has frame loss, which greatly improves the accuracy of frame loss detection.
[0204] 2. In the present invention, differentiated processing is performed according to the degree of frame loss. If the degree of frame loss is low, the audio data transmission quality can be reduced first, and priority can be given to checking whether the frame loss is caused by communication quality problems. If the degree of frame loss is high, frame supplementation processing is performed. Compared with frame supplementation, reducing the audio data transmission quality is more efficient in concealing frame loss, avoiding mechanical frame supplementation processing, which results in low processing efficiency.
[0205] 3. In the present invention, if the audio is of a musical nature, there will often be some repeated tunes. In the process of frame supplementation, it is possible to preferentially search whether there is audio data with very similar energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perceptual eigenvalues in the audio before the frame loss position, and preferentially use similar data for frame supplementation. If no similar audio exists, the audio contour is extracted for frame supplementation, and the audio waveform is kept consistent as much as possible to repair the distortion of the audio in auditory perception.
[0206] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0207] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for those of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A method for processing audio frame loss, applied to Bluetooth headsets, It is characterized in that A Bluetooth communication connection is established between the Bluetooth headset and the external device, and the method for processing audio frame loss includes: S101: Acquire audio data from the external device; S102: Extracting features of the audio data frame by frame to obtain energy feature values, time domain feature values, frequency domain feature values, music theory feature values, and perception feature values corresponding to the frame data of each frame; S103: Calculate the energy coherence parameter value between every two frames of the frame data according to the energy characteristic value, and when the energy coherence parameter value is less than a first preset value, set the energy coherence result to 1 to indicate that the energy characteristic is coherent, otherwise set the energy coherence result to 0 to indicate that the energy characteristic is incoherent; in this way, calculate the time domain coherence result, the frequency domain coherence result, the music theory coherence result and the perceptual coherence result; S104: Calculate a coherence result between the two frames of frame data, where the coherence result is the sum of the energy coherence result, the time domain coherence result, the frequency domain coherence result, the music theory coherence result and the perception coherence result. If the coherence result is less than a second preset value, determine that there is frame loss between the two frames of frame data, and set the frame loss result between the two frames of frame data to 1, otherwise, set it to 0; S105: taking a preset number of the frame data as a group, calculating a group frame loss result, wherein the group frame loss result is the sum of the frame loss results in the group; S106: if the group frame loss result is greater than a third preset value and less than a fourth preset value, reduce the quality of the audio data in a preset order, wherein the preset order is premium quality, lossless quality, high quality, and standard quality; S107: when the group frame loss result is greater than or equal to the fourth preset value, determining the number of frames N that need to be supplemented between the first frame data and the second frame data according to the first frame data and the second frame data between which frame loss occurs; S108: searching whether there is similar frame data that meets the conditions before the first frame data, wherein the similarity between the energy characteristic value, the time domain characteristic value, the frequency domain characteristic value, the music theory characteristic value and the perception characteristic value of the similar frame data and the first frame data is within a preset range; S109: if there is similar frame data that meets the condition, copy N frames of data following the similar frame data and insert them after the first frame of data; S110: In the case that similar frame data meeting the conditions does not exist, extracting energy feature values of each frame of the frame data to form an energy feature curve, extracting time domain feature values of each frame of the frame data to form a time domain feature curve, and extracting frequency domain feature values of each frame of the frame data to form a frequency domain feature curve; S111: In the energy characteristic curve, the time domain characteristic curve and the frequency domain characteristic curve, N horizontal coordinates are inserted between the horizontal coordinates corresponding to the first frame data and the horizontal coordinates corresponding to the second frame data, and the curve is fitted by using the least square method according to the known values in the curve to obtain the target energy characteristic value, the target time domain characteristic value and the target frequency domain characteristic value corresponding to the inserted N horizontal coordinates; S112: copying the first frame data, and transforming the first frame data so that the transformed frame data reaches the target energy characteristic value, the target time domain characteristic value, and the target frequency domain characteristic value; S113: inserting the transformed frame data between the first frame data and the second frame data.
2. The method for processing audio frame loss according to claim 1, It is characterized in that S102 specifically includes: S1021: performing windowing and frame division on the audio data, and extracting the energy eigenvalue and the time domain eigenvalue on a frame basis; S1022: Performing short-time Fourier transform on the audio data to obtain a short-time spectrum, and extracting the frequency domain eigenvalues on a frame-by-frame basis from the short-time spectrum; S1023: extracting a fundamental frequency from the short-time spectrum, and calculating the music theory characteristic value according to the overtone frequency of the frame data and the fundamental frequency; S1024: Extracting the perceptual feature value frame by frame from the audio data using an auditory perception model.
3. The method for processing audio frame loss according to claim 1, It is characterized in that The energy characteristic value is an amplitude, the time domain characteristic value is an autocorrelation value, the frequency domain characteristic value is a spectrum centroid value, the music theory characteristic value is a detuning value, and the perception characteristic value is a loudness value.
4. The method for processing audio frame loss according to claim 1, It is characterized in that The S103 specifically includes: S1031: Calculate the average value of the energy characteristic values of the frame data from the first frame to the Mth frame; S1032: using the energy characteristic value of the frame data of the M+1th frame minus the average value of the energy characteristic value of the previous M frames of data as the energy coherence parameter value between the Mth frame data and the M+1th frame data; S 1033: When the energy continuity parameter value is less than the first preset value, set the energy continuity result to 1; otherwise, set the energy continuity result to 0.
5. The method for processing audio frame loss according to claim 1, It is characterized in that The S107 specifically includes: S1071: extracting the perceptual feature values of the frame data of each frame to form a perceptual feature matrix S1; S1072: In the perceptual feature matrix S1, calculate the difference between the perceptual feature values of every two frames of the frame data to form a difference matrix S2; S1073: Remove the elements whose values are greater than the fifth preset value in the difference matrix S2 to form a matrix S3; S1074: Calculate the average value of the elements in the matrix S3 as the frame filling parameter value α; S1075: Let the perceptual feature value of the first frame of data be x 1 , the perceptual feature value of the second frame data is x 2 , then the number of frames N that need to be added between the first frame data and the second frame data can be calculated by the following formula: in, Indicates rounding up.
6. The method for processing audio frame loss according to claim 1, It is characterized in that Between S101 and S102, the method further includes: S114: Identify silent segments in the audio data, and delete the silent segments.
7. The method for processing audio frame loss according to claim 1, It is characterized in that After S113, the method further includes: S115: performing cross-fade processing on the first frame data, the inserted N frames of data, and the second frame data.
8. An audio frame loss processing device, applied to Bluetooth headsets, It is characterized in that A Bluetooth communication connection is established between the Bluetooth headset and the external device, and the audio frame loss processing device includes: An acquisition module, used for acquiring audio data from the external device; A first extraction module is used to perform feature extraction on the audio data frame by frame to obtain energy eigenvalues, time domain eigenvalues, frequency domain eigenvalues, music theory eigenvalues and perception eigenvalues corresponding to the frame data of each frame; A first calculation module is used to calculate the energy coherence parameter value between every two frames of the frame data according to the energy characteristic value, and when the energy coherence parameter value is less than a first preset value, the energy coherence result is set to 1 to indicate that the energy characteristic is coherent, otherwise the energy coherence result is set to 0 to indicate that the energy characteristic is incoherent; according to this method, the time domain coherence result, the frequency domain coherence result, the music theory coherence result and the perceptual coherence result are calculated; A second calculation module is used to calculate a coherence result between two frames of frame data, where the coherence result is the sum of the energy coherence result, the time domain coherence result, the frequency domain coherence result, the music theory coherence result and the perception coherence result. If the coherence result is less than a second preset value, it is determined that there is frame loss between the two frames of frame data, and the frame loss result between the two frames of frame data is set to 1, otherwise it is set to 0; A third calculation module is used to calculate a group frame loss result by taking a preset number of frame data as a group, wherein the group frame loss result is the sum of the frame loss results in the group; a reducing module, configured to reduce the quality of the audio data in a preset order when the group frame loss result is greater than a third preset value and less than a fourth preset value, wherein the preset order is premium quality, lossless quality, high quality, and standard quality; A determination module, configured to determine, when the group frame loss result is greater than or equal to the fourth preset value, the number of frames N that need to be supplemented between the first frame data and the second frame data according to the first frame data and the second frame data between which frame loss occurs; A retrieval module, used to retrieve whether there is similar frame data that meets the conditions before the first frame data, wherein the similarity between the energy characteristic value, the time domain characteristic value, the frequency domain characteristic value, the music theory characteristic value and the perception characteristic value of the similar frame data and the first frame data is within a preset range; A first copying module is used for copying N frames of data following the similar frame data and inserting them after the first frame of data when there is the similar frame data meeting the condition; A second extraction module is used for extracting energy characteristic values of each frame of the frame data to form an energy characteristic curve, extracting time domain characteristic values of each frame of the frame data to form a time domain characteristic curve, and extracting frequency domain characteristic values of each frame of the frame data to form a frequency domain characteristic curve when there is no similar frame data meeting the conditions; A first insertion module is used to insert N horizontal coordinates between the horizontal coordinate corresponding to the first frame data and the horizontal coordinate corresponding to the second frame data in the energy characteristic curve, the time domain characteristic curve and the frequency domain characteristic curve, and to perform curve fitting using the least squares method according to known values in the curve to obtain target energy characteristic values, target time domain characteristic values and target frequency domain characteristic values corresponding to the inserted N horizontal coordinates; A second copying module is used to copy the first frame data and transform the first frame data so that the transformed frame data reaches the target energy characteristic value, the target time domain characteristic value and the target frequency domain characteristic value; The second inserting module is used to insert the transformed frame data between the first frame data and the second frame data.
9. The apparatus for processing audio frame loss according to claim 8, It is characterized in that The first extraction module specifically includes: An energy time domain feature extraction submodule, used for windowing and framing the audio data, and extracting the energy feature value and the time domain feature value by frame; A frequency domain feature extraction submodule, used for performing short-time Fourier transform on the audio data to obtain a short-time spectrum, and extracting the frequency domain feature value on a frame-by-frame basis from the short-time spectrum; A music theory feature extraction submodule, used for extracting a fundamental frequency from the short-time spectrum, and calculating the music theory feature value according to the overtone frequency of the frame data and the fundamental frequency; The perceptual feature extraction submodule is used to extract the perceptual feature value for the audio data frame by frame through an auditory perception model.
10. A Bluetooth headset, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the method for processing audio frame loss according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and apparatus realizing compensation of frame loss in audio stream
CN104978966A
Frame loss judgment method and device, storage server and readable storage medium
CN114528171A