Method and system for mixing audio signals
By analyzing the audio signal quality through DSP circuits and machine learning models, the problem of audio interruption when the radio station switches signals is solved, achieving a smoother audio playback experience.
Patent Information
- Application Number
- CN202510446406.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-16
- Filing Date
- 2025-04-10
- Publication Date
- 2025-10-24
AI Technical Summary
When a radio station switches FM signals in its receiver, the audio quality often degrades and interruptions occur due to changes in signal strength, affecting the user experience.
Using digital signal processing (DSP) circuitry, a trained machine learning model is used to extract the parameters and quality scores of audio blocks, analyze and mix the audio blocks to output a high-quality mixed signal, and switch based on the audio quality score rather than signal strength.
By predicting audio quality changes early, it reduces interruptions caused by signal switching and improves the smoothness of the user's listening experience.
Smart Images

Figure CN120834880A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to frequency modulation (FM) communication, and more specifically to systems and methods for mixing audio signals. BACKGROUND
[0002] Radio stations often broadcast a single program simultaneously with alternate frequency modulation (FM) frequencies, such as a primary FM signal and an alternate FM signal. Based on the power strength of each of the primary FM signal and the alternate FM signal remaining above an acceptable power level, one of the primary FM signal or the alternate FM signal is selected as an output signal, which is played as audio by a dual FM radio receiver of a user. The audio quality of the FM signal often degrades before the power strength of the FM signal falls below an unacceptable level. In a scenario, the power strength of the primary FM signal and the power strength of the alternate FM signal are above the acceptable power level. However, the audio quality of the primary FM signal has degraded, while the audio quality of the alternate FM signal remains above an acceptable quality level. In this scenario, when the primary FM signal is selected as the output signal, multiple interruptions can occur in the played audio due to the degraded audio quality of the primary FM signal compared to the alternate FM signal. As a result, the listening experience of the user is affected. SUMMARY
[0003] In an embodiment of the present disclosure, a system for mixing audio signals is disclosed. The system can include a digital signal processing (DSP) circuit. The DSP circuit can be configured to extract a plurality of first audio parameters associated with a plurality of first audio blocks of a first audio signal and a plurality of second audio parameters associated with a plurality of second audio blocks of a second audio signal by executing a trained machine learning model. The system can be configured to receive the first audio signal and the second audio signal. The DSP circuit can be additionally configured to process the plurality of first audio parameters and the plurality of second audio parameters to generate a plurality of first audio quality scores and a plurality of second audio quality scores, respectively, by additionally executing the trained machine learning model, wherein each audio quality score of the plurality of first audio quality scores and the plurality of second audio quality scores can indicate an audio quality of a corresponding audio block of the plurality of first audio blocks and the plurality of second audio blocks, respectively. The DSP circuit can be additionally configured to analyze the plurality of first audio quality scores and the plurality of second audio quality scores. The DSP circuit can be additionally configured to output a plurality of mixed blocks upon analyzing the plurality of first audio quality scores and the plurality of second audio quality scores, wherein each of the plurality of mixed blocks can be at least one of a first audio block of the plurality of first audio blocks and a second audio block of the plurality of second audio blocks.
[0004] In some embodiments, the DSP circuit can be additionally configured to train the first machine learning model based on training data to obtain a trained machine learning model, where the training data can include a plurality of test audio recordings and a plurality of quality scores, such that the plurality of quality scores can include a first quality score for a first test audio recording of the plurality of test audio recordings, where each of the plurality of quality scores can indicate an audio quality of a corresponding test audio recording, and where a low score can indicate a low quality of a test audio recording of the plurality of test audio recordings, and a high score can indicate a high quality of the test audio recording.
[0005] In some embodiments, to train the first machine learning model, the DSP circuit can be additionally configured to extract a first plurality of training parameters for the first test audio recording. The DSP circuit can be additionally configured to determine a first test score based on processing of the first plurality of training parameters by way of a test scoring operation. The DSP circuit can be additionally configured to compare the first test score to the first quality score to determine a match between the first test score and the first quality score, where a match between the first test score and the first quality score can indicate to the first machine learning model that the determination of the first test score by way of the test scoring operation can be accurate, and a mismatch between the first test score and the first quality score can indicate to the first machine learning model that the determination of the first test score by way of the test scoring operation can be erroneous. The DSP circuit can be additionally configured to update the test scoring operation until a match between the first test score and the first quality score can be determined, where the first machine learning model can be trained based on the match between the first test score and the first quality score, and where the trained machine learning model can generate the plurality of first audio quality scores and the plurality of second audio quality scores based on the training of the first machine learning model.
[0006] In some embodiments, the DSP circuit can be additionally configured to perform a time-to-frequency domain conversion operation on the plurality of first audio blocks and the plurality of second audio blocks to generate a plurality of third audio blocks and a plurality of fourth audio blocks, respectively, where the plurality of third audio blocks and the plurality of fourth audio blocks in the frequency domain can be provided to the trained machine learning model to extract the plurality of first audio parameters and the plurality of second audio parameters, respectively.
[0007] In some embodiments, the system can additionally include a first receiver and a second receiver that can be configured to receive the first audio signal and the second audio signal from an audio source. The first and second receivers can be additionally configured to convert each of the first audio signal and the second audio signal to a digitized version of each of the first audio signal and the second audio signal, respectively.
[0008] In some embodiments, the system can additionally include a first buffer and a second buffer coupled to the first receiver and the second receiver, respectively, wherein the first buffer and the second buffer are configured to receive digitized versions of each of the first audio signal and the second audio signal from the first receiver and the second receiver, respectively. The first buffer and the second buffer can additionally be configured to store the digitized versions of each of the first audio signal and the second audio signal, respectively.
[0009] In some embodiments, wherein the DSP circuit can additionally be configured to read the plurality of first audio blocks and the plurality of second audio blocks from the first buffer and the second buffer, respectively, wherein the plurality of first audio parameters and the plurality of second audio parameters can be extracted after the plurality of first audio blocks and the plurality of second audio blocks are read, respectively.
[0010] In some embodiments, the system can additionally include a host processor, wherein the DSP circuit can additionally be configured to receive a first delay value from the host processor prior to analyzing the plurality of first audio quality scores and the plurality of second audio quality scores, the first delay value can be indicative of an expected delay between the plurality of first audio blocks and the plurality of second audio blocks. The DSP circuit can additionally be configured to convolve the plurality of first audio blocks and the plurality of second audio blocks based on a group consisting of the first delay value, the plurality of first audio quality scores, and the plurality of second audio quality scores to determine a second delay value that can be indicative of an actual delay between the plurality of first audio blocks and the plurality of second audio blocks.
[0011] In some embodiments, the DSP circuit can additionally be configured to generate a ready signal based on the second delay value. The DSP circuit can additionally be configured to provide the ready signal to the host processor, wherein the ready signal can be indicative of a request to initiate analysis of the plurality of first audio blocks and the plurality of second audio blocks. The DSP circuit can additionally be configured to receive a confirmation signal from the host processor based on the ready signal, wherein the confirmation signal can be indicative of initiation of analysis of the plurality of first audio blocks and the plurality of second audio blocks.
[0012] In some embodiments, the DSP circuit can additionally be configured to tune one of the first audio blocks and a corresponding second audio block of the plurality of second audio blocks based on detecting that the one of the first audio blocks and the corresponding second audio block can be out of sync due to the second delay value. The DSP circuit can additionally be configured to mix the one of the first audio blocks and the corresponding second audio block to output a mixed block of a plurality of mixed blocks, wherein the one of the first audio blocks can be mixed with the corresponding second audio block after tuning of the one of the first audio blocks and the corresponding second audio block.
[0013] In some embodiments, to analyze the plurality of first audio quality scores and the plurality of second audio quality scores, the DSP circuit can be further configured to identify whether each audio quality score of the plurality of first audio quality scores and each corresponding audio quality score of the plurality of second audio quality scores can be greater than a threshold score, wherein (i) the plurality of first audio blocks can include a first audio block and a second audio block such that the second audio block can be after the first audio block, (ii) the plurality of second audio blocks can include a third audio block such that the third audio block can correspond to the second audio block, (iii) the plurality of first audio quality scores can include a first audio quality score of the first audio block and a second audio quality score of the second audio block, and (iv) the plurality of second audio quality scores can include a third audio quality score of the third audio block, wherein upon identifying that at least one audio quality score of the plurality of first audio quality scores and the plurality of second audio quality scores can be greater than the threshold score, the DSP circuit can be further configured to detect at least one previous mixed block of the plurality of mixed blocks to output one of the plurality of mixed blocks.
[0014] In some embodiments, wherein upon identifying that (i) the second audio quality score and the third audio quality score can be greater than the threshold score and (ii) the third audio quality score can be greater than the second audio quality score, the DSP circuit can detect a previous mixed block of the plurality of mixed blocks, and wherein upon detecting that the first audio block can be the previous mixed block, the second audio block can be output as a current mixed block of the plurality of mixed blocks.
[0015] In some embodiments, upon identifying that (i) the first audio quality score and the second audio quality score can be lower than the threshold score and the third audio quality score can be greater than the threshold score, the DSP circuit can detect a previous sub-plurality of mixed blocks of the plurality of mixed blocks, wherein when the previous sub-plurality of mixed blocks can be detected to be a sub-plurality of audio blocks of the plurality of first audio blocks such that the sub-plurality of audio blocks (i) can include the first audio block and (ii) can have audio quality scores identified to be lower than the threshold score, the DSP circuit can mix the second audio block and the third audio block to output a current mixed block of the plurality of mixed blocks.
[0016] In some embodiments, each of the plurality of first audio parameters and the plurality of second audio parameters can include a group consisting of a spectral centroid, a spectral flux, and a noise floor of each of the plurality of first audio blocks and the plurality of second audio blocks, respectively.
[0017] In some embodiments, the data associated with the first audio signal and the data associated with the second audio signal can be substantially the same.
[0018] In another embodiment of the disclosure, a method is disclosed. The method can include extracting, by a digital signal processing (DSP) circuit, a plurality of first audio parameters associated with a plurality of first audio blocks of a first audio signal and a plurality of second audio parameters associated with a plurality of second audio blocks of a second audio signal by executing a trained machine learning model. The method can additionally include processing, by the DSP circuit, the plurality of first audio parameters and the plurality of second audio parameters to generate a plurality of first audio quality scores and a plurality of second audio quality scores, respectively, by additionally executing the trained machine learning model, wherein each of the plurality of first audio quality scores and the plurality of second audio quality scores can indicate an audio quality of a corresponding audio block of the plurality of first audio blocks and the plurality of second audio blocks, respectively. The method can additionally include analyzing, by the DSP circuit, the plurality of first audio quality scores and the plurality of second audio quality scores. The method can additionally include outputting, by the DSP circuit, a plurality of mixed blocks after analyzing the plurality of first audio quality scores and the plurality of second audio quality scores, wherein each of the plurality of mixed blocks can be at least one of a first audio block of the plurality of first audio blocks and a second audio block of the plurality of second audio blocks.
[0019] In some embodiments, the method can additionally include training, by the DSP circuit, the first machine learning model based on training data to obtain the trained machine learning model, wherein the training data can include a plurality of test audio recordings and a plurality of quality scores, such that the plurality of quality scores can include a first quality score of a first test audio recording of the plurality of test audio recordings, wherein each of the plurality of quality scores can indicate an audio quality of a corresponding test audio recording, and wherein a low score can indicate a low quality of a test audio recording of the plurality of test audio recordings and a high score can indicate a high quality of the test audio recording.
[0020] In some embodiments, the method can additionally include identifying, by the DSP circuit, whether each of the plurality of first audio quality scores and each of the plurality of second audio quality scores is greater than a threshold score, wherein (i) the plurality of first audio blocks can include a first audio block and a second audio block, such that the second audio block can be after the first audio block, (ii) the plurality of second audio blocks can include a third audio block, such that the third audio block can correspond to the second audio block, (iii) the plurality of first audio quality scores can include a first audio quality score of the first audio block and a second audio quality score of the second audio block, and (iv) the plurality of second audio quality scores can include a third audio quality score of the third audio block. The method can additionally include detecting, by the DSP circuit, at least one previous mixed block of the plurality of mixed blocks to output one of the plurality of mixed blocks after identifying that at least one of the plurality of first audio quality scores and the plurality of second audio quality scores can be greater than the threshold score.
[0021] In some embodiments, upon identifying that (i) the second audio quality score and the third audio quality score can be greater than the threshold score and (ii) the third audio quality score can be greater than the second audio quality score, a previous mixing block of the plurality of mixing blocks can be detected by the DSP circuit, and wherein upon detecting that the first audio block can be the previous mixing block, the second audio block can be output as a current mixing block of the plurality of mixing blocks.
[0022] In some embodiments, the method can additionally include mixing, by the DSP circuit, the second audio block and the third audio block to output a current mixing block of the plurality of mixing blocks, wherein upon identifying that (i) the first audio quality score and the second audio quality score can be lower than the threshold score and the third audio quality score can be greater than the threshold score, and when a previous sub-plurality of mixing blocks of the plurality of mixing blocks is detected as a sub-plurality of audio blocks of the first plurality of audio blocks such that the sub-plurality of audio blocks (i) includes the first audio block and (ii) has an audio quality score that can be identified as lower than the threshold score, the second audio block and the third audio block can be mixed. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 A schematic diagram illustrating a frequency modulation (FM) environment according to example embodiments of the present disclosure;
[0024] Figure 2 A block diagram representing an example scenario of mixing of a plurality of first audio blocks and a plurality of second audio blocks by a digital signal processing (DSP) circuit illustrating Figure 1 according to embodiments of the present disclosure;
[0025] Figure 3A A block diagram representing training of a first machine learning model illustrating Figure 1 according to example embodiments of the present disclosure;
[0026] Figure 3B A block diagram representing implementation of a trained machine learning model obtained after training Figure 3A a first machine learning model according to example embodiments of the present disclosure;
[0027] Figure 4 A block diagram representing a DSP circuit of an FM environment illustrating Figure 1 according to example embodiments of the present disclosure;
[0028] Figures 5A-5B Collectively, a flowchart representing training of a first machine learning model illustrating Figure 3A according to embodiments of the present disclosure; and
[0029] Figures 6A-6F Collectively, a flowchart representing implementation of a trained machine learning model obtained after training Figure 4a flowchart of an audio mixing method performed by a DSP circuit. DETAILED DESCRIPTION
[0030] The detailed description of the drawings is intended to describe embodiments of the present disclosure and is not intended to represent the only forms in which the present disclosure can be practiced. It should be understood that the same or equivalent functions can be accomplished by different embodiments that are intended to be encompassed within the spirit and scope of the present disclosure.
[0031] SUMMARY
[0032] Generally, a dual frequency modulation (FM) radio receiver can receive both an FM signal and an alternative FM signal broadcast from an audio source. The data associated with the alternative FM signal can be equivalent to the data associated with the FM signal. Due to the gradual fading nature of the FM signal, the dual FM radio receiver can determine whether to switch from the FM signal to the alternative FM signal, or vice versa, based only on the received signal strength, such as the power level of the FM signal. In a scenario where the FM signal can have a low power level while the alternative FM signal can have an acceptable power level, the dual FM radio receiver can switch from the FM signal to the alternative FM signal. However, the audio quality of the alternative FM signal can have already degraded before the switch from the FM signal to the alternative FM signal. As a result, the played audio is distorted due to an interruption, such as muting or pausing, which in turn affects the listening experience of the user.
[0033] Various embodiments of the present disclosure disclose a system for mixing audio signals. The system can include a digital signal processing (DSP) circuit, a first receiver, a second receiver, a first buffer, a second buffer, and a host processor. The first receiver and the second receiver can receive a first audio signal and a second audio signal, respectively, from an audio source. The audio source can be a radio station that broadcasts a program by transmitting the first audio signal and the second audio signal. The first receiver and the second receiver can convert the first analog signal and the second analog signal from an analog format to a digital format, and can provide digitized versions of the first audio signal and the second audio signal to the first buffer and the second buffer, respectively. The first buffer and the second buffer can store the digitized versions of the first audio signal and the second audio signal as first audio blocks and second audio blocks of the first and second audio signals, respectively. The DSP circuit can read the first audio blocks and the second audio blocks from the first buffer and the second buffer, respectively. The DSP circuit can extract first audio parameters associated with the first audio blocks and second audio parameters associated with the second audio blocks by executing a trained machine learning model. The trained machine learning model can be additionally executed to process the first audio parameters and the second audio parameters to generate first audio quality scores and second audio quality scores, respectively. Each audio quality score can indicate an audio quality of a corresponding audio block of the first audio blocks and the second audio blocks. The DSP circuit can analyze each of the first audio quality scores and the second audio quality scores. After analyzing the first audio quality scores and the second audio quality scores, the DSP circuit can mix the first audio blocks and the second audio blocks to output mixed blocks. Each of the mixed blocks is at least one of the first audio blocks and the second audio blocks. Furthermore, the audio blocks can be tuned based on an expected delay between the first audio blocks and the second audio blocks before mixing the first audio blocks with the second audio blocks.
[0034] In the present disclosure, the DSP circuit can output the audio blocks based on the audio quality scores that indicate qualities of the audio blocks, as compared to conventional solutions that rely only on power levels of analog FM signals. Thus, the signals are mixed based on analysis of the first audio quality scores, the second audio quality scores, and previous mixed blocks. The analysis of the first audio quality scores and the second audio quality scores can enable the DSP circuit to provide early prediction of mixing, as the audio quality of the audio blocks can decrease before detecting that a power level of a corresponding FM signal is lower than a threshold power level. The early prediction of mixing can additionally enable the DSP circuit to play audio relatively more smoothly as compared to conventional solutions that rely only on power levels of analog FM signals. Thus, the listening experience of a user is enhanced.
[0035] Figure 1A schematic diagram illustrating a frequency modulation (FM) environment 100 in accordance with an embodiment of the present disclosure is shown. The FM environment 100 can include an audio source 102 and a system 104 for mixing audio signals, hereinafter referred to as "system 104." The system 104 can be incorporated into a device (not shown). Examples of the device can include a dual FM radio receiver. A user (not shown) can operate or own the dual FM radio receiver.
[0036] The FM environment 100 can additionally include a communication network 106. The audio source 102 can communicate with the system 104 by means of the communication network 106. Examples of the communication network 106 can include the Internet, a local area network (LAN), a wide area network (WAN), and the like.
[0037] Audio source 102:
[0038] The audio source 102 can include suitable circuitry that can be configured to perform one or more operations. For example, the audio source 102 can broadcast or transmit a first audio signal FS and a second audio signal SS to the system 104. Each of the first audio signal FS and the second audio signal SS can be an analog FM radio signal that can indicate a radio program such as music, news, a podcast, and the like. Thus, the data associated with the first audio signal FS and the data associated with the second audio signal SS can be essentially the same. The audio source can transmit the first audio signal FS and the second audio signal SS to the system 104 in an analog format by means of the communication network 106. The second audio signal SS can have an alternative frequency with respect to the first audio signal FS. The frequencies of the first audio signal FS and the second audio signal SS can be in the range of 76 megahertz (MHz) to 108 MHz, but the range can be lower, higher, or different. In an example, the frequency of the first audio signal FS can be 93 MHz and the frequency of the second audio signal SS can be 105 MHz. The audio source 102 can include a source processor 108, a first transmitter circuit 110, and a second transmitter circuit 112. The source processor 108, the first transmitter circuit 110, and the second transmitter circuit 112 can communicate with each other by means of a first communication channel 114. Examples of the first communication channel 114 can include a fiber optic cable, an Ethernet cable, a coaxial cable, and the like. In an exemplary embodiment, the audio source 102 can be a radio station.
[0039] Source processor 108:
[0040] The source processor 108 can comprise suitable circuitry that can be configured to perform one or more operations. For example, the source processor 108 can be configured to generate and transmit the first audio signal FS and the second audio signal SS to the system 104 by means of the first transmitter circuit 110 and the second transmitter circuit 112, respectively. The source processor 108 can be additionally configured to set the frequency of each of the first audio signal FS and the second audio signal SS. An example of the source processor 108 can be a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application specific integrated circuit (ASIC), or the like.
[0041] First transmitter circuit 110:
[0042] The first transmitter circuit 110 can comprise suitable circuitry that can be configured to perform one or more operations. For example, the first transmitter circuit 110 can be configured to transmit the first audio signal FS to the system 104. The first transmitter circuit 110 can transmit the first audio signal FS to the system 104 by means of the communication network 106. An example of the first transmitter circuit 110 can comprise an FM transmitter.
[0043] Second transmitter circuit 112:
[0044] The second transmitter circuit 112 can be structurally and functionally similar to the first transmitter circuit 110. The second transmitter circuit 112 can comprise suitable circuitry that can be configured to perform one or more operations. For example, the second transmitter circuit 112 can be configured to transmit the second audio signal SS to the system 104. The second transmitter circuit 112 can transmit the second audio signal SS to the system 104 by means of the communication network 106. An example of the second transmitter circuit 112 can comprise an FM transmitter.
[0045] In another example embodiment, the audio source 102 can be a broadcast station that can comprise the source processor 108, the first transmitter circuit 110, and the second transmitter circuit 112. The source processor 108 can be, for example, a stationary or mobile live unit that can transmit a first analog signal (not shown) to the first transmitter circuit 110 and the second transmitter circuit 112. For example, the data associated with the first analog signal can be a live broadcast of a football match. The source processor 108, the first transmitter circuit 110, and the second transmitter circuit 112 can communicate with each other by means of the first communication channel 114. The first transmitter circuit 110 and the second transmitter circuit 112 can be radio units that can broadcast the first analog signal by transmitting the first audio signal FS and the second audio signal SS, respectively, to a number of electronic devices, for example, devices of the system 104. The devices can receive the first audio signal FS and the second audio signal SS.
[0046] System 104:
[0047] The system 104 can include suitable circuitry that can be configured to perform one or more operations. For example, the system 104 can be configured to receive the first audio signal FS and the second audio signal SS and output a plurality of mixing blocks B1-BN. The plurality of mixing blocks B1-BN can correspond to an audio output signal played by the system 104. The system 104 can include a first receiver 116, a second receiver 118, a first buffer 120, a second buffer 122, a host processor 124, and a digital signal processing (DSP) circuit 126.
[0048] First receiver 116:
[0049] The first receiver 116 can be coupled to the first buffer 120 and the DSP circuit 126. The first receiver 116 can include suitable circuitry that can be configured to perform one or more operations. For example, the first receiver 116 can be configured to receive the first audio signal FS from the audio source 102. The first receiver 116 can be additionally configured to convert the first audio signal FS from an analog format to a digital format and provide a digitized version of the first audio signal FS (e.g., a plurality of first audio packets FP1-FPN) to the first buffer 120. In an example, the first receiver 116 can convert the first audio signal FS to a digital format through pulse code modulation (PCM). The first receiver 116 can be additionally configured to detect and provide a plurality of first power levels FV1-FVN of the first audio signal FS to the DSP circuit 126. Each power level of the plurality of first power levels FV1-FVN can indicate a power intensity of the first audio signal FS. The first receiver 116 can be additionally configured to provide the plurality of first power levels FV1-FVN to the DSP circuit 126 at predefined time intervals. In an example, the first receiver 116 can provide the plurality of first power levels FV1-FVN to the DSP circuit 126 every 10 milliseconds (ms), although the time intervals can be different. Each of the plurality of first power levels FV1-FVN can be in a range of 10 decibels (dB) to 100 dB, although the power levels can be lower or higher. In an example, a first power level FV1 of the plurality of first power levels FV1-FVN can be 30 dB. Examples of the first receiver 116 can include an FM receiver.
[0050] Second receiver 118:
[0051] The second receiver 118 can be structurally and functionally similar to the first receiver 116. The second receiver 118 can be coupled to the second buffer 122 and the DSP circuit 126. The second receiver 118 can include suitable circuitry that can be configured to perform one or more operations. For example, the second receiver 118 can be configured to receive the second audio signal SS from the audio source 102. The second receiver 118 can be additionally configured to convert the second audio signal SS from an analog format to a digital format and provide the digitized version of the second audio signal SS (e.g., the plurality of second audio packets SP1-SPN) to the second buffer 122. The second receiver 118 can be additionally configured to detect and provide the plurality of second power levels SV1-SVN of the second audio signal SS to the DSP circuit 126. The second receiver 118 can be additionally configured to provide the plurality of second power levels SV1-SVN to the DSP circuit 126 at predefined time intervals. Examples of the second receiver 118 can include an FM receiver.
[0052] First buffer 120:
[0053] The first buffer 120 can be coupled with the first receiver 116, the host processor 124, and the DSP circuit 126. The first buffer 120 can include suitable circuitry that can be configured to store data. For example, the first buffer 120 can be configured to receive the digitized version of the first audio signal FS (e.g., the plurality of first audio packets FP1-FPN) from the first receiver 116. The first buffer 120 can be additionally configured to store the digitized version of the first audio signal FS (e.g., the plurality of first audio packets FP1-FPN) as the plurality of first audio blocks FB1-FBN. In an example, an audio block in the plurality of first audio blocks FB1-FBN can include five audio packets in the plurality of first audio packets FP1-FPN, but the number of audio packets can be different. Examples of the first buffer 120 can include a random access memory (RAM), a read only memory (ROM), a removable storage drive, a hard disk drive (HDD), a flash memory, a solid state memory, and the like.
[0054] Second buffer 122:
[0055] The second buffer 122 can be structurally and functionally similar to the first buffer 120. The second buffer 122 can be coupled with the second receiver 118, the host processor 124, and the DSP circuit 126. The second buffer 122 can include suitable circuitry that can be configured to store data. For example, the second buffer 122 can be configured to receive the digitized version of the second audio signal SS (e.g., the plurality of second audio packets SP1-SPN) from the second receiver 118. The second buffer 122 can be additionally configured to store the digitized version of the second audio signal SS (e.g., the plurality of second audio packets SP1-SPN) as the plurality of second audio blocks SB1-SBN. In an example, the audio blocks in the plurality of second audio blocks SB1-SBN can include five audio packets of the plurality of second audio packets SP1-SPN, but the number of audio packets can vary. Examples of the second buffer 122 can include random access memory (RAM), read only memory (ROM), removable storage drives, hard disk drives (HDD), flash memory, solid state memory, and the like.
[0056] Host processor 124:
[0057] The host processor 124 can be coupled to the first buffer 120, the second buffer 122, and the DSP circuit 126. The host processor 124 can include suitable circuitry that can be configured to perform one or more operations. For example, the host processor 124 can be configured to read the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN from the first buffer 120 and the second buffer 122, respectively. The host processor 124 can be additionally configured to determine the first delay value FD based on the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN. The first delay value FD indicates the expected delay between the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN. The host processor 124 can determine the first delay value FD based on historical data stored in a memory (not shown) associated with the host processor 124. In an embodiment, the historical data can include expected delay values between FM signals based on the geographical location of the radio station (e.g., the audio source 102). In an example, when the radio station (e.g., the audio source 102) can be close to a mountain, the expected delay between the FM signals can be 20 ms, whereas when the radio station (e.g., the audio source 102) can be close to a river bed, the expected delay between the FM signals can be 10 ms. The host processor 124 can be additionally configured to receive the ready signal RS from the DSP circuit 126. The ready signal RS can indicate to initiate the analysis of the plurality of first audio quality scores (shown in Figure 3B Figure 3B Figure 3B each of the plurality of first audio quality scores (shown in Figure 3B each of the plurality of first audio quality scores (shown in Figure 3B each of the plurality of first audio quality scores (shown in Figure 3B each of the plurality of first audio quality scores (shown in Figure 3B each of the plurality of first audio quality scores (shown in Figure 3B In an example, the host processor 124 can generate the confirmation signal CS based on determining that a previous function request of the system 104 is being executed. In another example, the host processor 124 can delay generation of the confirmation signal CS based on determining that an ongoing function request of the system 104 is being executed.
[0058] DSP circuit 126:
[0059] The DSP circuit 126 can be coupled to the first receiver 116, the second receiver 118, the first buffer 120, the second buffer 122, and the host processor 124. The DSP circuit 126 can include suitable circuitry that can be configured to perform one or more operations. For example, the DSP circuit 126 can be configured to read the plurality of first audio blocks FBl-FBN of the first audio signal FS and the plurality of second audio blocks SB 1-SBN of the second audio signal SS from the first buffer 120 and the second buffer 122, respectively. The DSP circuit 126 can be further configured to read the first buffer 120 and the second buffer 122 at predefined time intervals. In an example, the DSP circuit 126 can read the first buffer 120 and the second buffer 122 every 10 ms, although the time intervals can be different. The DSP circuit 126 can be further configured to read a sub-plurality of first audio blocks of the plurality of first audio packets FPl-FPN and a sub-plurality of second audio blocks of the plurality of second audio packets SPl-SPN. The sub-plurality of first audio blocks and the sub-plurality of second audio blocks are stored by the first buffer 120 and the second buffer 122 as corresponding audio blocks of the plurality of first audio blocks FBl-FBN and the plurality of second audio blocks SB 1-SBN, respectively, at the predefined time intervals. In an example, the audio blocks in the plurality of first audio blocks FBl-FBN and the plurality of second audio blocks SB 1-SBN can include five audio packets of the plurality of first audio packets FPl-FPN and the plurality of second audio packets SPl-SPN, respectively, although the number of audio blocks can be different.
[0060] The number of audio packets in an audio block (e.g., the length of an audio block) in the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN can be based on the frequency of the first audio signal FS and the second audio signal SS, respectively. For example, when the frequency of the first audio signal FS is 44.1 kilohertz (kHz), the number of audio packets in an audio block in the plurality of first audio blocks FB1-FBN can be 512, thereby achieving a higher frequency resolution. In such an example, the frequency resolution of the first audio signal FS can be 43 hz per frequency bin. The DSP circuit 126 can extract the plurality of first audio parameters and the plurality of second audio parameters after reading the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN, respectively.
[0061] The DSP circuit 126 can be additionally configured to perform a time-to-frequency domain conversion operation on the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN to generate a plurality of third audio blocks and a plurality of fourth audio blocks, respectively. In an example, the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN are in the time domain, and the plurality of third audio blocks and the plurality of fourth audio blocks are in the frequency domain. Prior to performing the time-to-frequency domain conversion operation, the DSP circuit 126 can be additionally configured to generate a plurality of first intermediate blocks (not shown) and a plurality of second intermediate blocks (not shown) based on the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN, respectively. Each of the plurality of first intermediate blocks and the plurality of second intermediate blocks can be a short audio block.
[0062] To generate the plurality of first intermediate blocks, the DSP circuit 126 can be additionally configured to combine (e.g., overlap) each of the plurality of first audio blocks FB1-FBN. The DSP circuit 126 can combine a current audio block in the plurality of first audio blocks FB1-FBN with one of a previous audio block in the plurality of first audio blocks FB1-FBN and a subsequent audio block in the plurality of first audio blocks FB1-FBN to generate a corresponding intermediate block. For example, the DSP circuit 126 can combine the first audio block FB1 with the second audio block FB2 to generate a first intermediate block in the plurality of first intermediate blocks. Additionally, the DSP circuit 126 can combine the second audio block FB2 with the third audio block FB3 to generate a second intermediate block in the plurality of first intermediate blocks. Similarly, the DSP circuit 126 can generate the plurality of second intermediate blocks in a similar manner as the plurality of first intermediate blocks.
[0063] The DSP circuit 126 can be further configured to perform a sine window operation on the plurality of first intermediate blocks and the plurality of second intermediate blocks to reduce window artifacts that can occur during the time-frequency domain conversion operation performed on the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN. The window artifacts can be reduced by smoothing the boundary frequencies of each of the plurality of first intermediate blocks and the plurality of second intermediate blocks. The time-frequency domain conversion operation is then performed on the plurality of first intermediate blocks and the plurality of second intermediate blocks to generate a plurality of third audio blocks and a plurality of fourth audio blocks, respectively.
[0064] The plurality of third audio blocks and the plurality of fourth audio blocks can be provided to a trained machine learning model (shown in FIG. 1) to extract a plurality of first audio parameters and a plurality of second audio parameters, respectively. Examples of the time-frequency domain conversion operation include a Fast Fourier Transform, a Cosine Transform, and the like. Figure 3B
[0065] The DSP circuit 126 can be further configured to extract a plurality of first audio parameters associated with the plurality of first audio blocks FB1-FBN (e.g., the plurality of third audio blocks) of the first audio signal FS and a plurality of second audio parameters associated with the plurality of second audio blocks SB1-SBN (e.g., the plurality of fourth audio blocks) of the second audio signal SS by performing a trained machine learning model (shown in FIG. 1). Figure 3B Figure 3B To obtain the trained machine learning model (shown in FIG. 1), the DSP circuit 126 can be further configured to train a first machine learning model (shown in FIG. 1). Each of the plurality of first and second audio parameters can include a spectral centroid, a spectral flux, and a noise floor of each of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN, respectively. Figure 3A
[0066] The spectral centroid can determine an average frequency of each of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN. The spectral centroid of each of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN can be extracted by Equation (1):
[0067]
[0068] where 'n' is a total number of frequency bins (e.g., a total number of different frequencies present in the audio blocks of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN), 'fi' is a frequency of the ith frequency bin, and 'A I' is a magnitude (or amplitude) of the ith frequency bin.
[0069] The values of the spectral centroids of the audio blocks in the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN can indicate presence of noise in the audio blocks. High values of the spectral centroids can indicate that the audio blocks can be affected by noise, and low values of the spectral centroids can indicate that the audio blocks can remain unaffected by noise. Thus, low spectral centroids of the audio blocks are desired. Accordingly, the values of the spectral centroids of each of the plurality of mixing blocks B1-BN can be lower.
[0070] The spectral flux (SF) can quantify the rate at which the energy distribution of the frequency bands of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN can change over a certain time period. The spectral flux can be extracted by Equation (2):
[0071]
[0072] where 'K' is the total number of frequency bins (e.g., the total number of different frequencies present in the audio blocks of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN), and 'S[k, n]' is the magnitude of the kth frequency bin in the nth time frame. The Nth time frame corresponds to a predefined number of time blocks required to determine the spectral flux.
[0073] High spectral flux of an audio block can indicate several transitions in the audio block compared to (i) the corresponding previous block of one of the plurality of first audio blocks FB1-FBN and (ii) the plurality of second audio blocks SB1-SBN. Low spectral flux can indicate similar or zero transitions or events compared to the previous block. Additionally, low spectral flux of an audio block can indicate that the audio block can be affected by noise. Thus, high spectral flux of the audio blocks is desired.
[0074] To determine the noise floor of the audio blocks, the DSP circuit 126 can be additionally configured to determine a global minimum and a local minimum of the plurality of first audio blocks FB1-FBN and the plurality of second blocks SB1-SBN. The value of the global minimum can indicate an amount of comfort noise that can be present in the first audio signal FS and the second audio signal SS. The comfort noise of the first audio signal FS and the second audio signal SS can indicate background noise (e.g., environmental noise) that can be added to the first audio signal FS and the second audio SS in the absence of audio in the first audio signal FS and the second audio signal SS. The DSP circuit 126 can update the value of the global minimum in real-time upon receiving each of the plurality of first audio blocks FB1-FBN and each of the plurality of second audio blocks SB1-SBN at a given point in time.
[0075] The local minimum can be indicative of a total amount of noise present in the first audio signal FS and the second audio signal SS in a predefined time interval, e.g., 3 seconds. Thus, the local minimum can be based on an amount of environmental interference or other unwanted artifacts. The DSP circuit 126 can determine the local minimum by detecting a minimum frequency of each frequency band of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN at each predefined time interval, e.g., 3 seconds, to avoid false peaks. Additionally, the DSP circuit 126 can determine the local minimum by determining a minimum frequency of a selective frequency band range of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN that can be affected by noise, e.g., a frequency band of 5 kHz to 19 kHz.
[0076] The DSP circuit 126 can determine a noise floor of an audio block of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN by subtracting the global minimum from the local minimum of the audio block. A high noise floor can be indicative that the audio block can be affected by noise. Thus, a low noise floor of the audio block is desired.
[0077] The DSP circuit 126 can additionally be configured to process the plurality of first audio parameters and the plurality of second audio parameters by additionally executing a trained machine learning model (shown in Figure 3B ) to generate a plurality of first audio quality scores (shown in Figure 3B ) and a plurality of second audio quality scores (shown in Figure 3B ), respectively. Each audio quality score of the plurality of first audio quality scores (shown in Figure 3B ) and the plurality of second audio quality scores (shown in Figure 3B ) can be indicative of an audio quality of a corresponding audio block of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN, respectively. The plurality of first audio quality scores (shown in Figure 3B ) and the plurality of second audio quality scores (shown in Figure 3BThe ITU five-grade impairment scale can be used in the evaluation of the audio quality of an audio signal. The grades are in the range of 1 to 5, where for each grade a different label is given. A higher audio quality score can indicate a higher audio quality. The highest audio quality score "5" can be labeled as "imperceptible", indicating that the effect of the noise on the audio signal is not perceptible. The audio quality score "4" can be labeled as "perceptible but not annoying", indicating that the effect of the noise on the audio signal is perceptible but not annoying. The audio quality score "3" can be labeled as "slightly annoying", indicating that the effect of the noise on the audio signal is slightly annoying. The audio quality score "2" can be labeled as "annoying", indicating that the effect of the noise on the audio signal is annoying. Further, the audio quality score "1" can be labeled as "very annoying", indicating that the effect of the noise on the audio signal is very annoying. The audio quality score "3" of the impairment quality label "slightly annoying" can be a threshold score.
[0078] The DSP circuit 126 can be additionally configured to receive, from the host processor 124, a first delay value FD that can indicate an expected delay between the plurality of first audio blocks FBl-FBN and the plurality of second audio blocks SBl-SBN, prior to analyzing the plurality of first audio quality scores (shown in Figure 3B Figure 3B The DSP circuit 126 can be additionally configured to receive, from the host processor 124, a first delay value FD that can indicate an expected delay between the plurality of first audio blocks FBl-FBN and the plurality of second audio blocks SBl-SBN, prior to analyzing the plurality of first audio quality scores (shown in Figure 3B Figure 3B The DSP circuit 126 can be additionally configured to receive, from the host processor 124, a first delay value FD that can indicate an expected delay between the plurality of first audio blocks FBl-FBN and the plurality of second audio blocks SBl-SBN, prior to analyzing the plurality of first audio quality scores (shown in
[0079]
[0080] where 'x[k]' is an audio block of the plurality of first audio blocks FB1-FBN and 'h[k]' is an audio block of the plurality of second audio blocks SB1-SBN, denotes an operator for convolution (e.g., circular convolution), and N corresponds to a predefined number of audio blocks of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN needed to determine the second delay value SD. The convolution can involve multiplying the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN shifted by index 'k' a number of times, and summing each result of the multiplication. The'mod' operation can ensure that the first index is the same as the last index, thereby making the convolution circular.
[0081] The DSP circuit 126 can be further configured to generate a ready signal RS based on the second delay value SD. The DSP circuit 126 can be further configured to provide the ready signal RS to the host processor 124. The DSP circuit 126 can be further configured to receive a confirmation signal CS from the host processor 124 based on the ready signal RS. The host processor 124 can generate the confirmation signal CS to confirm that the request to initiate the analysis of the plurality of first audio quality scores (shown in Figure 3B ) and the plurality of second audio quality scores (shown in Figure 3B ) is accepted.
[0082] The DSP circuit 126 can be further configured to receive the plurality of first power levels FV1-FVN and the plurality of second power levels SV1-SVN from the first receiver 116 and the second receiver 118, respectively. The DSP circuit 126 can be further configured to determine whether a first power level FV1 of the plurality of first power levels FV1-FVN and a corresponding second power level SV1 of the plurality of second power levels SV1-SVN is greater than a threshold power level prior to the analysis of the plurality of first audio quality scores (shown in Figure 3B ) and the plurality of second audio quality scores (shown in Figure 3B ). The threshold power level can be a minimum power level that indicates an acceptable intensity of the first audio signal FS and the second audio signal SS, respectively. In an example scenario, the system 104 can receive the first audio signal FS and the second audio signal SS for a predefined period of time (e.g., 10 seconds). For a first predefined period of time (e.g., four seconds), the DSP circuit 126 can determine whether the plurality of first power levels FV1-FVN and the plurality of second power levels SV1-SVN can be greater than the threshold power level. When the plurality of first power levels FV1-FVN and the plurality of second power levels SV1-SVN are determined to be greater than the threshold power level, the DSP circuit 126 can proceed with the analysis of the plurality of first audio quality scores (shown in Figure 3B ) and the plurality of second audio quality scores (shown in Figure 3BAt a second time point of the predefined time period (e.g., at the fifth second), the DSP circuit 126 can determine that at least one of the plurality of first power levels FV1-FVN is below the threshold power level. When at least one of the plurality of first power levels FV1-FVN is below the threshold power level at the second time point, the first receiver 116 is unable to receive the first audio signal FS at the second time point. Accordingly, the first receiver 116 can fail to provide a digitized version of the first audio signal FS (e.g., the plurality of first audio packets FP1-FPN) to the first buffer 120 at the second time point, resulting in an empty state of the first buffer 120 at the second time point. The DSP circuit 126 can output a corresponding second audio block of the plurality of second audio blocks SBl-SBN as the corresponding mix block at the second time point. For simplicity of the ongoing description, assume that each of the first power levels FV1-FVN and the plurality of second power levels SV1-SVN can be greater than or equal to the threshold power level.
[0083] The DSP circuit 126 can additionally be configured to analyze the plurality of first audio quality scores (shown in Figure 3B and the plurality of second audio quality scores (shown in Figure 2 ) by identifying whether each audio quality score of the plurality of first audio quality scores (shown in Figure 2 and each corresponding audio quality score of the plurality of second audio quality scores (shown in Figure 2 ) is greater than a threshold score. The threshold score can indicate an acceptable quality score of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SBl-SBN. The DSP circuit 126 can additionally be configured to tune a first audio block FB1 of the plurality of first audio blocks FB1-FBN and a corresponding second audio block SB1 of the plurality of second audio blocks SBl-SBN based on detecting that at least one of the first audio block FB1 and the corresponding second audio block SB1 are out of sync due to the second delay value SD. The DSP circuit 126 can detect that the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SBl-SBN can be out of sync during determining the second delay value SD. The DSP circuit 126 can tune the audio blocks by performing a second cross-correlation method. The DSP circuit 126 can perform the second cross-correlation method in accordance with the second delay value SD to tune the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SBl-SBN. The DSP circuit 126 can tune the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SBl-SBN based on the plurality of first audio quality scores (shown in Figure 4 and the plurality of second audio quality scores (shown in Figure 2The tuning of the first audio block FB1 and the corresponding second audio block SB1 can be performed based on the analysis of the plurality of first audio quality scores (shown in Figure 3B The mixing of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN is explained in detail in Figure 3B and 3B The plurality of mixed blocks B1-BN can be outputted by the DSP circuit 126 after analyzing the plurality of first audio quality scores (shown in Figure 3A and 3B The plurality of mixed blocks B1-BN can be outputted by the DSP circuit 126 after analyzing the plurality of first audio quality scores (shown in Figure 3B The DSP circuit 126 is explained in detail in
[0084] Figure 3B The exemplary scenario 200 showing the mixing of the plurality of mixed blocks B1-BN by the DSP circuit 126 according to an embodiment of the present disclosure is collectively represented.
[0085] The plurality of first audio blocks FB1-FBN is shown to include first to seventh audio blocks FB1-FB7. In an example, the seventh audio block FB7 is after the sixth audio block FB6. The plurality of first audio quality scores (shown in Figure 3B may include first to seventh audio quality scores FA1-FA7 of the first to seventh audio blocks FB1-FB7, respectively. The first to seventh audio quality scores FA1-FA7 can be '2', '5', '4', '2', '2', '1', and '1', respectively. The plurality of second audio blocks SB1-SBN is shown to include eighth to fourteenth audio blocks SB1-SB7. In an example, the fourteenth audio block SB7 is after the thirteenth audio block SB6. Further, each of the plurality of first audio blocks FB1-FBN corresponds to one of the plurality of second audio blocks SB1-SBN. In an example, the first audio block FB1 corresponds to the eighth audio block SB1. The plurality of second audio quality scores (shown in Figure 3BThe eighth through fourteenth audio quality scores SA1-SA7 can include eighth through fourteenth audio quality scores SA1-SA7 of eighth through fourteenth audio blocks SB1-SB7. The eighth through fourteenth audio quality scores SA1-SA7 are '1', '5', '5', '3', '4', '4', and '5', respectively. The DSP circuit 126 can identify whether each of the first through seventh audio quality scores FA1-FA7 and each corresponding one of the eighth through fourteenth audio quality scores SA1-SA7 is greater than a threshold score. In an example, to output a first mixing block B1 of the plurality of mixing blocks B1-BN, the DSP circuit 126 can identify whether the first audio quality score FA1 and the eighth audio quality score SA1 are greater than a threshold score. For example, the threshold score can be '3'. The first audio quality score FA1 '2' and the eighth audio quality score SA1 '1' are lower than the threshold score '3'. However, the first audio quality score FA1 '2' is greater than the eighth audio quality score SA1 '1'. Accordingly, the first audio block FB1 can be output by the DSP circuit 126 as the first mixing block B1. Similarly, to output a second mixing block B2 of the plurality of mixing blocks B1-BN, the DSP circuit 126 can identify whether the second audio quality score FA2 and the ninth audio quality score SA2 are greater than a threshold score. The second audio quality score FA2 '5' and the ninth audio quality score SA2 '5' are equal and greater than the threshold score '3'. In such a scenario, upon identifying that at least one of the plurality of first audio quality scores (e.g., the second audio quality score FA2) and at least one of the plurality of second audio quality scores (e.g., the tenth audio quality score SA2) is greater than the threshold score, the DSP circuit 126 can detect a previous mixing block of the plurality of mixing blocks B1-BN to output one of the plurality of mixing blocks B1-BN. The previous mixing block (e.g., the first mixing block B1) is detected to be the first audio block FB1. Upon detecting that the first mixing block B1 is the previous mixing block, a second audio block FB2 can be output as a current mixing block (e.g., the second mixing block B2) of the plurality of mixing blocks B1-BN.
[0086] To output a third mixing block B3 of the plurality of mixing blocks B1-BN, the DSP circuit 126 can identify whether the third audio quality score FA3 and the tenth audio quality score SA3 are greater than a threshold score. Upon identifying that the third audio quality score FA3 '4' and the tenth audio quality score SA3 '5' are greater than the threshold score '3' and that the tenth audio quality score SA3 '5' is greater than the third audio quality score FA3 '4', the DSP circuit 126 can detect a previous mixing block of the plurality of mixing blocks B1-BN. The previous mixing block (e.g., the second mixing block B2) is detected to be the second audio block FB2. Upon detecting that the second audio block FB2 is the previous mixing block, a third audio block FB3 can be output as a current mixing block (e.g., the third mixing block B3) of the plurality of mixing blocks B1-BN.
[0087] To output the fourth mixing block B4 of the plurality of mixing blocks B1-BN, the DSP circuit 126 can identify whether the fourth audio quality score FA4 and the eleventh audio quality score SA4 are greater than the threshold score. The fourth audio quality score FA4 '2' is lower than the threshold score '3', and the eleventh audio quality score SA4 '3' is equal to the threshold score '3'. The DSP circuit 126 can detect a previous mixing block of the plurality of mixing blocks B1-BN. The detected previous mixing block is the third audio block FB3. Upon detecting that the third audio block FB3 is the previous mixing block, the fourth audio block FB4 can be output as the current mixing block (e.g., the fourth mixing block B4) of the plurality of mixing blocks B1-BN. Despite the eleventh audio quality score SA4 '3' being equal to the threshold score and greater than the corresponding fourth audio quality score FA4 '2', the DSP circuit 126 can refrain from mixing the fourth audio block FB4 with the eleventh audio block SB4 until the audio quality scores of the number of sub-plurality of audio blocks are lower than the threshold score. Thus, the DSP circuit 126 can mix after the number of audio blocks whose audio quality scores are lower than the threshold score is greater than or equal to the number of such sub-plurality of audio blocks. In an example, the number of sub-plurality of audio blocks is two, but can be a different number (e.g., three, four, etc.). The DSP circuit 126 generally refrains from mixing the number of audio blocks that is lower than the number of sub-plurality of audio blocks to reduce processing overhead that can have occurred due to constant mixing between the audio blocks.
[0088] The DSP circuit 126 can predict that the fifth audio quality score FA5 and the sixth audio quality score FA6 of the subsequent audio blocks (e.g., the fifth audio block FB5 and the sixth audio block FB6) of the plurality of first audio blocks FB1-FBN can be lower than the threshold score because the fourth audio quality score FA4 of the fourth audio block FB4 is lower than the threshold score. In such a scenario, the DSP circuit 126 can additionally predict that mixing of the plurality of first audio blocks FB1-FBN with the plurality of second audio blocks SB1-SBN can be necessary when the audio quality scores of the number of sub-plurality of audio blocks of the plurality of first audio blocks FB1-FBN are lower than the threshold score. It will be apparent to those skilled in the art, however, that the number of sub-plurality of audio blocks is two in the present disclosure, but in various other embodiments, the number of sub-plurality of audio blocks can be greater than two. Additionally, the DSP circuit 126 can determine the number of sub-plurality of audio blocks based on a number of factors, such as the location of the device, data associated with mixing of previous audio signals, etc.
[0089] To output a fifth mixing block B5 of the plurality of mixing blocks B1-BN, the DSP circuit 126 can identify whether the fifth audio quality score FA5 and the twelfth audio quality score SA5 are greater than the threshold score. The fifth audio quality score FA5 '2' is lower than the threshold score '3', and the twelfth audio quality score SA5 '4' is greater than the threshold score '3'. The DSP circuit 126 can detect a previous mixing block of the plurality of mixing blocks B1-BN. Additionally, the DSP circuit 126 can detect whether the number of the sub-plurality of audio blocks having a lower threshold score is less than two. The previous mixing block (e.g., the first mixing block B1) is detected to be the fourth audio block FB4, and the number of audio blocks of the plurality of first audio blocks FB1-FBN having a lower threshold score remains below two. Upon detecting that the fourth mixing block B4 is the previous mixing block and the number of the sub-plurality of audio blocks is one, the fifth audio block FB5 can be output as a current mixing block (e.g., the fifth mixing block B5) of the plurality of mixing blocks B1-BN.
[0090] To output a sixth mixing block B6 of the plurality of mixing blocks B1-BN, the DSP circuit 126 can identify whether the sixth audio quality score FA6 and the thirteenth audio quality score SA6 are greater than the threshold score. The sixth audio quality score FA6 '1'is less than the threshold score '3', and the thirteenth audio quality score SA6 '4' is greater than the threshold score '3'. The DSP circuit 126 can detect a previous mixing block of the plurality of mixing blocks B1-BN. Additionally, the DSP circuit 126 can detect that the number of the sub-plurality of audio blocks is equal to the number two. The DSP circuit 126 can additionally detect whether the sixth audio block FB6 and the thirteenth audio block SB6 are out of sync by the second delay value SD. The DSP circuit 126 can tune one of the sixth audio block FB6 and the thirteenth audio block SB6 based on detecting that one of the sixth audio block FB6 and the thirteenth audio block SB6 is out of sync by the second delay value SD. The DSP circuit 126 can mix the sixth audio block FB6 with the thirteenth audio block SB6 to output a current mixing block (e.g., the sixth mixing block B6) of the plurality of mixing blocks FB1-FBN. The sixth mixing block B6 can be output after tuning one of the sixth audio block FB6 and the thirteenth audio block SB6 by the second delay value SD. In one embodiment, the sixth mixing block B6 can employ an equal composition of both the sixth audio block FB6 and the thirteenth audio block SB6 (e.g., B6 = FB6 + SB6). In another embodiment, the composition of the sixth audio block FB6 in the sixth mixing block B6 can be greater than the composition of the thirteenth audio block SB6 in the sixth audio block B6. The composition of the sixth audio block FB6 and the thirteenth audio block SB6 in the sixth mixing block B6 can be based on a level of smoothness required for playing the audio output signal. The DSP circuit 126 can terminate mixing of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN based on a predefined mixing block value. Thus, the DSP circuit 126 can output one of the plurality of second audio blocks SB1-SBN after mixing a number of audio blocks involved is greater than or equal to the number of the predefined mixing block value. In an example, the predefined mixing block value is one. The DSP circuit 126 can store the predefined mixing block value in a memory (not shown) associated with the DSP circuit 126. The predefined mixing block value can be determined based on a level of smoothness required for playing the audio output signal. It will be apparent to those skilled in the art, however, that the predefined mixing block value is one in the present disclosure, but in various other embodiments, the predefined mixing block value can be greater than one.
[0091] In another example embodiment, the sixth audio quality score FA6 and the thirteenth audio quality score SA6 are both below the threshold score. Thus, although the number of audio blocks in the plurality of first audio blocks FB1-FBN having a lower threshold score is equal to the number of sub-plurality of audio blocks, the DSP circuit 126 can output the sixth audio block FB6 as the current mixed output block (e.g., the sixth mixed output block B6) in the plurality of mixed audio blocks B1-BN based on detecting that the previous mixed output block (e.g., the fifth mixed output block B5) in the plurality of mixed output blocks B1-BN is the fifth audio block FB5 and the thirteenth audio quality score SA6 is below the threshold score.
[0092] To output the seventh mixed block B7 in the plurality of mixed blocks B1-BN, the DSP circuit 126 can identify whether the seventh audio quality score FA7 and the fifteenth audio quality score SA6 are greater than the threshold score. The seventh audio quality score FA7‘1’ is below the threshold score‘3’ and the fourteenth audio quality score SA7‘5’ is greater than the threshold score‘3’. The DSP circuit 126 can detect the previous mixed block in the plurality of mixed blocks B1-BN. Further, the DSP circuit 126 can detect whether the number of sub-plurality of audio blocks involved in mixing the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN is less than the predefined mixed block value. When the DSP circuit 126 detects that the number of sub-plurality of audio blocks is equal to the predefined mixed block value (e.g., the number one), the DSP circuit 126 can output the fourteenth audio block SB7 as the seventh mixed block B7 in the plurality of mixed blocks B1-BN. The mixing of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN to output the sixth mixed block B6 indicates mixing of the first audio signal FS and the second audio signal SS. Additionally, early prediction of mixing at the fourth mixed block B4 results in relatively smoother audio played by the device as compared to conventional solutions.
[0093] Figure 3B A block diagram 300A showing training of the first machine learning model 302, in accordance with example embodiments of the present disclosure.
[0094] The DSP circuit 126 can be additionally configured to train the first machine learning model 302 based on the training data to obtain a trained machine learning model (shown at Figure 4The training data can include a plurality of test audio recordings 304 (e.g., simulated FM data) and a plurality of quality scores 306 (e.g., audio quality labels of the simulated FM data), such that the plurality of quality scores 306 can include a first quality score of a first test audio recording of the plurality of test audio recordings 304. The DSP circuit 126 can be additionally configured to receive the training data from an external memory (not shown). Each of the plurality of quality scores 306 can indicate an audio quality of a corresponding audio recording, where a low score can indicate a low quality of a test audio recording of the plurality of test audio recordings 304, and a high score can indicate a high quality of the test audio recording. The plurality of quality scores 306 can be generated by a tester or an audio analysis model. The audio analysis model can be used in several commercial software (e.g., Opticom Opera) to enable generation of the plurality of quality scores 306. Each of the plurality of quality scores 306 can be a perceptual evaluation of audio quality (PEAQ) score. The PEAQ score can range from 1 to 5, where 1 represents the lowest quality, and 5 represents the highest quality. To generate the plurality of quality scores 306, the audio analysis model can compare a reference signal (e.g., an audio signal that remains unaffected by any external interference) with a signal under test (e.g., an audio signal affected by external interference). Each of the plurality of quality scores 306 can be generated based on an ITU five-grade impairment scale or some other scoring paradigm.
[0095] To train the first machine learning model 302, the DSP circuit 126 can be additionally configured to extract a first plurality of training parameters (e.g., perceptual audio quality descriptors) of the first test audio recording.
[0096] The first plurality of training parameters can be equivalent to the first plurality of audio parameters and the second plurality of audio parameters. The DSP circuit 126 can be further configured to determine, based on the processing of the first plurality of training parameters, a first test score of the plurality of test scores 308 by means of a test scoring operation. Each of the plurality of test scores 308 can be indicative of an audio quality (e.g., a perceptual audio quality) of a corresponding test audio recording that can be determined by the first machine learning model 302. The test scoring operation can determine the first test score based on a mathematical equation. The mathematical equation can be based on weights assigned to each of the plurality of training parameters. The DSP circuit 126 can be further configured to compare the first test score to the first quality score to determine a match between the first test score and the first quality score. The match between the first test score and the first quality score can indicate to the first machine learning model 302 that the determination of the first test score by means of the test scoring operation is accurate. A mismatch between the first test score and the first quality score can indicate to the first machine learning model 302 that the determination of the first test score by means of the test scoring operation is inaccurate. The DSP circuit 126 can be further configured to update the test scoring operation until the match between the first test score and the first quality score is determined. The test scoring operation can be updated by updating the weights of each of the plurality of training parameters of the mathematical equation. The first machine learning model 302 can be trained based on the match between the first test score and the first quality score. The trained machine learning model (shown in Memory 404: ) can produce the first plurality of audio quality scores (shown in DSP processor 402: ) and the second plurality of audio quality scores (shown in Synchronizer 406: ) based on the training of the first machine learning model 302. An example of the first machine learning model 302 can be a linear regression model, a logistic regression model, etc.
[0097] In an embodiment, the DSP circuit 126 can continuously predict a dependent output variable (e.g., each of the plurality of test scores 308) based on one or more independent variables (e.g., the plurality of training parameters). Additionally, the DSP circuit 126 can perform a linear regression model as the first machine learning model 302 when a relationship between each of the plurality of test scores 308 and the plurality of training parameters is linear. In another embodiment, the DSP circuit 126 can perform a logistic regression model as the first machine learning model 302 to predict whether an instance of a dependent variable (e.g., the first test score of the plurality of test scores 308) belongs to a given class (e.g., a range of PEAQ scores).
[0098] Controller 408:A block diagram 300B representing an implementation of a trained machine learning model 310 showing after training the first machine learning model 302 according to an example embodiment of the disclosure. The trained machine learning model 310 can extract a plurality of first audio parameters associated with the plurality of first audio blocks FB1-FBN and a plurality of second audio parameters associated with the plurality of second audio blocks SB1-SBN. The trained machine learning model 310 can process the plurality of first audio parameters and the plurality of second audio parameters to generate a plurality of first audio quality scores FA1-FAN and a plurality of second audio quality scores SA1-SAN, respectively, based on the plurality of first audio parameters and the plurality of second audio parameters.
[0099] Figure 2 A block diagram of the DSP circuit 126 according to an example embodiment of the disclosure. The DSP circuit 126 can include a DSP processor 402, a memory 404, a synchronizer 406, a controller 408, and a mixer 410.
[0100] Mixer 410:
[0101] The memory 404 can include suitable logic, circuitry, and / or interfaces to store data. For example, the memory 404 can be configured to store the trained machine learning model 310. Examples of the memory 404 can include a random access memory (RAM), a read only memory (ROM), a removable memory drive, a hard disk drive (HDD), a flash memory, a solid state memory, and the like.
[0102] Figure 2
[0103] The DSP processor 402 can be configured to train the first machine learning model 302 to obtain the trained machine learning model 310. The DSP processor 402 can read the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN from the first buffer 120 and the second buffer 122, respectively.
[0104] The DSP processor 402 can be additionally configured to perform a time-to-frequency domain conversion operation on the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN to generate a plurality of third audio blocks and a plurality of fourth audio blocks, respectively. In an example, the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN are in a time domain, and the plurality of third audio blocks and the plurality of fourth audio blocks are in a frequency domain. The DSP processor 402 can be additionally configured to execute the trained machine learning model 310 to extract a plurality of first audio parameters and a plurality of second audio parameters based on the plurality of third audio blocks and the plurality of fourth audio blocks, respectively. The DSP processor 402 can be additionally configured to generate a plurality of first audio quality scores FA1-FAN and a plurality of second audio quality scores SA1-SAN based on processing of the plurality of first audio parameters and the plurality of second audio parameters, respectively, by additionally executing the trained machine learning model 310. The trained machine learning model 310 can generate the plurality of first audio quality scores FA1-FAN and the plurality of second audio quality scores SA1-SAN, respectively, with a weighted mathematical equation that is based on a weight of each of the plurality of first audio parameters and the plurality of second audio parameters. An example of the DSP processor 402 can be a CPU, a GPU, a microcontroller, an ASIC, or the like.
[0105] Figure 2
[0106] The synchronizer 406 can be configured to read the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN from the first buffer 120 and the second buffer 122, respectively. The synchronizer 406 can be additionally configured to receive the plurality of first audio quality scores FA1-FAN and the plurality of second audio quality scores SA1-SAN from the DSP processor 402. The synchronizer 406 can be additionally configured to receive the first delay value FD from the host processor 124. The synchronizer 406 can be additionally configured to convolve the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN to determine the second delay value SD. The synchronizer 406 can perform a first cross-correlation method in accordance with the first delay value FD to convolve the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN. Prior to convolving the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN, the synchronizer 406 can be additionally configured to determine a plurality of rejected audio blocks from the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN. The plurality of rejected audio blocks can be determined such that a corresponding audio quality score of each of the plurality of rejected audio blocks is below a threshold score. The synchronizer 406 can be additionally configured to ignore the plurality of rejected audio blocks during determination of the second delay value SD, thereby ensuring that the second delay value SD is accurately determined. The second delay value SD can be determined by an amount of shifting required to determine a similarity between the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN. The synchronizer 406 can be additionally configured to generate and provide a completion signal CL to the controller 408 indicating that the second delay value SD is determined by the synchronizer 406. The synchronizer 406 can be additionally configured to provide the plurality of first audio blocks FB1-FBN, the plurality of second audio blocks SB1-SBN, and the second delay value SD to the mixer 410. In an embodiment, the synchronizer 406 can be additionally configured to tune one audio block from one of the plurality of first audio blocks FB1-FBN and one of the plurality of second audio blocks SB1-SBN by the second delay value SD prior to providing the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN to the mixer 410. The synchronizer 406 can perform a second cross-correlation method in accordance with the second delay value SD to tune the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN.
[0107] Figures 5A-5B
[0108] The controller 408 can be further configured to receive the plurality of first audio quality scores FA1-FA N and the plurality of second audio quality scores SA1-SAN from the DSP processor 402. The controller 408 can be further configured to receive the completion signal CL from the synchronizer 406. The controller 408 can receive the completion signal CL from the synchronizer 406 to indicate that the second delay value SD can be determined by the synchronizer 406. The controller 408 can be further configured to generate the ready signal RS. The ready signal RS can indicate a request to initiate analysis of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN. The controller 408 can be further configured to provide the ready signal RS to the host processor 124. The controller 408 can be further configured to receive the confirmation signal CS from the host processor 124. The confirmation signal CS can indicate initiation of analysis of the plurality of first audio quality scores FA1-FA N and the plurality of second audio quality scores SA1-SAN. The controller 408 can be further configured to receive the plurality of first power levels FV1-FVN and the plurality of second power levels SV1-SVN from the first receiver 116 and the second receiver 118, respectively.
[0109] The controller 408 can be further configured to determine whether the plurality of first power levels FV1-FVN and the plurality of second power levels SV1-SVN are greater than a threshold power level prior to analyzing the plurality of first audio quality scores FA1-FA N and the plurality of second audio quality scores SA1-SAN. For simplicity of the ongoing description, it is assumed that the power levels of the first audio signal FS and the second audio signal SS can be greater than or equal to the threshold power level. The controller 408 can be further configured to determine whether to mix the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN based on the analysis of the plurality of first audio quality scores FA1-FA N and the plurality of second audio quality scores SA1-SAN as explained in Figure 5A The controller 408 can be further configured to determine whether to mix the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN based on the analysis of the plurality of first audio quality scores FA1-FA N and the plurality of second audio quality scores SA1-SAN as explained in
[0110] Figure 5B
[0111] The mixer 410 can additionally be configured to receive the plurality of first audio blocks FB1-FBN, the plurality of second audio blocks SB1-SBN, and the second delay value SD from the synchronizer 406. The mixer 410 can be configured to receive the initiation signal IS from the controller 408. In one scenario, the mixer 410 can output one of the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN as one of the plurality of mixed blocks B1-BN (as explained in Figures 6A-6F Figure 6A In another scenario, the mixer 410 can mix the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN to output one of the plurality of mixed blocks B1-BN based on the initiation signal IS (as explained in Figure 6A In another embodiment, the mixer 410 can mix and output the plurality of mixed audio blocks B1-BN based on receiving the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN tuned from the synchronizer 406.
[0112] Figure 6B FIG. 5 collectively represents a flowchart 500 illustrating training of the first machine learning model 302, in accordance with an embodiment of the present disclosure.
[0113] Referring to Figure 6C At step 502, the DSP circuit 126 can store training data to train the first machine learning model 302, which can include the plurality of test audio recordings 304 and the plurality of quality scores 306. The plurality of test audio recordings 304 can include a first test audio recording, and the plurality of quality scores 306 can include a first quality score to train the first machine learning model 302. At step 504, the DSP circuit 126 can extract a first plurality of training parameters of the first test audio recording by executing the first machine learning model 302. At step 506, the DSP circuit 126 can determine a first test score based on processing of the plurality of training parameters by means of a test scoring operation.
[0114] Referring to Figure 6DAt step 508, the DSP circuit 126 can compare the first test score with the first quality score. At step 510, the DSP circuit 126 can determine whether the first test score matches the first quality score. If it is determined that the first test score matches the first quality score, step 512 is performed. At step 512, the DSP circuit 126 can train the first machine learning model 302 based on the match between the first test score and the first quality score to obtain a trained machine learning model 310. If it is determined that the first test score does not match the first quality score, step 514 is performed. At step 514, the DSP circuit 126 can update the first machine learning model 302 based on the mismatch between the first test score and the first quality score, and step 508 is performed. Steps 508, 510, and 514 are performed until a match between the first test score and the first quality score can be determined. Upon determining the match, step 512 is performed.
[0115] Figure 6E Collectively, a flowchart 600 showing a method of audio mixing by the DSP circuit 126 according to an embodiment of the present disclosure is shown.
[0116] Referring to Figure 6F At step 602, the DSP circuit 126 can read a plurality of first audio blocks FB1-FBN associated with the first audio signal FS and a plurality of second audio blocks SB1-SBN associated with the second audio signal SS from the first buffer 120 and the second buffer 122, respectively. At step 604, the DSP circuit 126 can perform a time-to-frequency domain conversion operation on the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN to generate a plurality of third audio blocks and a plurality of fourth audio blocks, respectively. At step 606, the DSP circuit 126 can extract a plurality of first audio parameters and a plurality of second audio parameters by performing the trained machine learning model 310 on the plurality of third audio blocks and the plurality of fourth audio blocks, respectively.
[0117] Referring to At step 608, the DSP circuit 126 can process the plurality of first audio parameters and the plurality of second audio parameters by additionally executing the trained machine learning model 310 to generate a plurality of first audio quality scores FA1-FA N and a plurality of second audio quality scores SA1-SAN, respectively. At step 610, the DSP circuit 126 can receive the first delay value FD from the host processor 124 based on the generation of the plurality of first audio quality scores FA1-FA N and the plurality of second audio quality scores SA1-SAN. At step 612, the DSP circuit 126 can convolve the plurality of first audio blocks FB1-FBN and the plurality of second audio blocks SB1-SBN based on the first delay value FD, the plurality of first audio quality scores FA1-FA N and the plurality of second audio quality scores SA1-SAN to determine a second delay value SD. At step 614, the DSP circuit 126 can be configured to generate a ready signal RS based on the determination of the second delay value SD. At step 616, the DSP circuit 126 can provide the ready signal RS to the host processor 124.
[0118] Referring now to At step 617, the DSP circuit 126 can receive a confirmation signal CS from the host processor 124 based on the ready signal RS. At step 618, the DSP circuit 126 can receive a plurality of first power levels FV1-FVN and a plurality of second power levels SV1-SVN from the first receiver 116 and the second receiver 118, respectively, based on the confirmation signal CS. At step 619, the DSP circuit 126 can determine whether each power level of the plurality of first power levels FV1-FVN and a corresponding power level of the plurality of second power levels SV1-SVN is greater than a threshold power level based on receiving the plurality of first power levels FV1-FVN and the plurality of second power levels SV1-SVN. For simplicity of the ongoing description, it is assumed that the power levels of the first audio signal FS and the second audio signal SS can be greater than or equal to the threshold power level. At step 620, the DSP circuit 126 can identify whether each audio quality score of the plurality of first audio quality scores FA1-FA N and each corresponding audio quality score of the plurality of second audio quality scores SA1-SAN is greater than a threshold score.
[0119] Referring now to Steps 620a-620c and steps 620d-620i are various exemplary scenarios illustrating the steps of outputting an audio block. At step 620a, the DSP circuit 126 can identify that the audio quality score of one of the plurality of first audio blocks FB1-FBN and the audio quality score of the corresponding audio block of the plurality of second audio blocks SB1-SBN are greater than the threshold score. At step 620b, the DSP circuit 126 can detect a previous mixed block. At step 620c, when the previous mixed block is one of (i) a previous first audio block of the plurality of first audio blocks FB1-FBN and (ii) a previous second audio block of the plurality of second audio blocks SB1-SBN, the DSP circuit 126 can output the audio block of (i) one of the plurality of first audio blocks FB1-FBN and (ii) one of the plurality of second audio blocks SB1-SBN, respectively, as one of the plurality of mixed blocks B1-BN. Step 622 is performed after step 620c. In another exemplary scenario, at step 620d, the DSP circuit 126 can identify that the audio quality score of one of the plurality of first audio blocks FB1-FBN is lower than the threshold score and the audio quality score of the corresponding audio block of the plurality of second audio blocks SB1-SBN is greater than the threshold score.
[0120] Referring now to At step 620e, the DSP circuit 126 can determine whether one of the plurality of first audio blocks FB1-FBN and the corresponding audio block of the plurality of second audio blocks SB1-SBN are out of sync by the second delay value SD. If it is determined that one of the plurality of first audio blocks FB1-FBN and the corresponding audio block of the plurality of second audio blocks SB1-SBN are out of sync by the second delay value SD, step 620f is performed. At step 620f, the DSP circuit 126 can tune one audio block from one of the plurality of first audio blocks FB1-FBN and the corresponding audio block of the plurality of second audio blocks SB1-SBN by the second delay value SD. Step 620g is performed after step 620f. If it is determined that one audio block from one of the plurality of first audio blocks FB1-FBN and the corresponding audio block of the plurality of second audio blocks SB1-SBN can be synchronized, step 620g is performed. At step 620g, the DSP circuit 126 can detect a previous sub-plurality of mixed blocks. At step 620h, the DSP circuit 126 can mix one of the plurality of first audio blocks FB1-FBN and the corresponding audio block of the plurality of second audio blocks SB1-SBN based on the previous sub-plurality of mixed blocks to output one of the plurality of mixed blocks B1-BN. Step 622 is performed after step 620h.
[0121] Referring now to At step 622, it is determined whether all the plurality of mixed blocks B1-BN is output. If it is determined that all the plurality of mixed blocks B1-BN is remaining to be output, one of steps 620a and 620d is performed. If it is determined that all the plurality of mixed blocks B1-BN is output, the process is paused.
[0122] Compared to conventional solutions that rely only on the power level of the audio signal, the DSP circuit 126 of the present disclosure outputs the plurality of mixed blocks B1-BN based on the audio quality of each of the audio blocks by means of the plurality of first audio quality scores FA1-FAN and the plurality of second audio quality scores SA1-SAN and the power level of the audio signal. Thus, the quality of the output audio signal, e.g. the plurality of mixed blocks B1-BN, is improved compared to conventional solutions that output the audio signal based only on the power level. Due to the analysis of the plurality of first audio quality scores FA1-FAN and the plurality of second audio quality scores SA1-SAN, the trained machine learning model 310 can enable the DSP circuit 126 to predict the mixing early, thereby enabling a seamless smooth transition of the mixed first audio signal FS and second audio signal SS. Thus, the listening experience of the user is improved. Furthermore, the DSP circuit 126 can ignore the plurality of rejected audio blocks during determining the second delay value SD to ensure that the second delay value SD is determined accurately. The DSP circuit 126 can perform the mixing of the first audio signal FS and the second audio signal SS based on the number of the previous sub-plurality of mixed blocks and the sub-plurality of audio blocks. Thus, the processing cycles for outputting each of the plurality of mixed blocks B1-BN is reduced and the plurality of mixed blocks B1-BN is output at a faster rate than conventional solutions.
[0123] While various embodiments of the present disclosure have been shown and described, it is to be understood that the present disclosure is not limited to the embodiments. Numerous modifications, changes, variations, alternatives, and equivalents will occur to those skilled in the art without departing from the spirit and scope of the present disclosure as described in the claims. In addition, terms such as "first" and "second" are used to arbitrarily distinguish one element from another element, unless otherwise specified. Therefore, these terms do not necessarily indicate priority or other priority of the elements.
Claims
1. A system for mixing audio signals, characterized by The system comprises: a digital signal processing (DSP) circuit configured to: extract, by executing a trained machine learning model, a plurality of first audio parameters associated with a plurality of first audio blocks of a first audio signal and a plurality of second audio parameters associated with a plurality of second audio blocks of a second audio signal, wherein the system is configured to receive the first audio signal and the second audio signal; process, by additionally executing the trained machine learning model, the plurality of first audio parameters and the plurality of second audio parameters to generate a plurality of first audio quality scores and a plurality of second audio quality scores, respectively, wherein each audio quality score of the plurality of first audio quality scores and the plurality of second audio quality scores indicates an audio quality of a corresponding audio block of the plurality of first audio blocks and the plurality of second audio blocks, respectively; analyze the plurality of first audio quality scores and the plurality of second audio quality scores; and output, upon analyzing the plurality of first audio quality scores and the plurality of second audio quality scores, a plurality of mixed blocks, wherein each of the plurality of mixed blocks is at least one of a first audio block of the plurality of first audio blocks and a second audio block of the plurality of second audio blocks.
2. The system of claim 1, wherein, The DSP circuit is additionally configured to train a first machine learning model based on training data to obtain the trained machine learning model, wherein the training data comprises a plurality of test audio recordings and a plurality of quality scores, such that the plurality of quality scores comprises a first quality score of a first test audio recording of the plurality of test audio recordings, wherein each of the plurality of quality scores indicates an audio quality of a corresponding test audio recording, and wherein a low score indicates a low quality of a test audio recording of the plurality of test audio recordings and a high score indicates a high quality of the test audio recording.
3. The system of claim 2, wherein, To train the first machine learning model, the DSP circuit is additionally configured to: extract a first plurality of training parameters of the first test audio recording; determine a first test score based on a processing of the first plurality of training parameters by means of a test scoring operation; compare the first test score with the first quality score to determine a match between the first test score and the first quality score, wherein the match between the first test score and the first quality score indicates to the first machine learning model that the determination of the first test score by means of the test scoring operation is accurate, and a mismatch between the first test score and the first quality score indicates to the first machine learning model that the determination of the first test score by means of the test scoring operation is erroneous; and updating the test scoring operation until a match between the first test score and the first quality score is determined, wherein the first machine learning model is trained based on the match between the first test score and the first quality score, and wherein the trained machine learning model produces the plurality of first audio quality scores and the plurality of second audio quality scores based on the training of the first machine learning model.
4. The system of claim 1, wherein, The DSP circuit is additionally configured to perform time-to-frequency domain conversion operations on the plurality of first audio blocks and the plurality of second audio blocks to produce a plurality of third audio blocks and a plurality of fourth audio blocks, respectively, and wherein the plurality of third audio blocks and the plurality of fourth audio blocks in the frequency domain are provided to the trained machine learning model to extract the plurality of first audio parameters and the plurality of second audio parameters, respectively.
5. The system of claim 1, wherein, Additionally comprising a first receiver and a second receiver configured to: receive the first audio signal and the second audio signal from an audio source; and convert each of the first audio signal and the second audio signal into a digitized version of each of the first audio signal and the second audio signal, respectively.
6. The system of claim 5, wherein, Additionally comprising a first buffer and a second buffer coupled to the first receiver and the second receiver, respectively, wherein the first buffer and the second buffer are configured to: receive the digitized version of each of the first audio signal and the second audio signal from the first receiver and the second receiver, respectively; and store the digitized version of each of the first audio signal and the second audio signal as the plurality of first audio blocks and the plurality of second audio blocks, respectively.
7. The system of claim 6, wherein, The DSP circuit is additionally configured to read the plurality of first audio blocks and the plurality of second audio blocks from the first buffer and the second buffer, respectively, wherein the plurality of first audio parameters and the plurality of second audio parameters are extracted after reading the plurality of first audio blocks and the plurality of second audio blocks, respectively.
8. The system of claim 1, wherein, Additionally comprising a host processor, wherein the DSP circuit is additionally configured to: receive, from the host processor, a first delay value indicative of an expected delay between the plurality of first audio blocks and the plurality of second audio blocks prior to analyzing the plurality of first audio quality scores and the plurality of second audio quality scores; and convolve the plurality of first audio blocks and the plurality of second audio blocks based on a group consisting of the first delay value, the plurality of first audio quality scores, and the plurality of second audio quality scores to determine a second delay value indicative of an actual delay between the plurality of first audio blocks and the plurality of second audio blocks.
9. The system of claim 1, wherein, To analyze the plurality of first audio quality scores and the plurality of second audio quality scores, the DSP circuit is further configured to identify whether each audio quality score in the plurality of first audio quality scores and each corresponding audio quality score in the plurality of second audio quality scores is greater than a threshold score, wherein (i) the plurality of first audio blocks comprises a first audio block and a second audio block such that the second audio block is after the first audio block, (ii) the plurality of second audio blocks comprises a third audio block such that the third audio block corresponds to the second audio block, (iii) the plurality of first audio quality scores comprises a first audio quality score for the first audio block and a second audio quality score for the second audio block, and (iv) the plurality of second audio quality scores comprises a third audio quality score for the third audio block, and wherein upon identifying that at least one audio quality score in the plurality of first audio quality scores and the plurality of second audio quality scores is greater than the threshold score, the DSP circuit is further configured to detect at least one previous mixing block in the plurality of mixing blocks to output one of the plurality of mixing blocks.
10. A method characterized by, Comprising: extracting, by a digital signal processing (DSP) circuit, a plurality of first audio parameters associated with a plurality of first audio blocks of a first audio signal and a plurality of second audio parameters associated with a plurality of second audio blocks of a second audio signal by executing a trained machine learning model; processing, by the DSP circuit, the plurality of first audio parameters and the plurality of second audio parameters to generate a plurality of first audio quality scores and a plurality of second audio quality scores, respectively, by further executing the trained machine learning model, wherein each audio quality score in the plurality of first audio quality scores and the plurality of second audio quality scores indicates an audio quality of a corresponding audio block in the plurality of first audio blocks and the plurality of second audio blocks, respectively; analyzing, by the DSP circuit, the plurality of first audio quality scores and the plurality of second audio quality scores; and outputting, by the DSP circuit, a plurality of mixing blocks upon analyzing the plurality of first audio quality scores and the plurality of second audio quality scores, wherein each of the plurality of mixing blocks is at least one of a first audio block in the plurality of first audio blocks and a second audio block in the plurality of second audio blocks.