Signal transmission methods, apparatus, electronic devices and computer-readable storage media
By dynamically updating the interval frame count and threshold value, the transmission of the encoded frame signal is adjusted according to the noise environment of the receiver, which solves the problem of low comfort noise quality in DTX transmission and improves user experience and bandwidth utilization efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-02
- Publication Date
- 2026-04-03
AI Technical Summary
The existing fixed interval frame count and threshold value result in low comfort noise quality during DTX voice transmission, which cannot be flexibly adjusted and affects the user's listening experience.
By acquiring the ambient noise energy of the receiver, dynamically updating the interval frame number and threshold value, and adjusting the transmission of the encoded frame signal according to the changes in target noise, more adaptive and comfortable noise is generated.
It improves the quality of comfortable noise, enhances the user's auditory experience, and reduces the bandwidth usage when noise levels change.
Smart Images

Figure CN116229993B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of signal processing technology, and more specifically to a signal transmission method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] With the development of science and technology, voice is being used more and more widely. Voice transmission is one technology used in voice applications.
[0003] Currently, to reduce bandwidth consumption, Discontinuous Transmission (DTX) technology is used for voice transmission. In DTX voice transmission, when no voice frame signal is detected, encoded non-voice frame signals are transmitted using a fixed interval of frames and a threshold value. However, the fixed interval of frames and threshold value cannot be flexibly changed, resulting in low-quality comfort noise generated from the encoded non-voice frame signals. Summary of the Invention
[0004] This application provides a signal transmission method, apparatus, electronic device, and computer-readable storage medium that can solve the technical problem of low quality comfort noise caused by fixed interval frame number and threshold value.
[0005] A signal transmission method, comprising:
[0006] Acquire the target noise level in the recipient's environment;
[0007] Determine the energy of the aforementioned target noise;
[0008] Based on the energy of the target noise, the current interval frame number and the current threshold value are updated respectively to obtain the target interval frame number and the target threshold value;
[0009] If the energy of the frame signal to be transmitted is less than the target threshold value, the frame signal to be transmitted is determined as the target non-speech frame signal, and the target non-speech frame signal is encoded to obtain the target encoded frame signal.
[0010] If the number of frames between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, then the target encoded frame signal will be transmitted to the receiving party.
[0011] Accordingly, embodiments of this application provide a signal transmission device, including:
[0012] The acquisition module is used to acquire the target noise in the recipient's environment;
[0013] The determination module is used to determine the energy of the aforementioned target noise;
[0014] The update module is used to update the current interval frame number and the current threshold value based on the energy of the target noise mentioned above, so as to obtain the target interval frame number and the target threshold value.
[0015] The encoding module is used to determine the frame signal to be transmitted as a target non-speech frame signal if the energy of the frame signal to be transmitted is less than the target threshold value, and to encode the target non-speech frame signal to obtain the target encoded frame signal.
[0016] The transmission module is configured to transmit the target encoded frame signal to the receiving party if the number of frames between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number.
[0017] Optionally, the energy of the target noise includes the target power spectrum of the target noise;
[0018] Accordingly, the above update module is specifically used to perform:
[0019] The initial interval frame number is determined based on the target power spectrum described above;
[0020] Get the maximum and minimum interval frame counts;
[0021] The target interval frame number is determined based on the above maximum interval frame number, the above minimum interval frame number, and the above initial interval frame number.
[0022] Optionally, the above update module is specifically used to perform:
[0023] The first interval frame number is determined based on the above maximum interval frame number and the above initial interval frame number;
[0024] The target interval frame number is determined based on the minimum interval frame number and the first interval frame number mentioned above.
[0025] Optionally, the above update module is specifically used to perform:
[0026] The smaller of the maximum interval frame number and the initial interval frame number is taken as the first interval frame number;
[0027] The larger of the minimum interval frame number and the first interval frame number is taken as the target interval frame number.
[0028] Optionally, the above update module is specifically used to perform:
[0029] The target power spectrum is substituted into the first preset function for calculation to obtain the initial interval frame number, which increases as the target power spectrum increases.
[0030] Optionally, the above update module is specifically used to perform:
[0031] Obtain the first mapping table;
[0032] The initial interval frame number corresponding to the target power spectrum is found in the first mapping table above. The initial interval frame number increases as the target power spectrum increases.
[0033] Optionally, the above update module is specifically used to perform:
[0034] The target power spectrum is substituted into the second preset function for calculation to obtain the target threshold value, which increases as the target power spectrum increases.
[0035] Furthermore, this application also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the signal transmission method provided in this application.
[0036] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a computer program adapted for loading by a processor to execute any of the signal transmission methods provided in embodiments of this application.
[0037] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the signal transmission methods provided in this application.
[0038] In this embodiment, the target noise of the receiver's environment is first acquired. Then, the energy of the target noise is determined. Next, based on the energy of the target noise, the current interval frame number and the current threshold value are updated to obtain the target interval frame number and the target threshold value. Then, if the energy of the frame signal to be transmitted is less than the target threshold value, the frame signal to be transmitted is determined as a target non-speech frame signal, and the target non-speech frame signal is encoded to obtain the target encoded frame signal. Finally, if the interval frame number between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, the target encoded frame signal is transmitted to the receiver.
[0039] In this embodiment, the current interval frame number and the current threshold value are not fixed. The sender updates the current interval frame number and the current threshold value based on the energy of the target noise of the receiver to obtain the target interval frame number and the target threshold value. Finally, when the energy of the frame signal to be transmitted is less than the target threshold value, the target non-speech frame signal is encoded to obtain the target encoded frame signal. If the interval frame number between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, the target encoded frame signal is transmitted to the receiver. This makes the target interval frame number and the target threshold value change with the change of the target noise, so that the comfort noise generated based on the target encoded frame signal changes with the change of the target noise, thereby improving the quality of the comfort noise and thus improving the listening experience of the receiver. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a schematic diagram of a signal transmission process provided in an embodiment of this application;
[0042] Figure 2 This is a schematic flowchart of the signal transmission method provided in an embodiment of this application;
[0043] Figure 3 This is a schematic diagram illustrating the determination of the minimum power spectrum provided in an embodiment of this application;
[0044] Figure 4 This is a schematic flowchart of another signal transmission method provided in an embodiment of this application;
[0045] Figure 5 This is an interactive schematic diagram of the signal transmission method provided in the embodiments of this application;
[0046] Figure 6 This is a schematic diagram of the structure of the signal transmission device provided in the embodiments of this application;
[0047] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0048] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0049] This application provides a signal transmission method, apparatus, electronic device, and computer-readable storage medium. The signal transmission apparatus can be integrated into an electronic device, which may be a server, a terminal, or other similar device.
[0050] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.
[0051] Furthermore, multiple servers can form a blockchain, with the servers being nodes on the blockchain.
[0052] The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, voice-interactive device, smart home appliance, or in-vehicle terminal, but is not limited to these. The terminal and server can be connected directly or indirectly via wired or wireless communication.
[0053] For example, such as Figure 1 As shown, the terminal can acquire the target noise of the recipient's environment; determine the energy of the target noise; based on the energy of the target noise, update the current interval frame number and the current threshold value respectively to obtain the target interval frame number and the target threshold value; if the energy of the frame signal to be transmitted is less than the target threshold value, the frame signal to be transmitted is determined as the target non-speech frame signal, and the target non-speech frame signal is encoded to obtain the target encoded frame signal; if the interval frame number between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, the target encoded frame signal is transmitted to the recipient.
[0054] Furthermore, in the embodiments of this application, "multiple" refers to two or more. The terms "first" and "second," etc., in the embodiments of this application are used for distinguishing descriptions and should not be construed as implying relative importance.
[0055] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.
[0056] In DTX voice transmission, a threshold value is used to determine whether the signal is a voice frame. If a non-voice frame signal is detected, the sender encodes the non-voice frame signal and sends the encoded non-voice frame signal to the receiver to notify the receiver's decoder to enter the noise decoding stage. Then, the sender sends one encoded non-voice frame signal to the receiver after a two-frame interval. This process continues after a seven-frame interval until the sender detects a voice frame signal and stops sending encoded non-voice frames. In other words, the frequency at which the sender sends encoded non-voice frames is fixed at this point, sending one frame every seven frames.
[0057] The fixed interval frame number and threshold value cannot be flexibly changed, resulting in low quality of comfort noise generated from the encoded non-speech frame signal.
[0058] To address the technical problem of low-quality comfort noise generated from non-speech frame signals, this application provides a signal transmission method. In this method, firstly, the target noise of the receiver's environment is acquired. Next, the energy of the target noise is determined. Then, based on the energy of the target noise, the current interval frame number and the current threshold value are updated to obtain the target interval frame number and the target threshold value. Next, if the energy of the frame signal to be transmitted is less than the target threshold value, the frame signal to be transmitted is identified as a target non-speech frame signal, and the target non-speech frame signal is encoded to obtain a target encoded frame signal. Finally, if the interval frame number between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, the target encoded frame signal is transmitted to the receiver.
[0059] In this embodiment of the application, during the DTX transmission of voice, the current interval frame number and the current threshold value are not fixed. The sender updates the current interval frame number and the current threshold value based on the energy of the target noise of the receiver to obtain the target interval frame number and the target threshold value. Finally, when the energy of the frame signal to be transmitted is less than the target threshold value, the target non-voice frame signal is encoded to obtain the target encoded frame signal. If the interval frame number between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, the target encoded frame signal is transmitted to the receiver. This makes the target interval frame number and the target threshold value change with the change of the target noise, so that the comfort noise generated based on the target encoded frame signal changes with the change of the target noise, thereby improving the quality of the comfort noise and thus improving the listening experience of the receiver.
[0060] In this embodiment, the description will be from the perspective of a signal transmission device, which can be integrated into a server or terminal. To facilitate the explanation of the signal transmission method of this application, the following will describe the process in detail with the signal transmission device integrated into the terminal, that is, with the terminal as the execution subject and the terminal as the sound sender.
[0061] Please see Figure 2 , Figure 2 This is a schematic flowchart of a signal transmission method provided in an embodiment of this application. The signal transmission method may include:
[0062] S201. Obtain the target noise of the receiver's environment.
[0063] The receiver can acquire the target noise of the receiver's environment at preset time intervals, and then send the target noise to the terminal, so that the terminal can acquire the target noise.
[0064] Alternatively, the receiving party can continuously acquire the target noise of the receiving party's environment and then send the target noise to the terminal, which then acquires the target noise.
[0065] Alternatively, the receiving party can continuously acquire the target noise in its environment. The receiving party then detects this target noise; if it changes, it transmits the target noise to the terminal, allowing the terminal to acquire it. If the target noise remains unchanged, it is not transmitted to the terminal. When the target noise changes, it is then transmitted to the terminal, thus reducing the amount of target noise transmitted to the terminal and consequently reducing the occupied transmission bandwidth.
[0066] Alternatively, the receiving party can obtain the target sound of their environment, then send the target sound to the terminal, and the terminal can then extract the target noise from the target sound, thus obtaining the target noise.
[0067] Users can choose the method for acquiring target noise based on their actual situation; this application does not impose any restrictions on it.
[0068] Alternatively, the receiving party may begin sending the target noise to the terminal after receiving the encoded frame signal. Or, the receiving party may begin sending the target noise to the terminal upon receiving the encoded voice frame signal; this application does not impose any limitations on this.
[0069] S202. Determine the energy of the target noise.
[0070] After acquiring the target noise, the terminal determines the energy of the target noise. The energy of the target noise may include the target power spectrum of the target noise or the energy of the target noise in the time domain (the energy in the time domain refers to the integral of the square of the amplitude of the target noise).
[0071] It should be noted that if the target noise includes multiple frames of noise signals, then the energy of the target noise refers to the energy of one frame of noise signal in the target noise.
[0072] When the energy of the target noise includes the target power spectrum of the target noise, the method for determining the energy of the target noise can be selected according to the actual situation. For example, the likelihood ratio method or the Minima-Controlled Recursive Averaging (MCRA) algorithm can be used to calculate the target power spectrum of the target noise.
[0073] The following describes the process of calculating the target power spectrum of the target noise using the minimum value controlled recursive averaging algorithm.
[0074] Perform a Fourier transform on the received target sound to obtain the initial power spectrum of the sub-band signal (one frame of the target sound signal can be divided into a preset number of sub-band signals):
[0075]
[0076] Where i represents the frame number of the sub-band signal, z represents the frequency index (the frequency index represents the number corresponding to each frequency interval after dividing the signal into a preset number of frequency intervals), and k represents the sub-band number. X(i,z) represents the complex frequency domain value obtained by performing a Fourier transform on the z-th frequency point of the i-th frame signal, freq1(k) is the starting frequency index of the k-th sub-band, and freq2(k) is the ending frequency index of the k-th sub-band. S(i,k) represents the initial power spectrum of the sub-band signal.
[0077] Next, the initial power spectrum of the sub-band signal is smoothed in the frequency domain by adjacent sub-band signals and in the time domain by historical frame signals to obtain the historical power spectrum. The frequency domain smoothing process of adjacent sub-band signals is as follows:
[0078]
[0079] This represents the frequency domain smoothing weighting factor group, for example, x[5] = [0.1, 0.2, 0.4, 0.2, 0.1].
[0080] The process of time-domain smoothing of historical frame signals is as follows:
[0081]
[0082] in, Let c0 represent the historical power spectrum, and c0 represent the time-domain smoothing factor, for example, c0 = 0.9.
[0083] Next, the minimum power spectrum of the subband signal is solved using the minimum tracking method. The specific process is as follows: Figure 3 As shown, the specific process is as follows:
[0084] If the remainder between the current frame i to which the sub-band signal belongs and the power spectrum update period T is zero, then the intermediate power spectrum (S) of the previous frame of the current frame will be used. tmp (i-1,k)) and historical power spectrum The smaller value in the spectrum is taken as the minimum power spectrum (S). min (i,k)), and uses the historical power spectrum as the intermediate power spectrum (S) of the current frame. tmp (i,k)), that is, the intermediate power spectrum of the current frame is the historical power spectrum of the current frame.
[0085] If the remainder between the current frame to which the subband signal belongs and the power spectrum update period is not zero, then the smaller value between the intermediate power spectrum of the previous frame and the historical power spectrum is taken as the minimum power spectrum, and the smaller value between the intermediate power spectrum of the previous frame and the historical power spectrum is taken as the intermediate power spectrum of the current frame.
[0086] After obtaining the minimum power spectrum, the target probability of the speech frame signal in the sub-band signal is calculated. The specific process is as follows:
[0087]
[0088]
[0089] p'(i,k)=α p p'(i-1,k)+(1-α p p(i,k)
[0090] Among them, S r (i,k) represents the ratio of the historical power spectrum to the minimum power spectrum, p(i,k) represents the intermediate probability, p'(i,k) represents the target probability, and α p δ represents the smoothing constant and the preset threshold value.
[0091] Finally, the target power spectrum of the target noise is calculated:
[0092] S'(i,k)=p'(i,k)S'(i-1,k)+(1-p'(i,k))S(i,k)
[0093]
[0094] Where S'(i,k) represents the power spectrum of the noise signal in the sub-band signal, β represents the noise estimation smoothing coefficient, β<1, K0 and K1 represent the sub-band signal index, and N(i) represents the power spectrum of the noise in the frame signal to which the sub-band signal belongs, i.e. the power spectrum of the target noise, i.e. the power spectrum of the target noise is composed of the power spectra of the noise in each sub-band signal in the frame signal.
[0095] It should be noted that noise signals can be calculated only in sub-band signals within the frequency range of human voice. For example, to calculate the power spectrum of noise signals in sub-band signals with frequencies below 100 Hz to 6 kHz, the frequency range of the sub-band signals represented by K0 and K1 is 100 Hz to 6 kHz.
[0096] S203. Based on the energy of the target noise, update the current interval frame number and the current threshold value respectively to obtain the target interval frame number and the target threshold value.
[0097] After obtaining the energy of the target noise, the terminal can update the current interval frame count and the current threshold value based on the energy of the target noise, thereby obtaining the target interval frame count and the target threshold value. The current threshold value can be a preset threshold value, and the current interval frame count can be a preset interval frame count.
[0098] The process of updating the current interval frame number and the current threshold value based on the energy of the target noise to obtain the target interval frame number and the target threshold value can be described as follows:
[0099] Based on the energy of the target noise, determine the target interval frame number and the target threshold value, and replace the current interval frame number with the target interval frame number and the current threshold value with the target threshold value.
[0100] Alternatively, based on the energy of the target noise, the process of updating the current interval frame number and the current threshold value to obtain the target interval frame number and the target threshold value can also be as follows:
[0101] Based on the energy of the target noise, the target interval frame number and the target threshold value are determined. The target interval frame number is subtracted from the current interval frame number to obtain the first difference. Then, the first difference is added to the current interval frame number to replace the current interval frame number with the target interval frame number.
[0102] Similarly, the target threshold value is subtracted from the current threshold value to obtain the second difference value. Then, the second difference value is added to the current threshold value to replace the current threshold value with the target threshold value.
[0103] Because the receiver is affected by ambient noise, a masking effect occurs—the perception of a weaker sound is masked by a stronger one. Therefore, lower-energy signals in the sound are masked by the target noise of the receiver, and the user cannot perceive these lower-energy signals.
[0104] Therefore, when the ambient noise level is high, a fixed current threshold and current interval frame number may cause low-energy signals that the receiver cannot hear to be transmitted to the receiver, thus wasting transmission bandwidth. When the ambient noise level is low, a fixed current threshold and current interval frame number may cause the user to perceive discontinuity in the sound.
[0105] Therefore, in some embodiments, the current frame interval and the current threshold value are updated based on the energy of the target noise in the recipient's environment to obtain the target frame interval and the target threshold value. This ensures that the target frame interval and the target threshold value change with the energy of the target noise; that is, the higher the energy of the target noise, the higher the target frame interval and the target threshold value, and vice versa. When the energy of the target noise is lower and the target frame interval and the target threshold value are smaller, the recipient can perceive a continuity of sound, thus improving their auditory experience.
[0106] When the energy of the target noise is greater, and the target interval frame number and target threshold value are also larger, the target interval frame number is greater than the current interval frame number, and the target threshold value is greater than the current threshold value. This reduces the number of encoded frame signals transmitted to the receiver, thereby reducing the occupied transmission bandwidth. Therefore, in this embodiment, the occupied transmission bandwidth can be reduced while the user perceives the continuity of sound.
[0107] In some embodiments, when the energy of the target noise includes the target power spectrum of the target noise, updating the current interval frame number based on the energy of the target noise to obtain the target interval frame number includes:
[0108] The initial interval frame number is determined based on the target power spectrum;
[0109] Get the maximum and minimum interval frame counts;
[0110] The target interval frame number is determined based on the maximum interval frame number, the minimum interval frame number, and the initial interval frame number.
[0111] The process of determining the initial interval frame number based on the target power spectrum can be as follows:
[0112] The target power spectrum is substituted into the first preset function for calculation to obtain the initial interval frame number. Since the initial interval frame number increases with the increase of the target power spectrum, the first preset function can be a monotonically increasing function.
[0113] Alternatively, the mapping relationship between the target power spectrum and the initial interval frame number can be stored in the first mapping table. After the terminal obtains the target power spectrum, the terminal then looks up the initial interval frame number corresponding to the target power spectrum in the first mapping table. The initial interval frame number increases as the target power spectrum increases.
[0114] The maximum and minimum interval frame counts are preset interval frame counts. In this embodiment, the maximum and minimum interval frame counts are preset, and then the target interval frame count is determined based on the maximum, minimum, and initial interval frame counts, instead of directly using the initial interval frame count as the target interval frame count.
[0115] Furthermore, the initial interval frame number increases as the target power spectrum increases, which in turn increases the target interval frame number, thereby reducing the number of encoded frame signals transmitted to the receiver and thus reducing the occupied transmission bandwidth.
[0116] In other embodiments, determining the target interval frame number based on the maximum interval frame number, the minimum interval frame number, and the initial interval frame number includes:
[0117] The first interval frame number is determined based on the maximum interval frame number and the initial interval frame number;
[0118] The target interval frame number is determined based on the minimum interval frame number and the first interval frame number.
[0119] In this embodiment, the first interval frame number is first determined based on the maximum interval frame number and the initial interval frame number, and then the target interval frame number is determined based on the minimum interval frame number and the first interval frame number.
[0120] The process of determining the first interval frame number based on the maximum interval frame number and the initial interval frame number can be as follows: take the smaller value between the maximum interval frame number and the initial interval frame number as the first interval frame number.
[0121] The process of determining the target interval frame number based on the minimum interval frame number and the first interval frame number can be as follows: take the larger value between the minimum interval frame number and the first interval frame number as the target interval frame number.
[0122] That is, at this time, the target interval frame number d = max(DMIN, min(DMAX, f1(N(i)))), where DMIN represents the minimum interval frame number, DMAX represents the maximum interval frame number, f1(N(i)) represents the initial interval frame number, f1() represents the first preset function, and min(DMAX, f1(N(i))) represents the first interval frame number.
[0123] In this embodiment, the smaller of the maximum interval frame number and the initial interval frame number is used as the first interval frame number, and the larger of the minimum interval frame number and the first interval frame number is used as the target interval frame number, thereby preventing the target interval frame number from becoming infinitely large or infinitely small.
[0124] In other embodiments, when the energy of the target noise includes the target power spectrum of the target noise, the current threshold value is updated based on the energy of the target noise to obtain a target threshold value, including:
[0125] The target power spectrum is substituted into the second preset function for calculation to obtain the target threshold value, which increases as the target power spectrum increases.
[0126] Since the target threshold value increases as the target power spectrum increases, the second preset function can be a monotonically increasing function.
[0127] Alternatively, the mapping relationship between the target power spectrum and the target threshold value can be stored in a second mapping table. After the terminal obtains the target power spectrum, the terminal then looks up the target threshold value corresponding to the target power spectrum from the second mapping table.
[0128] The target threshold increases as the target power spectrum increases, which reduces the number of encoded frame signals transmitted to the receiver, thereby reducing the transmission bandwidth used.
[0129] The specific types of the first preset function and the second preset function may be the same or different, and this application embodiment does not limit them.
[0130] S204. If the energy of the frame signal to be transmitted is less than the target threshold, the frame signal to be transmitted is determined as the target non-speech frame signal, and the target non-speech frame signal is encoded to obtain the target encoded frame signal.
[0131] The terminal can determine whether the frame signal to be transmitted is a target non-speech frame signal using the Voice Activity Detection (VAD) algorithm. If the energy of the frame signal to be transmitted is less than the current threshold value in the VAD algorithm, the frame signal to be transmitted is determined to be a target non-speech frame signal; if the energy of the frame signal to be transmitted is greater than or equal to the current threshold value in the VAD algorithm, the frame signal to be transmitted is determined to be a speech frame signal.
[0132] If the frame signal to be transmitted is a voice frame signal, a voice coding algorithm is used to encode the frame signal to be transmitted, thereby obtaining the encoded voice frame signal. The encoded voice frame signal is then transmitted to the receiving party, and the transmission of the encoded frame signal is stopped.
[0133] It should be noted that in the process of using the speech endpoint detection algorithm to determine whether the frame signal to be transmitted is the target non-speech frame signal, other parameters can also be used to determine whether the frame signal to be transmitted is the target non-speech frame signal. For example, the zero-crossing rate and spectral distribution characteristics of the frame signal to be transmitted can be used to determine whether the frame signal to be transmitted is the target non-speech frame signal. That is, the zero-crossing rate, spectral distribution characteristics and energy of the frame signal to be transmitted can be used to determine whether the frame signal to be transmitted is the target non-speech frame signal. In this case, the frame signal to be transmitted is determined to be the target non-speech frame signal only when the zero-crossing rate and spectral distribution characteristics of the frame signal to be transmitted meet the preset conditions and the energy of the frame signal to be transmitted is less than the target threshold.
[0134] Once the terminal determines that the frame signal to be transmitted is a target non-speech frame signal, it can use the noise coding algorithm in DTX to encode the target non-speech frame signal, thereby obtaining the target encoded frame signal. The target encoded frame signal includes the spectral envelope and energy of the frame signal to be transmitted.
[0135] S205. If the number of frames between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, then the target encoded frame signal is transmitted to the receiver.
[0136] The previously transmitted encoded post-frame signal refers to the last encoded post-frame signal transmitted before the target encoded post-frame signal.
[0137] Each time the terminal acquires a frame signal to be transmitted, if the frame signal to be transmitted is a non-voice frame signal, it encodes the frame signal to be transmitted to obtain an encoded frame signal. However, the encoded frame signal is only transmitted to the receiver when the interval between two encoded frame signals reaches the current interval.
[0138] For example, if the current frame interval is 5 frames, when the terminal receives the first encoded frame signal, it transmits the first encoded frame signal to the recipient. When the terminal receives the second to sixth encoded frame signals, it does not transmit the second to sixth encoded frame signals to the recipient. The seventh encoded frame signal is transmitted to the recipient only when the terminal receives the seventh encoded frame signal. Similarly, when the terminal receives the thirteenth encoded frame signal, it transmits the thirteenth encoded frame signal to the recipient.
[0139] If, after transmitting the encoded seventh frame signal to the recipient, the terminal obtains the target interval frame number and the target threshold value, and the target interval frame number is 6 frames, then the terminal will next determine whether the frame signal to be transmitted is a target non-voice frame signal based on the target threshold value. If the frame signal to be transmitted is a target non-voice frame signal, then the target non-voice frame signal is encoded to obtain the target encoded frame signal.
[0140] At this point, the previously transmitted encoded frame signal is the seventh encoded frame signal. Therefore, when the target encoded frame signal is one of the frames from the eighth to the thirteenth encoded frame signals, the target encoded frame signal will not be transmitted to the receiver. The target encoded frame signal will only be transmitted to the receiver when it is the fourteenth encoded frame signal.
[0141] Similarly, if the terminal obtains the target interval frame number and target threshold value after transmitting the thirteenth frame encoded signal to the recipient, and the previously transmitted encoded frame signal is the thirteenth frame encoded signal, then the target encoded frame signal will only be transmitted to the recipient when the target encoded frame signal is the twentieth frame encoded signal.
[0142] It should be noted that the above is just an example of transmitting the target encoded frame signal to the receiver. If the terminal continues to collect frame signals to be transmitted that are non-voice frame signals of the target, the target encoded frame signal can continue to be transmitted based on the target interval frame number.
[0143] Additionally, after transmitting the target encoded frame signal to the recipient, if no voice frame signal is received, the terminal can return to acquiring the target noise of the recipient's environment. If the terminal does not reacquire the target noise of the recipient's environment, it will then determine whether to transmit the acquired target non-voice frame signal to the recipient based on the target threshold value and the target interval frame number.
[0144] If the terminal reacquires the target noise of the recipient's environment, it updates the current interval frame number and the current threshold value based on the energy of the reacquired target noise to obtain a new target interval frame number and a new target threshold value (at this time, the current interval frame number is the target interval frame number, and the current threshold value is the target threshold value). Based on the new target threshold value and the new target interval frame number, it determines whether to transmit the acquired target non-voice frame signal to the recipient.
[0145] After receiving the target encoded frame signal, the receiver uses the Comfort Noise Generation (CNG) algorithm to decode the target encoded frame signal to obtain the target non-speech frame signal.
[0146] Furthermore, when the target interval is 8 frames and the duration of one frame is 20 milliseconds, the receiver updates the spectral envelope and energy of the non-voice frame signal every 160 milliseconds. Within 160 milliseconds, 8 non-voice frame signals with the same characteristics but not exactly the same are generated based on the same spectral envelope and energy as well as random values.
[0147] As described above, in this embodiment, the target noise of the receiver's environment is first obtained. Then, the energy of the target noise is determined. Next, based on the energy of the target noise, the current interval frame number and the current threshold value are updated to obtain the target interval frame number and the target threshold value. Then, if the energy of the frame signal to be transmitted is less than the target threshold value, the frame signal to be transmitted is determined as a target non-speech frame signal, and the target non-speech frame signal is encoded to obtain the target encoded frame signal. Finally, if the interval frame number between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, the target encoded frame signal is transmitted to the receiver.
[0148] In this embodiment, the current interval frame number and the current threshold value are not fixed. The sender updates the current interval frame number and the current threshold value based on the energy of the target noise of the receiver to obtain the target interval frame number and the target threshold value. Finally, when the energy of the frame signal to be transmitted is less than the target threshold value, the target non-speech frame signal is encoded to obtain the target encoded frame signal. If the interval frame number between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, the target encoded frame signal is transmitted to the receiver. This makes the target interval frame number and the target threshold value change with the change of the target noise, so that the comfort noise generated based on the target encoded frame signal changes with the change of the target noise, thereby improving the quality of the comfort noise and thus improving the listening experience of the receiver.
[0149] Based on the methods described in the above embodiments, the following examples will provide further detailed explanations.
[0150] This embodiment uses the signal transmission device integrated into the terminal as an example. Please refer to [link to example]. Figure 4 , Figure 4 This is a schematic flowchart illustrating a signal transmission method provided in an embodiment of this application. The signal transmission method may include:
[0151] S401. The sender determines the power spectrum, zero-crossing rate, and spectral distribution characteristics of the first frame signal to be transmitted in the sound. When the zero-crossing rate and spectral distribution characteristics of the first frame signal to be transmitted meet the preset conditions and the power spectrum is less than the current threshold value, the first frame signal to be transmitted is determined as a non-speech frame signal.
[0152] S402. Encode the first frame signal to be transmitted to obtain the first encoded frame signal, and transmit the first encoded frame signal to the receiver.
[0153] The initiator can use the Voice Activity Detection (VAD) algorithm to determine whether the first frame signal to be transmitted is a non-voice frame signal. In this case, the current threshold value is the energy threshold value in the VAD algorithm. The current threshold value and the current interval frame number can be preset.
[0154] For example, such as Figure 5 As shown, the sender obtains the first frame signal to be transmitted in the sound, uses the voice endpoint detection algorithm to determine whether the first frame signal to be transmitted is a non-voice frame signal. If the first frame signal to be transmitted is a non-voice frame signal, the noise coding algorithm in DTX is used to encode the first frame signal to be transmitted to obtain the first coded frame signal, and the first coded frame signal is transmitted to the receiver through the channel.
[0155] If the first frame signal to be transmitted is a voice frame signal, then the voice coding algorithm is used to encode the first frame signal to be transmitted, and the encoded first frame signal to be transmitted is sent to the receiver through the channel.
[0156] S403, The sender receives the target power spectrum of the target noise returned by the receiver, where the target noise is the ambient noise of the receiver.
[0157] like Figure 5 As shown, after receiving the first encoded frame signal, the receiving party uses a comfortable noise generation algorithm to decode the first encoded frame signal and synthesize comfortable noise, so that the user of the receiving party can hear the noise corresponding to the first frame signal to be transmitted.
[0158] The receiver then acquires the ambient noise, i.e., the target noise. The receiver then determines the target power spectrum of the target noise and returns it to the sender.
[0159] S404. The sender substitutes the target power spectrum into the second preset function for calculation to obtain the target threshold value. The target threshold value increases as the target power spectrum increases.
[0160] S405. The sender substitutes the target power spectrum into the first preset function for calculation to obtain the initial interval frame number. The initial interval frame number increases as the target power spectrum increases.
[0161] S406. The sender uses the smaller of the maximum interval frame number and the initial interval frame number as the first interval frame number, and uses the larger of the minimum interval frame number and the first interval frame number as the target interval frame number.
[0162] The maximum and minimum interval frame counts are preset interval frame counts. In this embodiment, the smaller of the maximum and initial interval frame counts is used as the first interval frame count, and the larger of the minimum and first interval frame counts is used as the target interval frame count, thereby preventing the target interval frame count from becoming infinitely large or infinitely small.
[0163] S407. The sender determines the power spectrum, zero-crossing rate, and spectral distribution characteristics of the second frame signal to be transmitted in the sound. When the zero-crossing rate and spectral distribution characteristics of the second frame signal to be transmitted meet the preset conditions and the power spectrum is less than the target threshold, the second frame signal to be transmitted is determined as a non-speech frame signal.
[0164] S408. Encode the second frame signal to be transmitted to obtain the second encoded frame signal. If the number of frames between the second encoded frame signal and the first encoded frame signal reaches the target number of frames, then transmit the second encoded frame signal to the receiver and return to execute S403.
[0165] For example, if the target interval is 6 frames, the first frame signal to be transmitted is the first non-speech frame signal in the sound, and the second frame signal to be transmitted is the eighth non-speech frame signal in the sound, that is, the interval between the first coded frame signal and the second coded frame signal is 6 frames, then the second coded frame signal will be transmitted to the receiver.
[0166] If the sender transmits a third encoded frame signal before receiving the second frame signal, then if the interval between the second and first encoded frame signals reaches the target interval, the second encoded frame signal is transmitted to the receiver, and execution returns to step S403, including:
[0167] If the number of frames between the second and third encoded frame signals reaches the target number of frames, the second encoded frame signal is transmitted to the receiver, and the process returns to execute S403.
[0168] For example, if the target interval is 6 frames, the first frame to be transmitted is the first non-speech frame in the audio, the third frame to be transmitted is the eighth non-speech frame in the audio, and the second frame to be transmitted is the fifteenth non-speech frame in the audio. After receiving the third frame to be transmitted, the sender will encode it to obtain the third encoded frame signal and transmit it to the receiver.
[0169] The number of frames between the second and third encoded frame signals reaches the target interval of 6, so the second encoded frame signal is transmitted to the receiver.
[0170] If the number of frames between the second coded frame signal and the first coded frame signal does not reach the target number of frames, the second coded frame signal will not be transmitted to the receiver.
[0171] For example, if the target interval is 4 frames, and the second frame to be transmitted is the fourth non-speech frame in the audio, then the sender will not transmit the second encoded frame to the receiver. Subsequent frames to be transmitted are the third and fourth frames. The third frame is the fifth non-speech frame in the audio, and the fourth frame is the sixth non-speech frame in the audio. In this case, the sender will transmit the fourth encoded frame corresponding to the fourth frame to the receiver.
[0172] If a voice frame signal appears in the subsequent frame signal to be transmitted, the voice frame is encoded and the encoded voice frame signal is sent to the receiver, and the execution based on the target interval frame number is stopped, and the second encoded frame signal is sent to the receiver.
[0173] Since the target noise in the environment of the receiver changes, the receiver can periodically acquire the target noise, determine the target power spectrum of the target noise, and return the target power to the sender. This allows the sender to update the target threshold and the target interval frame number based on the newly received target power spectrum, thereby making the target threshold and the target interval frame number change with the change of the target power spectrum.
[0174] The larger the target power spectrum of the target noise, the larger the target threshold and the larger the target interval frame number. The smaller the target power spectrum of the target noise, the smaller the target threshold and the smaller the target interval frame number, thus enabling the receiving user to perceive the continuity of the sound.
[0175] Furthermore, when the target power spectrum of the target noise is larger, the target threshold value is larger, and the target interval frame number is larger, the number of non-voice frame signals transmitted to the receiver is smaller, and the transmission bandwidth occupied is smaller, thereby reducing the occupied transmission bandwidth. Therefore, in the embodiments of this application, the occupied transmission bandwidth can be reduced while the user perceives the continuity of sound.
[0176] In this embodiment, the current interval frame number and the current threshold value are not fixed. The sender updates the current interval frame number and the current threshold value based on the energy of the target noise of the receiver to obtain the target interval frame number and the target threshold value. Finally, when it is determined based on the target threshold value that the received second frame signal to be transmitted is a non-voice frame signal, and the interval frame number between the second frame signal to be transmitted and the first frame signal to be transmitted reaches the target interval frame number, the second frame signal to be transmitted is sent to the receiver. This makes the target interval frame number and the target threshold value change with the change of the target noise, so that the comfort noise generated based on the target non-voice frame signal changes with the change of the target noise, thereby improving the quality of the comfort noise and thus improving the listening experience of the receiver.
[0177] Other specific implementation methods and corresponding beneficial effects in this embodiment can be referred to the content of the above signal transmission method embodiment, and this embodiment does not make specific limitations here.
[0178] To facilitate better implementation of the signal transmission method provided in the embodiments of this application, the embodiments of this application also provide an apparatus based on the above-described signal transmission method. The meanings of the terms used are the same as in the above-described signal transmission method, and specific implementation details can be found in the descriptions in the method embodiments.
[0179] For example, such as Figure 6 As shown, the signal transmission device may include:
[0180] The acquisition module 601 is used to acquire the target noise of the receiver's environment.
[0181] The determination module 602 is used to determine the energy of the target noise.
[0182] The update module 603 is used to update the current interval frame number and the current threshold value based on the energy of the target noise, respectively, to obtain the target interval frame number and the target threshold value.
[0183] The encoding module 604 is used to determine the frame signal to be transmitted as a target non-speech frame signal if the energy of the frame signal to be transmitted is less than the target threshold value, and to encode the target non-speech frame signal to obtain the target encoded frame signal.
[0184] The transmission module 605 is used to transmit the target encoded frame signal to the receiver if the number of frames between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number.
[0185] Optionally, the energy of the target noise includes the target power spectrum of the target noise.
[0186] Accordingly, update module 603 is specifically used to perform:
[0187] The initial interval frame number is determined based on the target power spectrum;
[0188] Get the maximum and minimum interval frame counts;
[0189] The target interval frame number is determined based on the maximum interval frame number, the minimum interval frame number, and the initial interval frame number.
[0190] Optionally, update module 603 is specifically used to perform:
[0191] The first interval frame number is determined based on the maximum interval frame number and the initial interval frame number;
[0192] The target interval frame number is determined based on the minimum interval frame number and the first interval frame number.
[0193] Optionally, update module 603 is specifically used to perform:
[0194] The smaller of the maximum interval frame number and the initial interval frame number is used as the first interval frame number;
[0195] The larger of the minimum interval frame number and the first interval frame number is used as the target interval frame number.
[0196] Optionally, update module 603 is specifically used to perform:
[0197] The target power spectrum is substituted into the first preset function for calculation to obtain the initial interval frame number, which increases as the target power spectrum increases.
[0198] Optionally, update module 603 is specifically used to perform:
[0199] Obtain the first mapping table;
[0200] Find the initial interval frame number corresponding to the target power spectrum from the first mapping table. The initial interval frame number increases as the target power spectrum increases.
[0201] Optionally, update module 603 is specifically used to perform:
[0202] The target power spectrum is substituted into the second preset function for calculation to obtain the target threshold value, which increases as the target power spectrum increases.
[0203] In practice, each of the above modules can be implemented as an independent entity or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation methods and corresponding beneficial effects of each of the above modules, please refer to the previous method embodiments, which will not be repeated here.
[0204] This application also provides an electronic device, which may be a server or a terminal, etc. Figure 7 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically:
[0205] The electronic device may include components such as a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, a power supply 703, and an input unit 704. Those skilled in the art will understand that... Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0206] The processor 701 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing computer programs and / or modules stored in the memory 702, and by calling data stored in the memory 702, thereby providing overall monitoring of the electronic device. Optionally, the processor 701 may include one or more processing cores; preferably, the processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 701.
[0207] The memory 702 can be used to store computer programs and modules. The processor 701 executes various functional applications and data processing by running the computer programs and modules stored in the memory 702. The memory 702 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, computer programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 702 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 702 may also include a memory controller to provide the processor 701 with access to the memory 702.
[0208] The electronic device also includes a power supply 703 that supplies power to the various components. Preferably, the power supply 703 can be logically connected to the processor 701 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 703 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0209] The electronic device may also include an input unit 704, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0210] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 701 in the electronic device loads the executable files corresponding to the processes of one or more computer programs into the memory 702 according to the following instructions, and the processor 701 runs the computer programs stored in the memory 702 to realize various functions, such as:
[0211] Acquire the target noise in the environment of the receiver of the sound;
[0212] Determine the energy of the target noise;
[0213] Based on the energy of the target noise, the current interval frame number and the current threshold value are updated respectively to obtain the target interval frame number and the target threshold value;
[0214] If the energy of the frame signal to be transmitted is less than the target threshold, the frame signal to be transmitted is determined to be the target non-speech frame signal, and the target non-speech frame signal is encoded to obtain the target encoded frame signal.
[0215] If the number of frames between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, then the target encoded frame signal will be transmitted to the receiver.
[0216] For details on the specific implementation methods and corresponding beneficial effects of the above operations, please refer to the detailed description of the image processing methods above, which will not be repeated here.
[0217] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0218] Therefore, embodiments of this application provide a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the signal transmission methods provided in embodiments of this application. For example, the computer program can execute the following steps:
[0219] When the energy of the frame signal to be transmitted is detected to be less than the current threshold, the frame signal to be transmitted is determined as the initial non-voice frame signal and transmitted to the receiver.
[0220] Acquire the target noise in the environment of the receiver of the sound;
[0221] Determine the energy of the target noise;
[0222] Based on the energy of the target noise, the current interval frame number and the current threshold value are updated respectively to obtain the target interval frame number and the target threshold value;
[0223] If the energy of the frame signal to be transmitted is less than the target threshold, the frame signal to be transmitted is determined to be the target non-speech frame signal, and the target non-speech frame signal is encoded to obtain the target encoded frame signal.
[0224] If the number of frames between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, then the target encoded frame signal will be transmitted to the receiver.
[0225] For details on the specific implementation methods and corresponding beneficial effects of the above operations, please refer to the previous embodiments, which will not be repeated here.
[0226] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0227] Since the computer program stored in the computer-readable storage medium can execute the steps of any of the signal transmission methods provided in the embodiments of this application, the beneficial effects that any of the signal transmission methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0228] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned signal transmission method.
[0229] The foregoing has provided a detailed description of a signal transmission method, apparatus, electronic device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A signal transmission method, characterized in that, include: Acquire the target noise level in the recipient's environment; Determine the energy of the target noise; Based on the energy of the target noise, the current interval frame number and the current threshold value are updated respectively to obtain the target interval frame number and the target threshold value; If the energy of the frame signal to be transmitted is less than the target threshold, the frame signal to be transmitted is determined as the target non-speech frame signal, and the target non-speech frame signal is encoded to obtain the target encoded frame signal. If the number of frames between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number, then the target encoded frame signal is transmitted to the receiving party. Wherein, the greater the energy of the target noise, the greater the target interval frame number and the target threshold value; the smaller the energy of the target noise, the smaller the target interval frame number and the target threshold value. When the energy of the target noise is greater, the number of target interval frames and the target threshold value are greater, the number of target interval frames is greater than the current number of interval frames and the target threshold value is greater than the current threshold value, so as to reduce the number of frames of the target encoded frame signal transmitted to the receiver.
2. The signal transmission method according to claim 1, characterized in that, The energy of the target noise includes the target power spectrum of the target noise; The step of updating the current interval frame number based on the energy of the target noise to obtain the target interval frame number includes: The initial interval frame number is determined based on the target power spectrum; Get the maximum and minimum interval frame counts; The target interval frame number is determined based on the maximum interval frame number, the minimum interval frame number, and the initial interval frame number.
3. The signal transmission method according to claim 2, characterized in that, Determining the target interval frame number based on the maximum interval frame number, the minimum interval frame number, and the initial interval frame number includes: The first interval frame number is determined based on the maximum interval frame number and the initial interval frame number; The target interval frame number is determined based on the minimum interval frame number and the first interval frame number.
4. The signal transmission method according to claim 3, characterized in that, Determining the first interval frame number based on the maximum interval frame number and the initial interval frame number includes: The smaller value between the maximum interval frame number and the initial interval frame number is used as the first interval frame number; Accordingly, determining the target interval frame number based on the minimum interval frame number and the first interval frame number includes: The larger of the minimum interval frame number and the first interval frame number is taken as the target interval frame number.
5. The signal transmission method according to claim 2, characterized in that, Determining the initial interval frame number based on the target power spectrum includes: The target power spectrum is substituted into the first preset function for calculation to obtain the initial interval frame number, which increases as the target power spectrum increases.
6. The signal transmission method according to claim 2, characterized in that, Determining the initial interval frame number based on the target power spectrum includes: Obtain the first mapping table; The initial interval frame number corresponding to the target power spectrum is found from the first mapping table. The initial interval frame number increases as the target power spectrum increases.
7. The signal transmission method according to claim 1, characterized in that, The energy of the target noise includes the target power spectrum of the target noise; The step of updating the current threshold value based on the energy of the target noise to obtain the target threshold value includes: The target power spectrum is substituted into the second preset function for calculation to obtain the target threshold value, which increases as the target power spectrum increases.
8. A signal transmission device, characterized in that, include: The acquisition module is used to acquire the target noise in the recipient's environment; The determination module is used to determine the energy of the target noise; The update module is used to update the current interval frame number and the current threshold value based on the energy of the target noise, respectively, to obtain the target interval frame number and the target threshold value; The encoding module is used to determine the frame signal to be transmitted as a target non-speech frame signal if the energy of the frame signal to be transmitted is less than the target threshold value, and to encode the target non-speech frame signal to obtain the target encoded frame signal. The transmission module is configured to transmit the target encoded frame signal to the receiving party if the number of frames between the target encoded frame signal and the previously transmitted encoded frame signal reaches the target interval frame number. Wherein, the greater the energy of the target noise, the greater the target interval frame number and the target threshold value; the smaller the energy of the target noise, the smaller the target interval frame number and the target threshold value. When the energy of the target noise is greater, the number of target interval frames and the target threshold value are greater, the number of target interval frames is greater than the current number of interval frames and the target threshold value is greater than the current threshold value, so as to reduce the number of frames of the target encoded frame signal transmitted to the receiver.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a computer program, and the processor running the computer program in the memory to perform the signal transmission method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the signal transmission method according to any one of claims 1 to 7.
11. A computer program product, characterized in that, The computer program product stores a computer program adapted for loading by a processor to execute the signal transmission method according to any one of claims 1 to 7.
Citation Information
Patent Citations
System and method for adaptive transmission of comfort noise parameters during discontinuous speech transmission
US20060293885A1