A sound compensation method and device
By detecting the waiting time and abnormal time of sound capture during video recording or live broadcast, and performing mute compensation, the system solves the problem of audio and video out-synchronization caused by skipping the sound-free time period, improving user experience and saving time costs.
Patent Information
- Application Number
- CN202211713567.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-12-29
AI Technical Summary
During video recording or live broadcast, the system may skip the soundless time period, causing the audio and video to be out of sync.
By determining the waiting time and abnormal time of sound capture, it is determined whether silent compensation is required. The specific steps include performing sound capture, determining whether sound data is detected, calculating the abnormal time, and performing mute compensation when the abnormal time length is greater than or equal to the threshold.
It effectively avoids the problem of audio and video out-of-synchronization, ensures that the audio sampling time is consistent with the video sampling time, improves the user experience, and saves the time cost of manual re-recording in the later stage.
Smart Images

Figure CN116170632B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of communication equipment, and in particular, relates to a sound compensation method and device. Background Art
[0002] With the rise of short videos, live video and recorded video industries, more and more users are beginning to communicate through recorded or live broadcasts. However, when recording screens and live video, the system does not capture sound throughout the entire process (for example, when the user pauses in their speech or remains silent for a long time). This situation can easily cause the captured audio sampling duration to be less than the video sampling duration, resulting in audio and video being out of sync when the video file is played. Audio and video out of sync, as the name implies, refers to the fact that the played video screen is out of sync with the played sound. Specifically, it is the phenomenon that the video screen has not been updated, but the audio has started playing.
[0003] Figure 1 is an example of the total duration of an audio sample. Figure 1 As shown, the total sound capture time starts from segment A, passes through segment B, and ends at segment C. Among them, segments A and C are the time periods for normal capture of system sounds (for example, the computer is continuously playing songs during this time period), and segment B is the time period when the sound cannot be captured normally, that is, the time period without audio output. Under normal circumstances, the acquired audio information only includes segments A and C. That is, the actual audio file only has the voice package of segment A and the voice package of segment C. When playing audio and video, due to the lack of audio data of segment B, the audio of segment C will be played at the same time as the video screen of segment B is played.
[0004] In the existing technology, operators can only handle the problem after the person who posted the recorded video or watched the live broadcast reported that the audio and video were out of sync. The recorded video can be re-recorded by re-recording the sound data later, or lengthening or shortening the audio duration of segment A or segment C, to align the correct video time point and audio time point. However, manual re-recording is time-consuming and labor-intensive, and is prone to errors. Summary of the invention
[0005] The embodiments of the present application provide a sound compensation method and device, which can solve the problem that the system automatically skips the silent period during video recording or live broadcast, resulting in audio and video asynchronism.
[0006] In a first aspect, an embodiment of the present application provides a sound compensation method, comprising:
[0007] Determine a first waiting time according to a first parameter, a first cache capacity, and a cache duration, wherein the first waiting time refers to a waiting time for capturing sound, the first parameter is used to characterize a sampling rate of the sound, and the first cache capacity is used to characterize a space capacity for caching sound data;
[0008] Execute sound capture and record a first time, wherein the first time refers to a start time of the sound capture;
[0009] After the first waiting period ends, determining whether sound data is detected;
[0010] If no sound data is detected, record a second time and determine a first abnormal duration, wherein the second time refers to the time when no sound data is detected, the first abnormal duration refers to the duration when no sound data is captured, and the first abnormal duration is determined based on the first time and the second time;
[0011] Determine whether the first abnormal duration is greater than or equal to a first time threshold;
[0012] When the first abnormal duration is greater than or equal to a first time threshold, determining a first compensation frame, where the first compensation frame refers to a voice data frame that needs to be silence compensated;
[0013] Based on the first compensation frame, silence compensation is performed on the system sound. In an embodiment of the present application, after the sound output format is obtained, the sound capture time required for sound capture at each stage is calculated in reverse; the process of sound capture and detection of whether the sound data packet is captured is executed in turn in each stage until the entire sound capture time period ends. Among them, after each determination that the sound data packet is not captured, the abnormal capture duration is determined, and on the premise that it is determined to be greater than the first time threshold, audio data compensation is performed for the time period corresponding to the abnormal capture duration. Compared with the prior art, the sound compensation method provided in the embodiment of the present application effectively avoids the problem that during the video recording or live broadcast process, the system skips the time period without audio data due to the existence of a time period in which the audio is not captured during audio sampling, and the final audio sampling duration cannot be synchronized with the video screen, that is, the problem of audio and video being out of sync is caused, and the captured audio data and video screen data are accurately matched, which greatly improves the user experience, and does not require manual re-recording in the later stage, saving a lot of time costs.
[0014] In a possible implementation, the method also includes: the method also includes: obtaining a sound output format, the sound output format includes: the first parameter, the second parameter, and the third parameter, wherein the second parameter is used to characterize the number of sampling channels of the sound, and the third parameter is used to characterize the number of sampling bytes.
[0015] In a possible implementation, the method further includes:
[0016] If a sound data packet is detected after the first waiting period ends, a third time is recorded, where the third time is the time when the sound data packet is detected;
[0017] Updating the first time to the third time;
[0018] According to the first waiting time, starting to perform sound capture from the third time;
[0019] After the first waiting period ends, determining whether sound data is detected;
[0020] If no sound data is detected, a fourth time is recorded and a second abnormal duration is determined, wherein the fourth time refers to a time when no sound data is detected, and the second abnormal duration is determined based on the third time and the fourth time.
[0021] In a possible implementation, after the first waiting period ends, determining whether sound data is detected includes: after the first waiting period ends, calling an interface; and determining whether sound data is detected according to a value returned by the interface.
[0022] In a possible implementation, the performing silence compensation on the system sound based on the first compensation frame includes: after determining the number of the first compensation frames, inserting the first compensation frames into corresponding positions in a sound data queue stored in the system according to time attributes of the first compensation frames;
[0023] Among them, determining the number of first compensation frames includes: determining the unit data amount of the first compensation frame according to the first parameter and the first waiting time, and the unit data amount is used to characterize the number of audio frames required to supplement the audio of the first waiting time in a single sampling channel; determining the total number of the first compensation frames according to the second parameter, the third parameter, and the unit data amount.
[0024] That is to say, because the first compensation frame itself has a time attribute, after receiving the first compensation frame, it will be inserted into the sound data packet queue arranged in order according to the time information carried on it, and then played or stored as a video file in sequence by the system's audio output module. Through the above operation, the audio sampling duration can be made consistent with the video sampling duration, and under the premise that the sound data packet is not captured for a long time, the time period without audio is automatically supplemented into the compensation packet, ensuring that the sound data packet captured later can accurately correspond to the correct video screen time point, and avoiding the problem of audio and video being out of sync.
[0025] In a possible implementation, the unit data amount of the first compensation frame satisfies the following formula: SF = sleeptime* Among them, SF represents the unit data amount of the first compensation frame, sleeptime represents the first waiting time, and S represents the first parameter.
[0026] In a possible implementation, the number of the first compensation frames satisfies the following formula:
[0027] CN=C*NB*SF, wherein CN represents the number of the first compensation frames, C represents the second parameter, NB represents the third parameter, and SF represents the unit data amount of the first compensation frame.
[0028] In a possible implementation manner, the method further includes: the first time threshold is determined according to the first parameter and the unit data amount.
[0029] In a possible implementation manner, the method further includes: the first time threshold satisfies the following formula: Among them, MaxT represents the first time threshold, SF represents the unit data amount of the first compensation frame, and S represents the first parameter.
[0030] That is to say, by presetting the first time threshold, the accuracy of the system in identifying the abnormal time period in which the sound data is not captured is improved.
[0031] In a second aspect, an embodiment of the present application provides a sound compensation device, including: a method for executing the above-mentioned first aspect or any possible implementation of the first aspect. Specifically, the device includes a module (or unit) for executing the above-mentioned first aspect or any possible implementation of the first aspect.
[0032] In a third aspect, an embodiment of the present application provides a sound compensation device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. Specifically, when the processor executes the computer program, the method in the first aspect or any possible implementation of the first aspect is implemented.
[0033] In a fourth aspect, an embodiment of the present application provides an audio system, which can capture and output audio, and perform audio data compensation operations on time periods without audio data during the audio capture process. For example, compensation is performed on the audio-free time period in the method described in the first aspect, and audio is output that precisely matches the time point of the video picture.
[0034] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0035] In a sixth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables a sound compensation device to execute the method in the first aspect or any possible implementation of the first aspect.
[0036] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0037] Compared with the prior art, the embodiments of the present application have the following beneficial effects: the embodiments of the present application aim at the problem of asynchrony between audio and video during video recording or live broadcasting, and propose a sound compensation method. After the sound capture of each stage is completed, it is determined whether a sound data packet is detected, thereby determining the abnormal time period in which the audio is not captured, and supplementing the audio data for the abnormal time period. The problem of the inability to match the later audio sampling time point with the video sampling time point is avoided, the cost and effort of video recording are reduced, the quality of video files is improved, and the user experience is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0039] Figure 1 This is an example graph of the total duration of an audio sample;
[0040] Figure 2 It is a schematic diagram of a scenario for an audio system to capture sound provided by an embodiment of the present application;
[0041] Figure 3 is a flow chart of a sound compensation method provided in an embodiment of the present application;
[0042] Figure 4 is an example flow chart of the sound compensation method provided in the embodiment of the present application;
[0043] Figure 5 is a structural block diagram of a sound compensation device provided in an embodiment of the present application;
[0044] Figure 6 It is a schematic diagram of the structure of the sound compensation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0046] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0047] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0048] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0049] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0050] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0051] The sound compensation method provided in the embodiment of the present application can be applied to terminal devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc. The embodiment of the present application does not impose any restrictions on the specific type of the terminal device.
[0052] In one possible application scenario, the sound compensation method provided in the embodiment of the present application can be applied to application scenarios involving screen recording and desktop video live broadcast, especially involving audio and video recording and audio and video playback.
[0053] In the audio and video recording stage, the user records audio and video, and the system encodes the recorded audio data and video data, and synthesizes the encoded audio data and video data into a video file in a certain format. In the audio and video playback stage, the system separates the video file in a certain format, obtains the encoded audio data and video data, decodes the audio data and video data respectively, and plays the audio and video synchronously.
[0054] Specifically, during the recording stage of audio and video, the sounds of nature are collected through an audio input module (e.g., a microphone) to obtain an analog signal. Then the system captures the analog signal emitted by the audio input module to obtain a digital signal. In this process, the analog signal needs to go through three steps of sampling, quantization, and encoding to be converted into a digital signal. Among them, sampling refers to the collection and recording of continuous analog signals at a certain frequency; quantization refers to the number of binary digits used to represent the collected data; encoding refers to recording the sampled and quantized data in a certain format (e.g., sequential storage or compressed storage). In short, the process of audio collection is the process of converting continuous analog signals into discrete digital signals.
[0055] From the perspective of the audio system, during screen recording or live broadcast, the sound capture process of the audio system can be roughly divided into the following four steps:
[0056] (1) Initialize the buffer inside the audio system according to the configured parameters;
[0057] (2) Collecting original audio data;
[0058] (3) Create a management project and continuously read the collected audio data from the buffer;
[0059] (4) Stop acquisition and release resources (for example, send to the system's audio output module for real-time playback, or save as an audio file)
[0060] The following combination Figure 2 The specific scenario in the example is described.
[0061] Figure 2 1 is a schematic diagram of a scenario in which the audio system captures sound during screen recording and desktop video live broadcast provided by the embodiment of the present application. Figure 2 As shown, the audio system includes an audio input module, an audio processing module, an audio output module and a user.
[0062] The audio input module is used to convert the sound emitted by the user into an electrical signal and input it into the audio processing module.
[0063] The audio processing module is used to capture the audio stream input by the audio input module and output the captured audio data to the audio output module.
[0064] The audio output module is used to convert the audio data output by the audio processing module into corresponding sound signals and radiate them into the space for users to listen directly.
[0065] It should be understood that the audio system provided in the embodiment of the present application is applicable to Windows system, Linux system, Android system, iPhone Operating System (iOS), etc. The embodiment of the present application is not specifically limited.
[0066] Combined with the previous article Figure 1 The total time of sound capture starts from segment A, passes through segment B, and ends at segment C. However, since segment B does not capture any sound, the system sampling will directly skip segment B where the sound cannot be captured normally. The sound data finally obtained includes the sound data packets of segment A and segment C.
[0067] In the embodiment of the present application, in the sound capture stage, if the audio system fails to capture a sound data packet for a certain period of time (for example, the speaker pauses or is silent for a long time), in order to ensure that the subsequently captured audio data can correspond to the correct video screen time point, the audio data of the abnormal period without audio output is supplemented (for example, the system determines during the capture process that the audio data is not captured for a certain period of time). Figure 1If an exception occurs in segment B, the compensation packet generated by the embodiment of the present application is inserted between the captured segment A sound data packet and segment C sound data packet according to the time attribute, so that the audio data output later includes the segment A sound data packet, the segment B compensation packet and the segment C sound data packet), so as to avoid the problem of audio and video being out of sync during recording or live broadcasting.
[0068] In view of this, the present application deals with a sound compensation method, which can automatically compensate for the natural time without sound output, not only to ensure that the audio sampling duration is consistent with the video sampling duration, but also to accurately match the audio sampling and video sampling, thereby solving the problem of audio and video out of sync, and at the same time improving the user's use and viewing experience.
[0069] To facilitate understanding by those skilled in the art, Figure 3 to Figure 4 The process of performing sound compensation according to the embodiment of the present application is described.
[0070] Figure 3 is a flow chart of the sound compensation method provided in the embodiment of the present application. It should be understood that it is an example rather than a limitation. Figure 3 The method in can be applied to Figure 1 In the application scenario shown in Figure 3 As shown, the method comprises the following steps:
[0071] Step S110: Determine a first waiting period based on a first parameter, a first cache capacity, and a cache duration, wherein the first waiting period refers to the waiting period for capturing sound, the first parameter is used to characterize the sampling rate of the sound, and the first cache capacity is used to characterize the spatial capacity of caching sound data.
[0072] It should be understood that the first waiting period here is only for the convenience of description and does not limit the scope of protection of the embodiments of the present application.
[0073] Exemplarily, the first parameter represents the sampling rate of the sound. The sampling rate refers to the number of samples taken per second on each channel.
[0074] The first cache capacity refers to the space capacity for caching sound data, and the space capacity for caching sound data is actually the actual cache capacity allocated by the system.
[0075] The cache duration is the maximum duration for storing audio data. Just as the storage limit of MP3 is 100 songs, if more than 100 songs are stored, no more songs can be cached. If the cache duration is set to five hours, the total duration of all audio stored in the buffer must not exceed five hours. The cache duration is equivalent to the maximum capacity or maximum length of the buffer.
[0076] Optionally, step S110 includes: calculating the first waiting time by using the following formula:
[0077]
[0078] Among them, sleeptime represents the first waiting time, T represents the cache duration, B represents the first cache capacity, and S represents the first parameter.
[0079] The embodiment of the present application does not specifically limit the unit of the first waiting time. In a possible implementation, the unit of the first sound capture time is milliseconds.
[0080] It should be understood that the above is only an example of the method for determining the first waiting time, and the embodiment of the present application does not limit the specific method for determining the first waiting time.
[0081] Optionally, in a possible implementation, before step S110, the method further includes: obtaining a sound output format, the sound output format including: the first parameter, a second parameter, and a third parameter, wherein the second parameter is used to characterize the number of sampling channels of the sound, and the third parameter is used to characterize the number of sampling bytes.
[0082] It should be noted that the sound output format is obtained so as to obtain relevant parameters of the sound output format for calculation. Moreover, the obtained sound output format needs to be consistent with the sound format captured by the system (or the same format) to avoid sound distortion.
[0083] For example, in one possible implementation, a system interface, such as IAudioClient::Initialize interface, is called to obtain the sound output format. In another possible implementation, the system interface is also used to initialize the buffer inside the audio.
[0084] For example, in a possible implementation, after calling the system interface, the first buffer capacity is obtained through the system interface, for example, IAduioClient::GetBufferSize. The first buffer capacity here is the buffer size or buffer space capacity actually used by the system, and its common units include KB, MB, etc.
[0085] Optionally, in a possible implementation, the cache duration is set by modifying the buffer duration in the system. Setting the cache duration here is equivalent to specifying the maximum value of the buffer storage space.
[0086] It should be understood that the above is only an example of a method for obtaining the sound output format, and the embodiment of the present application does not limit the specific method for obtaining the sound output format.
[0087] It should be understood that the above is only an example of how to obtain the cache duration, and the embodiment of the present application does not limit the specific method for obtaining the cache duration.
[0088] It should be understood that the above is only an example of describing the method for obtaining the first cache capacity, and the embodiment of the present application does not limit the specific method for obtaining the first cache capacity.
[0089] It should be understood that the first waiting time here is only an exemplary description, and the first waiting time may also be named in other ways, such as system sound capture waiting time. The embodiment of the present application does not specifically limit this.
[0090] It should be understood that the first cache capacity here is only an exemplary description, and the first cache capacity may also be named in other ways, such as the actual cache allocated by the system. The embodiment of the present application does not specifically limit this.
[0091] Step S120: Execute sound capture and record a first time, wherein the first time refers to the start time of the sound capture.
[0092] Exemplarily, before starting sound capture, a first time is recorded. During the first waiting time, the system performs sound capture.
[0093] It should be understood that the embodiments of the present application do not impose any specific limitation on the execution method of sound capture.
[0094] Exemplarily, in a possible implementation, during the first waiting period, a continuous analog signal is captured by an audio input module (eg, a microphone) in the system, converted into a digital signal and stored in an internal audio buffer.
[0095] It should be understood that the above is only an example of a method for performing sound capture, and the embodiments of the present application do not limit the specific method for performing sound capture.
[0096] It should be understood that the first time here is only an exemplary description, and the first time may also be named in other ways, such as the current capture start time or beginT. The embodiment of the present application does not specifically limit this.
[0097] Step S130: After the first waiting period ends, determine whether sound data is detected.
[0098] If no sound data is detected, execute step S140-1; if sound data is detected, execute steps S140-2-1 to S140-2-4 (these steps will be described later).
[0099] It should be noted that after the first waiting period ends, the system needs to determine whether sound data is captured so as to perform subsequent operations based on the determination result.
[0100] The specific means of detecting sound data is not limited in the embodiment of the present application. Optionally, in a possible implementation, after the sound capture is completed, a system interface is called, the system interface is used to obtain the sound data; and according to the null value returned by the system interface, it is determined that the sound data is not detected.
[0101] Exemplarily, the system interface is IAudioClient::GetNextPacketSize. The IAudioClient::GetNextPacketSize interface is used to obtain the number of audio frames in the newly saved sound data.
[0102] Exemplarily, after the first waiting period ends, a system interface, for example, the IAudioClient::GetNextPacketSize interface, is called to determine whether sound data is detected based on the value returned by the system interface.
[0103] For example, if a null value is returned, it is determined that the sound data is not detected; if the returned value is greater than 0, that is, the IAudioClient::GetNextPacketSize interface obtains a voice frame, it is determined that the sound data is detected. In other words, different values can be used to determine whether the system captures sound data within the first waiting time.
[0104] It should be understood that the above is only described as an example to determine whether sound data is detected, and the embodiment of the present application does not limit the specific method of determining whether sound data is detected.
[0105] Step S140-1: If no sound data is detected, record a second time and determine a first abnormal duration, wherein the second time refers to the time when no sound data is detected, and the first abnormal duration refers to the duration when the sound data is not captured, and the first abnormal duration is determined based on the first time and the second time.
[0106] It should be understood that the terms "first", "second", "third", "fourth", etc. are introduced here only to distinguish the abnormal duration, capture time, etc. in different stages in the embodiments of the present application, and are not limited to this.
[0107] Exemplarily, during the sound capture time periods of two consecutive stages, if the system detects that the sound data cannot be detected, the system enters the exception handling process and obtains the time when the sound data is not detected.
[0108] It should be understood that the embodiment of the present application does not impose any specific limitation on the method of calculating the first abnormality duration.
[0109] Optionally, in a possible implementation, the first abnormal duration is determined based on the first time and the second time, including: the first abnormal duration is the duration of the second time minus the first time. Exemplarily, the first abnormal duration refers to a time period during which no sound data is captured after exceeding the first waiting time.
[0110] For example, suppose that after completing the sound capture of section A, the system starts to perform the sound capture of section B. The first time represents the time when the sound capture starts, the second time represents the time when no sound data is detected, and the first waiting time that is an integer multiple of the first time interval can also be understood as the time when the system detects the current abnormality. The first time here can be understood as Figure 1 The left endpoint of segment B in the middle, the second time can be understood as Figure 1 After obtaining the first time and the second time, the length of the first abnormal duration (for example, the length of segment B) can be determined by the difference between the first time and the second time.
[0111] It should be understood that the examples here are only for ease of understanding. In actual situations, the length of the first abnormal duration is much smaller than the length of segment B. Optionally, in one possible manner, the first abnormal duration is calculated using the following formula:
[0112] First abnormal duration = second time - first time
[0113] For example, in another possible implementation, the first time is also called beginT, the second time is also called endT, and the first abnormal duration is also called duration. The above formula for calculating the first abnormal duration can also be expressed as:
[0114] duration = endT - bedinT
[0115] It should be understood that the above is only an example of a method for determining the first abnormality duration, and the embodiment of the present application does not limit the specific method for determining the first abnormality duration.
[0116] Optionally, in another possible implementation, the actual duration of each sampling may be timed, and a value may be directly returned by calling a system command, thereby obtaining the first abnormal duration. The embodiment of the present application does not specifically limit this.
[0117] It should be understood that the above is only an example of a method for determining the first abnormality duration, and the embodiment of the present application does not limit the specific method for determining the first abnormality duration.
[0118] It should be understood that the first abnormality duration here is only an example description, and the first abnormality duration may also be named in other ways, such as capturing abnormality duration, and the present application does not limit this.
[0119] It should be understood that the second time here is only an exemplary description, and the second time may also be named in other ways, such as the current abnormal start time. This embodiment of the present application does not specifically limit this.
[0120] Step S150: Determine whether the first abnormal duration is greater than or equal to a first time threshold.
[0121] Exemplarily, if it is determined that the first abnormal duration determined in step S150 is greater than or equal to the first time threshold, it is necessary to subsequently perform a silence compensation operation on the corresponding time period, or in other words, supplement sound data to the corresponding time period.
[0122] Optionally, in a possible implementation, the method further includes: the first time threshold is determined according to the first parameter and the unit data amount of the first compensation frame. The unit data amount of the first compensation frame refers to the number of audio frames required to supplement the audio of the first waiting time length in a single sampling channel, and the unit data amount of the first compensation frame is determined according to the first parameter and the first waiting time length.
[0123] Optionally, step S150 includes: calculating the first time threshold using the following formula:
[0124]
[0125] Among them, MaxT represents the first time threshold, SF represents the unit data amount of the first compensation frame, and S represents the sampling rate in the sound output format.
[0126] It should be understood that the above is only an example of the method for determining the first time threshold, and the embodiment of the present application does not limit the specific method for determining the first time threshold.
[0127] The embodiment of the present application does not specifically limit the unit of the first time threshold. In a possible implementation, the unit of the first time threshold is microseconds.
[0128] It should be understood that although the concept of the unit data amount of the first compensation frame is introduced in step S150, and the calculation formula of the unit data amount of the first compensation frame is given later, in actual operation, the unit data amount of the first compensation frame has been determined when calculating the first waiting time in step S110. It is introduced here only to facilitate understanding of how to obtain the first time threshold.
[0129] It should be noted that the first waiting time determined in step S110 is a calculated fixed value. The first time threshold here is the maximum time value that is allowed not to add audio compensation data.
[0130] It should be noted that when it is determined that the first abnormal duration is less than or equal to the first time threshold, the sound compensation operation will not be performed, but the next sound capture process will be entered to perform sound capture and determine whether sound data is detected. The reason is that, first, the unit of the first time threshold is microseconds, for example, the first time threshold is 50 microseconds, that is, 0.00005 seconds. Secondly, during the sound capture stage, the sound is continuously and stably output, and there will not be multiple discontinuous silent periods or sound periods in microseconds that do not exceed the first time threshold. And for the user, it is almost impossible to perceive the existence of tens of microseconds of silent blanks. Therefore, even if a blank that does not exceed the first time threshold is generated in a certain judgment cycle, it will not affect the user; that is, if the first abnormal duration is less than the first time threshold, sound compensation can be performed without execution. It should be understood that the embodiment of the present application does not specifically limit the method for determining the first time threshold.
[0131] It should be understood that the first time threshold here is only an exemplary description, and the first time threshold may also be named in other ways, such as maximum waiting time. The embodiment of the present application does not specifically limit this.
[0132] By setting the first time threshold, it is possible to better determine the abnormal time period in which the audio data cannot be captured, thereby increasing the possibility of continuously acquiring the system sound, thereby ensuring the stable output of the audio sampling file.
[0133] Step S160: when the first abnormal duration is greater than or equal to a first time threshold, determining a first compensation frame, where the first compensation frame refers to a voice data frame that needs to be silence compensated.
[0134] Exemplarily, when the first abnormal duration is greater than or equal to the first time threshold, silence compensation needs to be performed on the first abnormal duration, that is, sound data is supplemented for the abnormal event segment without audio output. Before performing silence compensation, the amount of supplemented sound data, that is, the number of first compensation frames, needs to be determined.
[0135] It should be understood that the embodiment of the present application does not impose any specific limitation on the method of determining the first compensation frame.
[0136] Optionally, in a possible implementation, when the first abnormal duration is greater than or equal to a first time threshold, determining a first compensation frame includes: determining a unit data amount of the first compensation frame based on the first parameter and the first waiting duration, the unit data amount being used to characterize the number of audio frames required to supplement the audio of the first waiting duration in a single sampling channel; and determining a total number of the first compensation frames based on the second parameter, the third parameter, and the unit data amount.
[0137] Exemplarily, in a possible implementation, step S160, according to the sampling rate of the sound and the first waiting time, determines the unit data volume of the first compensation frame; according to the sampling channel of the sound, the number of sampling bytes of the sound, and the unit data volume, determines the total number of the first compensation frames. Optionally, it includes: using the following formula to calculate the total number of the first compensation frames:
[0138] CN=C*NB*SF
[0139] Among them, CN represents the total number of the first compensation frames, C represents the second parameter, NB represents the third parameter, and SF represents the unit data amount of the first compensation frame.
[0140] The calculation formula of the unit data amount of the first compensation frame is as follows:
[0141]
[0142] Among them, SF represents the unit data amount of the first compensation frame, sleeptime represents the first waiting time, and S represents the first parameter.
[0143] It should be understood that the above is only an example of describing the method of determining the first compensation frame, and the embodiment of the present application does not limit the specific method of determining the first compensation frame.
[0144] It should be understood that the unit data volume of the first compensation frame here is only an exemplary description, and the unit data volume of the first compensation frame can also be named in other ways, such as the number of silence compensation frames. This embodiment of the application does not specifically limit this.
[0145] It should be understood that the first compensation frame here is only an exemplary description, and the first compensation frame may also be named in other ways, such as compensation packet. This embodiment of the present application does not specifically limit this.
[0146] Step S170: Based on the first compensation frame, perform silence compensation on the system sound.
[0147] Exemplarily, an audio compensation operation is performed for the time period corresponding to the first abnormal duration, or in other words, sound data is added to the corresponding time period. The sound data may be silent data or white noise data, which is not specifically limited in the present embodiment of the application.
[0148] Optionally, in one possible manner, the silence compensation for system sound is performed based on the first compensation frame, including: sending the first compensation frame to a sound data queue stored in the system according to a time attribute of the first compensation frame; and filling a first abnormal duration according to the first compensation frame.
[0149] For example, Figure 1 Taking the example for description, segment B is the time period when the sound cannot be captured normally, that is, the time period without audio output. After determining that the duration of segment B exceeds the first time threshold, perform a silent compensation operation on segment B. Determine the first compensation frame required for the compensation operation, and send the first compensation frame to the sound data queue stored in the system. Because the data of the first compensation frame itself has a time attribute (because the sampling rate involved in the calculation can be converted to the corresponding specific time), after the system receives the first compensation frame, it will insert the first compensation frame into the sound data queue arranged in the order of capture time according to the time information carried thereon (that is, insert it between the captured sound data of segments A and C, and fill segment B through the first compensation frame).
[0150] For example, in the sound capture phase, the system has captured segment A sound data packets (at the beginning of the 1st to 5th milliseconds) and segment C sound data packets (at the 6th to 15th milliseconds). After confirming that segment B (5th to 6th milliseconds) is of abnormal duration, the number of first compensation frames required to fill segment B is calculated according to the method of an embodiment of the present application, and the first compensation frame is sent to the queue for storing sound data in the system. In the queue for storing sound data, all sound data packets are strictly arranged in chronological order according to the capture time, and the first captured sound data is ranked first to be output. After the first compensation frame enters the queue for storing sound data, it is also stored according to the time information thereon. The above-mentioned operation of inserting the first compensation frame between sound data packets can also be understood as the so-called filling of segment B audio.
[0151] It should be noted that the above examples are only for ease of understanding, and the embodiments of the present application are not limited thereto. In actual situations, due to some reasons (for example, the songs that are continuously and stably output in segment A are delayed by the network), the playback is stuck, so segment B, or the time when the external sound is in a silent state, will be longer, for example, tens of seconds or minutes. According to the method provided in the embodiment of the present application, the compensation frame of the length of the first waiting time will be continuously and repeatedly supplemented in segment B (the time period of the playback stuck), until segment C (the song is stably played again) starts.
[0152] Through the above operations, it is possible to ensure that the audio sampling duration is consistent with the video sampling duration during the sound capture stage, and automatically add sound data to the time period without audio if the sound is not captured for a long time, to ensure that the sound data captured later corresponds to the correct video picture time point, thereby avoiding the problem of audio and video being out of sync.
[0153] In an embodiment of the present application, after the sound output format is obtained, the sound capture time required for sound capture in each stage is calculated in reverse; the process of sound capture and detection of whether sound data is captured is executed in sequence in each stage until the entire sound capture time period ends. Wherein, after each determination that sound data is not captured, the actual abnormal duration of sound capture is determined, and on the premise that the abnormal duration of sound capture is determined to be greater than the first time threshold, audio data compensation is performed on the abnormal duration of capture. Compared with the prior art, the sound compensation method provided in the embodiment of the present application effectively avoids the situation in which the system skips the time period without audio data due to the existence of a time period in which audio is not captured during video recording or live broadcasting, and the final audio sampling duration cannot be synchronized with the video screen, that is, the problem of audio and video being out of sync is caused, and the captured audio data and video screen data are accurately matched, which greatly improves the user experience.
[0154] It should be noted that, for ease of understanding, the embodiment of the present application is described by taking the first waiting time in the entire sound capture time period as an example, and does not limit the entire sound capture time period to only one sound capture. In actual operation, the entire sound capture time period includes multiple sound capture time periods. The embodiment of the present application does not limit this.
[0155] exist Figure 3 Based on the method shown, another embodiment of the present application provides an example in which, in step 130, a sound data packet is detected after the first waiting period ends.
[0156] Exemplarily, after the first waiting period ends, the system interface is called, and based on a returned value greater than 0, it is determined that the sound data packet is captured.
[0157] Optionally, in one possible implementation, Figure 3 The method shown also includes:
[0158] Step S140-2-1, if a sound data packet is detected after the first waiting period ends, a third time is recorded, where the third time is the time when the sound data packet is detected;
[0159] Step S140-2-2, updating the first time to the third time;
[0160] Step S140-2-3, performing sound capture from the third time according to the first waiting time; after the first waiting time is over, determining whether sound data is detected;
[0161] Step S140-2-4, if no sound data is detected, record the fourth time and determine the second abnormal duration, wherein the fourth time refers to the time when no sound data is detected, and the second abnormal duration is determined based on the third time and the fourth time.
[0162] Exemplarily, if a sound data packet is detected after the sound capture is completed in step 120, the next round of sound capture process will be entered, the sound capture start time will be updated, and the sound capture will be re-executed within the first waiting period, and the above process will be repeated. In the process of re-executing the sound capture, similarly, after the first waiting period ends, it is still necessary to detect whether there is a sound data packet. Among them, after each round of the sound capture process ends, it is necessary to determine whether there is an abnormal duration based on the returned detection results, and determine whether it is necessary to perform silence compensation for the abnormal duration based on the first time threshold. If the returned detection result shows that it is normal, the next round of sound capture process will be re-entered until the entire sound capture stage ends.
[0163] It should be understood that the specific implementation process of step S140-2-3 and step S140-2-4 can refer to the above-mentioned step S120 to step S140-1, which will not be described here for brevity.
[0164] It should be understood that the sound data packet here is only an exemplary description, and the sound data packet may also have other naming methods, such as capture packet. The embodiment of the present application does not specifically limit this.
[0165] To facilitate understanding by those skilled in the art, the following Figure 4 The specific example in describes the sound compensation method provided by the present application. It should be understood that Figure 4 The examples are only for the convenience of those skilled in the art to understand the sound compensation method provided by the embodiments of the present application, and are not intended to limit the embodiments of the present application to the specific scenarios illustrated. Figure 4 It is obvious that various equivalent modifications or changes can be made to the examples in the present application, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0166] Figure 4 It is an example flow chart of the sound compensation method provided in the embodiment of the present application. It can be understood that it is an example rather than a limitation. Figure 4 The method in can be applied to Figure 1 In the application scenario shown. Figure 4 The relevant terms or explanations involved in the above description can be found in the previous section and will not be repeated below. Figure 4 As shown, the specific steps include:
[0167] Step S1: Obtain the sound output format.
[0168] It should be understood that in order to avoid the sound quality of the audio output module being damaged or affecting the user experience, the sound output format should be consistent with the sound input format. Therefore, to determine the sound output format, it is necessary to first obtain the sound input format captured by the system.
[0169] Exemplarily, the sound input format captured by the system is obtained by calling the IAudioClient::GetMixFormat interface. Optionally, the sound input format includes one or more of the following parameters: sampling rate, sampling channel, and sampling byte number.
[0170] Exemplarily, according to the acquired sound input format, the IAudioClient::Initialize interface is called to determine the system sound output format. Specifically, the system sound output format also includes one or more of the following parameters: sampling rate, sampling channel, and sampling byte number.
[0171] It should be understood that the specific method of obtaining the sound output format can refer to the embodiment of the present application. Figure 3 And step S110 is not repeated here for the sake of brevity.
[0172] Step S2: Determine the cache duration and the actual cache size.
[0173] It should be understood that the embodiment of the present application does not impose any specific restrictions on the method of setting the cache time. For the specific method of determining the cache time, reference can be made to step S110 provided in the embodiment of the present application, which will not be described here for brevity.
[0174] It should be understood that the embodiments of the present application do not impose any specific limitation on the method of determining the actual cache.
[0175] Exemplarily, after the system is started, the IAduioClient::GetBufferSize interface is called to obtain the actual buffer size allocated by the system.
[0176] For the specific method of determining the actual cache, reference may be made to step S110 provided in the embodiment of the present application, which will not be described again here for the sake of brevity.
[0177] Step S3: Calculate the system sound capture waiting time.
[0178] It should be understood that the "system sound capture waiting time" here is another way of expressing the term "first waiting duration" in the above step S110.
[0179] It should be noted that the system sound capture waiting time calculated here is not the entire sound capture waiting time period. The entire sound capture waiting time period includes multiple sound capture waiting times. During each sound capture waiting time, the system captures external sounds or receives electrical signals generated by the audio input module.
[0180] Optionally, the system sound capture waiting time is determined according to the sound output format acquired in step S1 and the actual buffering and buffering time in step S2.
[0181] It should be understood that the specific method of calculating the system sound capture waiting time can refer to step S110 provided in the embodiment of the present application, and for the sake of brevity, it will not be repeated here.
[0182] Step S4: Start sound capture.
[0183] Exemplarily, after the system sound capture waiting time is calculated in step S3, the first phase of sound capture begins. Before starting the sound capture, the current capture time is recorded.
[0184] Step S5: After the current sound capture waiting time ends, determine whether a capture packet is detected.
[0185] It should be understood that the "capture packet" here is another way of expressing the term "sound data packet" in step S130 above.
[0186] Exemplarily, after the first phase of sound capture is completed, the IAudioClient::GetNextPacketSize interface is called to detect whether there is a capture packet, that is, to detect whether a system sound is generated.
[0187] Exemplarily, if a capture packet is detected in step S5, step S9 is directly entered to send the capture packet to the audio output module in sequence, and the current capture start time is updated, thereby indicating that the sound capture before this time point is normal. The system enters the second phase of the sound capture process. The sound capture process includes: step S4: performing a capture operation within the sound capture time; step S5: after the sound capture time ends, determining whether a capture packet is detected. Specifically, the system continuously repeats the process of step S4 and step S5 until the entire sound capture time period ends.
[0188] Exemplarily, if no voice packet is detected, the exception handling process is entered into step S6.
[0189] It should be understood that the specific method of determining whether a capture packet is detected can refer to step S130 provided in the embodiment of the present application, and for the sake of brevity, it will not be repeated here.
[0190] Step S6: Calculate the duration of capturing the exception.
[0191] It should be understood that the "abnormal capture duration" here is another way of expressing the term "first abnormality duration" in the above step S140.
[0192] Exemplarily, the current abnormal capture time and the current capture time are obtained to determine the abnormal capture duration.
[0193] It should be understood that the specific method for calculating the duration of capturing anomalies can refer to step S140 provided in the embodiment of the present application, and for the sake of brevity, it will not be repeated here.
[0194] Step S7: Determine whether the abnormality capture time is greater than the maximum waiting time.
[0195] It should be understood that the “maximum waiting time” here is another way of expressing the term “first time threshold” in the above step S150.
[0196] Exemplarily, in a possible implementation, the maximum capture waiting time is calculated according to the sound output format acquired in step S1 and the buffer time set in step S2.
[0197] Exemplarily, if the capture abnormality duration determined in step S6 is greater than the maximum capture waiting time, the audio data is compensated for the time period corresponding to the capture abnormality duration. The audio data may be either silent data or white noise data.
[0198] It should be understood that the specific calculation of the maximum capture waiting time and the method of determining whether the maximum capture waiting time is exceeded can refer to step S150 provided in the embodiment of the present application, and for the sake of brevity, it will not be repeated here.
[0199] Step S8: Calculate the compensation quantity required in the compensation package.
[0200] It should be understood that the “compensation packet” here is another way of expressing the term “first compensation frame” in the above step S160 .
[0201] It should be understood that the specific method for calculating the amount of compensation required in the compensation package can refer to step S160 provided in the embodiment of the present application, and for the sake of brevity, it will not be repeated here.
[0202] Step S9: Send compensation packet / capture packet.
[0203] Exemplarily, in step S5, when a capture packet is detected, the capture packet is directly sent to the buffer area of the system, wherein the capture packet includes captured sound data.
[0204] Exemplarily, in step S5, if no capture packet is detected, then step S6 is used to calculate the capture abnormality duration, step S7 is used to determine whether it is greater than the abnormal duration, step S8 is used to calculate the compensation quantity of the compensation packet, and step S9 is used to send the compensation packet to the buffer area of the system. The compensation packet includes silence data or white noise data for silence compensation.
[0205] It should be understood that the embodiments of the present application do not specifically limit the manner of sending compensation packets and capture packets.
[0206] Exemplarily, the compensation packet is accurately inserted into the audio data playback queue according to the time attribute thereon, and fills the abnormal time period in which the system does not capture the audio.
[0207] It should be understood that the specific method of sending the compensation packet can refer to step S170 provided in the embodiment of the present application. For the sake of brevity, it will not be repeated here.
[0208] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0209] Corresponding to the sound compensation method described in the above embodiment, Figure 5 A structural block diagram of a sound compensation device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0210] Reference Figure 5 The device 500 includes: a determination unit 510, a sound capturing unit 520, a detection unit 530, and a silence compensation unit 540.
[0211] In some possible implementations, the determining unit 510 is used to determine a first waiting time according to a first parameter, a first cache capacity, and a cache duration, wherein the first waiting time refers to a waiting time for capturing sound, the first parameter is used to characterize a sampling rate of the sound, and the first cache capacity is used to characterize a space capacity for caching sound data;
[0212] In some possible implementations, the sound capturing unit 520 is used to perform sound capturing and record a first time, wherein the first time refers to a start time of the sound capturing;
[0213] In some possible implementations, the detection unit 530 is configured to determine whether sound data is detected after the first waiting period ends;
[0214] Optionally, the detection unit 530 is configured to determine whether sound data is detected after the first waiting period ends, including:
[0215] After the first waiting period ends, calling the interface;
[0216] Based on the value returned by the interface, determine whether sound data is detected.
[0217] In some possible implementations, the determining unit 510 is further configured to, if no sound data is detected, record a second time and determine a first abnormal duration, wherein the second time refers to a time when no sound data is detected, the first abnormal duration refers to a duration during which no sound data is captured, and the first abnormal duration is determined based on the first time and the second time;
[0218] In some possible implementations, the determining unit 510 is further configured to determine whether the first abnormal duration is greater than or equal to a first time threshold;
[0219] In some possible implementations, the determining unit 510 is further configured to, when the first abnormal duration is greater than or equal to a first time threshold, determine a first compensation frame, where the first compensation frame refers to a voice data frame that requires silence compensation;
[0220] In some possible implementations, the silence compensation unit 540 is configured to perform silence compensation on system sound based on the first compensation frame.
[0221] Optionally, the silence compensation unit 540 is configured to perform silence compensation on system sound based on the first compensation frame, including:
[0222] After determining the number of the first compensation frames, inserting the first compensation frames into corresponding positions in a sound data queue stored in the system according to time attributes of the first compensation frames;
[0223] Wherein, determining the first compensation frame includes:
[0224] Determine, according to the first parameter and the first waiting time, a unit data amount of the first compensation frame, where the unit data amount is used to represent the number of audio frames required to supplement the audio of the first waiting time in a single sampling channel;
[0225] The total number of the first compensation frames is determined according to the second parameter, the third parameter, and the unit data amount.
[0226] Optionally, the device 500 further includes:
[0227] The sound output format is obtained, where the sound output format includes: the first parameter, the second parameter, and the third parameter, wherein the second parameter is used to characterize the number of sampling channels of the sound, and the third parameter is used to characterize the number of sampling bytes.
[0228] Optionally, the apparatus 500 further comprises: if a sound data packet is detected after the first waiting period ends, recording a third time, the third time being the time when the sound data packet is detected;
[0229] Updating the first time to the third time;
[0230] According to the first waiting time, starting to perform sound capture from the third time;
[0231] After the first waiting period ends, determining whether sound data is detected;
[0232] If no sound data is detected, a fourth time is recorded and a second abnormal duration is determined, wherein the fourth time refers to a time when no sound data is detected, and the second abnormal duration is determined based on the third time and the fourth time.
[0233] Optionally, the unit data amount of the first compensation frame satisfies the following formula: Among them, SF represents the unit data amount of the first compensation frame, sleeptime represents the first waiting time, and S represents the first parameter.
[0234] Optionally, the total number of the first compensation frames satisfies the following formula: CN=C*NB*SF, wherein CN represents the total number of the first compensation frames, C represents the second parameter, NB represents the third parameter, and SF represents the unit data amount of the first compensation frames.
[0235] Optionally, the device 500 further includes: the first time threshold is determined according to the first parameter and the unit data volume.
[0236] Optionally, the first time threshold satisfies the following formula: Among them, MaxT represents the first time threshold, SF represents the unit data amount of the first compensation frame, and S represents the first parameter.
[0237] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0238] Figure 6 This is a schematic diagram of the structure of a sound compensation device provided in one embodiment of the present application. Figure 6As shown, the sound compensation device 6 of this embodiment includes: at least one processor 60 ( Figure 6 Only one is shown in the figure) a processor, a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60, wherein the processor 60 implements the steps of any of the above-mentioned sound compensation method embodiments when executing the computer program 62.
[0239] The sound compensation device / terminal device 6 can be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The sound compensation device / terminal device can include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 6 It is merely an example of the sound compensation device / terminal device 6 and does not constitute a limitation on the sound compensation device / terminal device 6. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components, for example, it may also include input and output devices, network access devices, etc.
[0240] The processor 60 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0241] In some embodiments, the memory 61 may be an internal storage unit of the sound compensation device / terminal device 6, such as a hard disk or memory of the sound compensation device / terminal device 6. In other embodiments, the memory 61 may also be an external storage device of the sound compensation device / terminal device 6, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the sound compensation device / terminal device 6. Further, the memory 61 may also include both the internal storage unit of the sound compensation device / terminal device 6 and an external storage device. The memory 61 is used to store an operating system, an application program, a boot loader (BootLoader), data and other programs, such as the program code of the computer program, etc. The memory 61 may also be used to temporarily store data that has been output or is to be output.
[0242] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0243] An embodiment of the present application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps in any of the above-mentioned method embodiments when executing the computer program.
[0244] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0245] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0246] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.
[0247] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0248] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0249] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0250] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0251] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A sound compensation method, applied to the sound capture stage, characterized in that: include: Determine a first waiting time according to a first parameter, a first cache capacity, and a cache duration, wherein the first waiting time refers to a waiting time for capturing sound, the first parameter is used to characterize a sampling rate of the sound, and the first cache capacity is used to characterize a space capacity for caching sound data; Execute sound capture and record a first time, wherein the first time refers to a start time of the sound capture; After the first waiting period ends, determining whether sound data is detected; If no sound data is detected, record a second time and determine a first abnormal duration, wherein the second time refers to the time when no sound data is detected, the first abnormal duration refers to the duration when no sound data is captured, and the first abnormal duration is determined based on the first time and the second time; Determine whether the first abnormal duration is greater than or equal to a first time threshold, where the first time threshold is determined according to the first parameter and a unit data volume; When the first abnormal duration is greater than or equal to a first time threshold, determining a first compensation frame, where the first compensation frame refers to a voice data frame that needs to be silence compensated; Based on the first compensation frame, silence compensation is performed on the system sound.
2. The method according to claim 1, characterized in that The method further comprises: The sound output format is obtained, where the sound output format includes: the first parameter, the second parameter, and the third parameter, wherein the second parameter is used to characterize the number of sampling channels of the sound, and the third parameter is used to characterize the number of sampling bytes.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: If a sound data packet is detected after the first waiting period ends, a third time is recorded, where the third time is the time when the sound data packet is detected; Updating the first time to the third time; According to the first waiting time, starting to perform sound capture from the third time; After the first waiting period ends, determining whether sound data is detected; If no sound data is detected, a fourth time is recorded and a second abnormal duration is determined, wherein the fourth time refers to a time when no sound data is detected, and the second abnormal duration is determined based on the third time and the fourth time.
4. The method according to claim 1 or 2, characterized in that: After the first waiting period ends, determining whether sound data is detected includes: After the first waiting period ends, calling the interface; Based on the value returned by the interface, determine whether sound data is detected.
5. The method according to claim 2, characterized in that: The performing silence compensation on the system sound based on the first compensation frame includes: After determining the number of the first compensation frames, inserting the first compensation frames into corresponding positions in a sound data queue stored in the system according to time attributes of the first compensation frames; Wherein, determining the first compensation frame includes: Determine, according to the first parameter and the first waiting time, a unit data amount of the first compensation frame, where the unit data amount is used to represent the number of audio frames required to supplement the audio of the first waiting time in a single sampling channel; The total number of the first compensation frames is determined according to the second parameter, the third parameter, and the unit data amount.
6. The method according to claim 5, characterized in that The unit data amount of the first compensation frame satisfies the following formula: Among them, SF represents the unit data amount of the first compensation frame, sleeptime represents the first waiting time, and S represents the first parameter.
7. The method according to claim 5 or 6, characterized in that: The number of the first compensation frames satisfies the following formula: CN=C*NB*SF, wherein CN represents the number of the first compensation frames, C represents the second parameter, NB represents the third parameter, and SF represents the unit data amount of the first compensation frame.
8. The method according to claim 1, characterized in that The first time threshold satisfies the following formula: Among them, MaxT represents the first time threshold, SF represents the unit data amount of the first compensation frame, and S represents the first parameter.
9. A sound compensation device, applied to the sound capture stage, characterized in that: It includes a determination unit, a sound capture unit, a detection unit, and a silence compensation unit, wherein: The determining unit is used to determine a first waiting time according to a first parameter, a first cache capacity, and a cache duration, wherein the first waiting time refers to a waiting time for capturing sound, the first parameter is used to characterize a sampling rate of the sound, and the first cache capacity is used to characterize a space capacity for caching sound data; The sound capturing unit is used to perform sound capturing and record a first time, wherein the first time refers to a start time of the sound capturing; The detection unit is used to determine whether sound data is detected after the first waiting period ends; The determining unit is further configured to, if no sound data is detected, record a second time and determine a first abnormal duration, wherein the second time refers to a time when no sound data is detected, the first abnormal duration refers to a duration when no sound data is captured, and the first abnormal duration is determined based on the first time and the second time; The determining unit is further used to determine whether the first abnormal duration is greater than or equal to a first time threshold, where the first time threshold is determined according to the first parameter and the unit data volume; The determining unit is further configured to, when the first abnormal duration is greater than or equal to a first time threshold, determine a first compensation frame, wherein the first compensation frame refers to a voice data frame that needs to be silence compensated; The silence compensation unit is used to perform silence compensation on system sound based on the first compensation frame.
Citation Information
Patent Citations
Audio stream compensation method and device, storage medium and equipment
CN113436639A