Audio transcoding method and device, electronic equipment, storage medium and program product
By detecting and dynamically adjusting the volume in the audio stream, the problem of volume inconsistency in the prior art is solved, and the consistency of volume and user experience are improved.
Patent Information
- Application Number
- CN202510585894.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-27
AI Technical Summary
The existing audio transcoding technology lacks a volume adjustment mechanism during the transcoding process, resulting in inconsistent volumes in different scenarios, affecting the user's audio-visual experience.
By detecting the audio content type and source stream volume of the original audio stream, adding this information to the audio stream using timestamp alignment and encapsulation insertion techniques, dynamically adjusting the target volume of the transcoded audio stream to ensure consistency of volume.
It realizes dynamic adjustment of volume during the transcoding process, unify the volume gain in a single and/or multiple audio streams, maintains the consistency of volume when switching different scenes, and improves the user's audio-visual experience.
Smart Images

Figure CN120220704A_ABST
Abstract
Description
Background Art
[0002] In the current context where the dissemination of digital multimedia content is becoming increasingly popular, audio transcoding technology, as the core technology for realizing the efficient transmission and adaptation of audio content, is widely used in many fields such as live streaming, video on demand, and online music. Currently, during audio transcoding operations, only fixed parameters such as audio sampling rate, bit depth, and number of channels are set at the beginning of transcoding. During the transcoding process, these parameters remain unchanged. If the audio source to be transcoded comes from different scenarios and the volumes of each scenario are inconsistent, the sound in the live stream will jitter, and the live stream generated by transcoding will also jitter, encountering the problem of fluctuating volume, which affects the user's audio-visual experience.
[0003] It should be noted that the information disclosed in the above Background Art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of the present disclosure is to provide an audio transcoding method, an audio transcoding device, an electronic device, and a computer-readable storage medium, which can at least to some extent improve the problem in related technologies that affects the user's audio-visual experience due to fluctuating volume.
[0005] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be learned in part through the practice of the present disclosure.
[0006] According to one aspect of the present disclosure, there is provided an audio transcoding method, including: in response to an acquired original audio stream, detecting the audio content type and the source stream volume of the original audio stream; based on an alignment operation with the timestamp of the original audio stream, adding the audio content type and the source stream volume to the original audio stream to obtain an enhanced audio stream; determining a target volume of the transcoded audio stream based on the relationship between the source stream volume and a reference volume, wherein when a switch of the audio content type is detected during the transcoding process, adjusting the second source stream volume after the switch based on the first source stream volume before the switch, so as to determine the target volume based on the relationship between the adjusted second source stream volume and the reference volume.
[0007] In one embodiment of the present disclosure, in response to the acquired original audio stream, detecting the audio content type and the source stream volume of the original audio stream includes: dividing the original audio stream based on a specified time window, collecting audio features at audio sampling points within each time window, inputting the audio features into a feature detection model, and outputting, by the feature detection model, the audio content type, the noise detection result, and the corresponding timestamp information, where the feature detection model is generated by relying on deep learning; and inputting the original audio stream into a volume detection model to correspondingly output a volume curve as the source stream volume.
[0008] In one embodiment of the present disclosure, based on an alignment operation with the timestamp of the original audio stream, adding the audio content type and the source stream volume to the original audio stream to obtain an enhanced audio stream includes: aligning, in units of the time window, the audio content type, the noise detection result, and the volume curve to obtain aligned information; encapsulating the aligned information into an SEI data unit based on the format of supplementary enhancement information SEI; and inserting the SEI data unit into the original audio stream based on the corresponding timestamp information to obtain the enhanced audio stream.
[0009] In one embodiment of the present disclosure, determining the target volume of the transcoded audio stream based on the relationship between the source stream volume and a reference volume includes: determining a reference period corresponding to the transcoding process, and determining the reference volume based on the volume attributes of transcoded audio files belonging to the reference period recorded in a big data module; calculating the ratio of the absolute difference between the source stream volume and the reference volume to the reference volume; if the ratio is less than or equal to a reference ratio, keeping the target volume as the source stream volume; and if the ratio is greater than the reference ratio, adjusting the target volume to the reference volume.
[0010] In one embodiment of the present disclosure, when a switch of the audio content type is detected during the transcoding process, adjusting the second source stream volume after the switch based on the first source stream volume before the switch includes: when a switch of the audio content type is detected, determining the median of the first source stream volume and the second source stream volume; and using the median as the adjustment result of the first source stream volume.
[0011] In one embodiment of the present disclosure, it further includes: parsing the noise detection result in the SEI data unit; and if it is determined based on the noise detection result that there is noise information in the original audio stream, performing a noise reduction function during the transcoding process.
[0012] In one embodiment of the present disclosure, before detecting the audio content type and the source stream volume of the obtained original audio stream in response to the obtained original audio stream, it further includes: collecting audio data in different scenarios; annotating the audio data based on the scenario type, the content type, and the presence or absence of noise to obtain annotated audio; converting the annotated audio into a Mel spectrogram, where the Mel spectrogram characterizes the time-frequency domain features of the annotated audio; training a deep learning model based on the Mel spectrogram to obtain the feature detection model.
[0013] In one embodiment of the present disclosure, before detecting the audio content type and the source stream volume of the obtained original audio stream in response to the obtained original audio stream, it further includes: constructing the volume detection model based on a volume detection formula, where the volume detection formula is: L p = 20 * log 10 (Prms / Pref) dB, where Prms is the sound amplitude value at any moment in the original audio stream, and Pref is the maximum reference value of the sound amplitude.
[0014] In one embodiment of the present disclosure, it further includes: reporting the transcoded audio stream including the target volume to the big data module based on the timestamp dimension to update the transcoded audio file in the big data module.
[0015] According to another aspect of the present disclosure, there is provided an audio transcoding device, including: a detection module, configured to detect the audio content type and the source stream volume of the obtained original audio stream in response to the obtained original audio stream; an adding module, configured to add the audio content type and the source stream volume to the original audio stream based on an alignment operation with the timestamp of the original audio stream to obtain an enhanced audio stream; a determining module, configured to determine the target volume of the transcoded audio stream based on the relationship between the source stream volume and the reference volume, where when a switch of the audio content type is detected during the transcoding process, the second source stream volume after the switch is adjusted based on the first source stream volume before the switch, so as to determine the target volume based on the relationship between the adjusted second source stream volume and the reference volume.
[0016] According to still another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the audio transcoding method of any one of the above via executing the executable instructions.
[0017] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the audio transcoding method of any one of the above.
[0018] The audio transcoding solution provided by the embodiments of the present disclosure obtains the audio content type and the source stream volume by detecting the audio content type and the source stream volume of the original audio stream, adds the audio content type and the source stream volume information to the original audio stream by using timestamp alignment and encapsulation insertion technology to form an enhanced audio stream, determines the target volume of the transcoded audio stream based on the reference volume setting, volume comparison, and volume adjustment strategy when the audio content is switched, thereby constructing an audio transcoding processing structure to dynamically adjust the volume without interruption during the transcoding process, unify the volume gain in a single and / or multiple audio streams, maintain the consistency of the volume when switching between different scenarios, and improve the user's audio-visual experience.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0021] Figure 1 A schematic diagram showing the structure of an audio transcoding system in an embodiment of the present disclosure;
[0022] Figure 2 A flowchart showing an audio transcoding method in an embodiment of the present disclosure;
[0023] Figure 3 A schematic diagram showing another structure of an audio transcoding system in an embodiment of the present disclosure;
[0024] Figure 4 A flowchart showing another audio transcoding method in an embodiment of the present disclosure;
[0025] Figure 5 A schematic diagram showing an audio transcoding device in an embodiment of the present disclosure;
[0026] Figure 6 A schematic diagram showing an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0028] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0029] At present, with the increasing popularity of digital multimedia content dissemination, audio transcoding technology, as the core technology for realizing efficient transmission and adaptation of audio content, is widely used in many fields such as live broadcast, video on demand, and online music. Currently, mainstream audio transcoding technologies generally follow a fixed parameter setting mode during operation. That is, at the start-up stage of the transcoding operation, technicians will set key parameters such as audio sampling rate, bit depth, and number of channels at one time according to conditions such as the target playback device and network bandwidth, and these parameters remain unchanged throughout the transcoding process. At the same time, the volume size of the live stream or on-demand file generated by transcoding completely depends on the volume level of the original stream or source file, and the transcoding system lacks an active volume adjustment mechanism.
[0030] Taking the live broadcast industry as an example, with the increasing richness of live broadcast content forms, it covers various types such as game live broadcast, talent show, online education, and e-commerce live streaming. During the live broadcast, the same live stream may experience multiple scene switches. For example, in a live broadcast party, there is an alternation between live performances and promotional pads. The types and volume sizes of audio content (such as human voices, music, ambient sounds, etc.) in different scenes are significantly different. In addition, in the scenario of switching between multiple live broadcast rooms, since each live stream is transcoded independently and no volume calibration is performed, when the user switches between live broadcast rooms, the sound in the live stream will jitter, and the live stream generated by transcoding will also jitter. Therefore, the user will also encounter the problem of fluctuating volume. This existing audio transcoding technology based on fixed parameter settings and lacking volume adjustment is difficult to meet the user's demand for a stable and comfortable audio-visual experience and urgently needs improvement and innovation.
[0031] Figure 1 It is a schematic structural diagram of a computer system provided by an exemplary embodiment of the present application. The system includes: a plurality of terminals 120 and a server cluster 140.
[0032] The terminal 120 may be a mobile terminal such as a mobile phone, a game console, a tablet computer, an e - book reader, smart glasses, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a smart home device, an AR (Augmented Reality) device, a VR (Virtual Reality) device, etc. Alternatively, the terminal 120 may also be a personal computer (PC), such as a laptop computer and a desktop computer, etc.
[0033] Among them, an application program for providing audio transcoding may be installed in the terminal 120.
[0034] The terminal 120 is connected to the server cluster 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0035] The server cluster 140 is a single server, or consists of several servers, or is a virtualization platform, or is a cloud computing service center. The server cluster 140 is used to provide background services for the application program that provides audio transcoding. Optionally, the server cluster 140 undertakes the main computing work, and the terminal 120 undertakes the secondary computing work; or, the server cluster 140 undertakes the secondary computing work, and the terminal 120 undertakes the main computing work; or, a distributed computing architecture is adopted between the terminal 120 and the server cluster 140 for collaborative computing.
[0036] In some alternative embodiments, the server cluster 140 is used to store audio transcoding program information.
[0037] Optionally, the clients of the application programs installed in different terminals 120 are the same, or the clients of the application programs installed on two terminals 120 are clients of the same type of application program on different control system platforms. Based on the differences in the terminal platforms, the specific forms of the clients of the application program may also be different. For example, the client of the application program may be a mobile phone client, a PC client, or a World Wide Web (Web) client, etc.
[0038] Those skilled in the art can know that the number of the above - mentioned terminals 120 may be more or less. For example, there may be only one of the above - mentioned terminals, or there may be dozens or hundreds of the above - mentioned terminals, or even more. The embodiments of the present application do not limit the number and device types of the terminals.
[0039] Optionally, the system may further include a management device ( Figure 1(not shown), the management device is connected to the server cluster 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0040] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0041] Next, the audio transcoding method in this exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.
[0042] As Figure 2 shown, an audio transcoding method according to an embodiment of the present disclosure includes:
[0043] Step S202, in response to the acquired original audio stream, detect the audio content type and the source stream volume of the original audio stream.
[0044] In some embodiments, the audio content type of the original audio stream can be detected based on the method of feature extraction and classification model, or based on the method of audio fingerprint, or based on the method of rule and pattern matching.
[0045] In some embodiments, the source stream volume is detected based on the method of peak detection, or based on the method of root mean square (RMS) energy, or based on the VAD (volume detection) algorithm, etc.
[0046] In some embodiments, taking a live broadcast party as an example, when the party starts the live broadcast and the system receives the original audio stream, for the detection of the audio content type, a pre-trained audio classification model can be used. For example, when the host starts speaking, the model identifies that the audio content type is the host's explanation; when the singer starts singing, the model identifies it as a song performance.
[0047] In some embodiments, for the detection of the source stream volume, it can be determined by calculating the average energy of the audio signal over a period of time. For example, when the host is speaking normally, the detected source stream volume is 70 dB; when the singer is singing at a high pitch, the source stream volume reaches 85 dB, etc.
[0048] Step S204, based on the alignment operation with the time stamp of the original audio stream, add the audio content type and the source stream volume to the original audio stream to obtain an enhanced audio stream.
[0049] In some embodiments, in a live broadcast party, the original audio stream is continuously updated, and the system records the time stamp of each audio segment. When the audio content type (such as the host's transition, the singer's performance) and the source stream volume are detected, these information are aligned with the corresponding time stamps, and in accordance with the format specification of SEI (Supplemental Enhancement Information), these information are encapsulated into data units, and the SEI data units are inserted into the corresponding positions of the original audio stream according to the time stamps to form an enhanced audio stream.
[0050] For example, when the host is introducing the next program, the system adds the audio content type (host's introduction) and the source stream volume (72 dB) information at this time to the audio stream position corresponding to the time stamp, which is convenient for subsequent audio analysis and processing.
[0051] Step S206, determine the target volume of the transcoded audio stream based on the relationship between the source stream volume and the reference volume. Among them, when it is detected that there is a switch in the audio content type during the transcoding process, adjust the second source stream volume after the switch based on the first source stream volume before the switch, so as to determine the target volume based on the relationship between the adjusted second source stream volume and the reference volume.
[0052] In some embodiments, the volume attributes (such as average volume, peak volume, etc.) of the transcoded audio file in the reference period can be analyzed through a big data module to determine a representative reference volume.
[0053] In some embodiments, during the transcoding process of a live broadcast party, the reference volume can also be set as the average volume of the audio of the same type of programs in the past of this party.
[0054] In some embodiments, the absolute difference ratio between the source stream volume and the reference volume can be calculated, and it is judged whether the source stream volume meets the transcoding requirements according to a preset reference ratio, so as to adjust the target volume.
[0055] In some embodiments, the ratio r of the source stream volume to the reference volume can also be calculated, and the target volume of the transcoded audio stream can be adjusted according to the calculated ratio r. If r >, it indicates that the source stream volume is greater than the reference volume, and the target volume can be set to the reference volume or reduced by a certain ratio. If r < 1, the target volume is set to the reference volume or enlarged proportionally.
[0056] In some embodiments, the source stream volume range and the reference volume range can also be linearly mapped, and the target volume can be calculated for the source stream volume through a linear interpolation formula.
[0057] In some embodiments, the target volume can also be dynamically adjusted by comprehensively considering the source stream volume, the reference volume, and other characteristics of the audio.
[0058] In some embodiments, when the audio content type is switched, the median value of the volume before and after the switch can be calculated and used as the adjusted volume, which can make the volume transition smooth.
[0059] In some embodiments, when the audio content type is switched, the volume can also linearly transition from the first source stream volume to the second source stream volume.
[0060] In this embodiment, by detecting the audio content type and the source stream volume of the original audio stream, the audio content type and the source stream volume are obtained. The audio content type and the source stream volume information are added to the original audio stream by using timestamp alignment and encapsulation insertion technology to form an enhanced audio stream. The target volume of the transcoded audio stream is determined based on the reference volume setting, volume comparison, and volume adjustment strategy during audio content switching, thereby constructing an audio transcoding processing structure to dynamically adjust the volume without interruption during the transcoding process, unify the volume gain in a single and / or multiple audio streams, maintain the consistency of the volume during different scene switches, and improve the user's audio-visual experience.
[0061] As Figure 3 shown, in some embodiments, the audio transcoding system 300 includes an upstream subsystem, a transcoding subsystem, and a downstream subsystem.
[0062] Among them, the upstream subsystem inputs the source file as the source audio stream of the transcoding system.
[0063] The transcoding subsystem includes an AI detection module 302, a big data module 304, and a transcoding module 306. The AI detection module 302 is used to detect the attributes of the source in real time, including sound attributes such as content type, volume, and noise, so as to detect the audio content type and the source stream volume of the original audio stream.
[0064] The big data module 304 is used to record, analyze, and summarize the volume attributes of the source and transcoded streams / files in the transcoding system in real time to determine the reference volume.
[0065] The transcoding module 306 is used to dynamically configure the volume attribute values of the transcoded stream / file in real time according to AI detection and big data statistics information and continuously perform transcoding in real time.
[0066] The downstream subsystem is used to output the transcoded audio stream.
[0067] In an embodiment of the present disclosure, in response to the acquired original audio stream, the audio content type and the source stream volume of the original audio stream are detected, including: dividing the original audio stream based on a specified time window, collecting audio features at the audio sampling points within each time window, inputting the audio features into a feature detection model, and outputting the audio content type, noise detection result, and corresponding timestamp information by the feature detection model, where the feature detection model is generated by relying on deep learning; and inputting the original audio stream into a volume detection model to output the corresponding volume curve as the source stream volume.
[0068] In some embodiments, audio stream data is acquired in real time, divided according to a certain time window. Usually, each window contains several audio sampling points. Features are extracted from the audio data within each time window, and then the features are input into a trained large AI model, that is, the feature detection model. The model outputs the classification result of the current audio content type and the judgment of whether there is noise. For example, the output result may be "human voice, no noise", or "music, with noise", etc. If the model has the ability of multi-class classification, it can further subdivide the types of human voices (such as male voice, female voice, children's voice, etc.) or the types of music (such as pop music, classical music, rock music, etc.), so as to output the classification result and timestamp.
[0069] In some embodiments, a model constructed based on the VAD volume detection formula quantifies the volume of the original audio stream at each time point by calculating the logarithmic relationship between the root mean square amplitude of the audio signal and the reference amplitude, generating a continuous volume curve, and realizing the all-round information detection of the original audio stream.
[0070] In this embodiment, through the time window division and audio feature extraction technology, combined with the deep learning feature detection model and the model constructed by the volume detection formula, a complete audio information detection structure is formed to obtain the audio content type, noise situation, and volume change information of the original audio stream, providing a reliable data basis for subsequent audio processing, improving the automation degree and detection accuracy of audio processing, and meeting the requirements of comprehensive audio information detection in application scenarios such as audio transcoding.
[0071] In one embodiment of the present disclosure, based on the alignment operation with the timestamps of the original audio stream, adding the audio content type and the source stream volume to the original audio stream to obtain an enhanced audio stream includes: taking time windows as units, aligning the audio content type, the noise detection result, and the volume curve to obtain aligned information; encapsulating the aligned information into an SEI data unit based on the format of the supplementary enhancement information SEI; inserting the SEI data unit into the original audio stream based on the corresponding timestamp information to obtain the enhanced audio stream.
[0072] In some embodiments, taking time windows as indexes, matching the audio content type, the noise detection result, and the volume curve corresponding to each window according to timestamps, establishing the corresponding relationship between the information and the time position of the audio stream, encoding and encapsulating the aligned information according to the SEI format specification to form an SEI data unit including header information (for identifying data types, lengths, etc.) and content data. Finally, according to the timestamp information, inserting the SEI data unit accurately into a specific position of the original audio stream (such as between audio frames or in the metadata area), so that on the basis of maintaining the original playback function of the original audio stream, rich auxiliary information is added, providing more comprehensive data support for subsequent processing.
[0073] In this embodiment, through techniques such as time window alignment, SEI format encapsulation, and timestamp-based insertion, a structure for enhancing the information of the original audio stream is constructed. Based on this structure, the audio content type, the noise situation, and the volume information can be added to the original audio stream losslessly, ensuring the time consistency between the information and the audio stream. Without affecting the normal playback of the original audio stream, the data dimension of the audio stream is expanded, providing richer information resources for subsequent operations such as audio transcoding and analysis, and improving the flexibility and functionality of audio processing.
[0074] In one embodiment of the present disclosure, determining the target volume of the transcoded audio stream based on the relationship between the source stream volume and the reference volume includes: determining the reference period corresponding to the transcoding process to determine the reference volume based on the volume attributes of the transcoded audio files belonging to the reference period recorded in the big data module; calculating the absolute difference between the source stream volume and the reference volume and the ratio to the reference volume; if the ratio is less than or equal to the reference ratio, keeping the target volume as the source stream volume; if the ratio is greater than the reference ratio, adjusting the target volume to the reference volume.
[0075] In some embodiments, the appropriate reference period (such as the last hour, the transcoding period of similar audio within a day, etc.) can be determined according to the characteristics and requirements of the transcoding task.
[0076] In some embodiments, volume data (such as statistical values like average volume, maximum volume, minimum volume, etc.) of the transcoded audio file during this period is extracted from the big data module, and a representative reference volume can be determined through statistical analysis or machine learning algorithms.
[0077] In some embodiments, the absolute difference between the source stream volume and the reference volume is calculated and compared with the reference volume. The ratio is compared with a preset reference ratio. When the ratio is within a reasonable range, the source stream volume is maintained as the target volume to preserve the audio source stream volume characteristics. When the ratio exceeds the range, the target volume is adjusted to the reference volume so that the transcoded audio volume conforms to the big data statistical law and generally accepted standards, ensuring that the audio volume is appropriate in different playback scenarios.
[0078] In some embodiments of red date rice cakes, the source stream volume A in the SEI information is parsed in real time, and the overall audio volume value of the current real-time transcoding of the transcoding system obtained from the big data module is compared as the reference volume B, where
[0079] If abs(A, B) / B <= 2%, the target volume remains the source stream volume A.
[0080] If abs(A, B) / B > 2%, the transcoded volume is set to the reference volume B.
[0081] In this embodiment, the reference volume is determined through data analysis of the big data module, and a dynamic adjustment structure for the target volume of the transcoded audio is constructed by combining volume difference ratio calculation and comparison techniques. This structure can automatically adjust the target volume of the transcoded audio based on the relationship between the source stream volume and the big data statistical volume, realizing adaptive optimization of the volume, enabling the volume of the transcoded audio to conform to the statistical law of the overall audio volume, and improving the volume adaptability and auditory experience of the transcoded audio in different playback environments.
[0082] Such as Figure 4 As shown, in an embodiment of the present disclosure, when a switch in the audio content type is detected during transcoding, the second source stream volume after the switch is adjusted based on the first source stream volume before the switch, including:
[0083] Step S402, when a switch in the audio content type is detected, determine the median of the first source stream volume and the second source stream volume.
[0084] Step S404, use the median as the adjustment result of the first source stream volume.
[0085] Step S406, determine the reference volume based on the volume attributes of the transcoded audio files belonging to the reference period recorded in the big data module.
[0086] Step S408, calculate the ratio of the absolute difference between the adjustment result and the reference volume to the reference volume.
[0087] Step S410, if the ratio is less than or equal to the reference ratio, keep the target volume as the adjustment result.
[0088] Step S412, if the ratio is greater than the reference ratio, adjust the target volume to the reference volume.
[0089] In some embodiments, during the transcoding process, the change of the audio content type is monitored in real time. When a content type switch is detected (such as switching from speech to music), the source stream volume values before and after the switch (i.e., the first source stream volume and the second source stream volume) are obtained. The median value is obtained by calculating the average of the two volume values. This median value synthesizes the characteristics of the volume before and after the switch, and this median value is used as the adjustment target for the volume after the switch, so that the volume can achieve a smooth transition at the moment of content switching.
[0090] In this embodiment, through real-time detection of the audio content type, combined with the calculation of the volume median value and the adjustment strategy, a structure for smooth volume transition during audio content switching is constructed. This structure can quickly calculate and apply an appropriate intermediate volume value when the audio content type changes, effectively preventing the problem of sudden volume change, improving the coherence of audio playback and the user's auditory comfort, and providing a reliable solution for the volume processing of different content types during the audio transcoding process.
[0091] In an embodiment of the present disclosure, it further includes: parsing the noise detection result in the SEI data unit; if it is determined based on the noise detection result that there is noise information in the original audio stream, perform a noise reduction function during the transcoding process.
[0092] In some embodiments, during transcoding, first parse the SEI data unit in the enhanced audio stream to extract the noise detection result information therein. Through a preset judgment rule, such as whether it is marked as having noise, a noise intensity threshold, etc., determine whether there is noise in the original audio stream. If noise is detected, call the noise reduction algorithm integrated in the transcoding system, such as a filtering algorithm based on spectrum analysis, a deep learning denoising model, etc., to process the audio stream, and suppress or eliminate the noise component on the premise of retaining the effective signal of the audio, so as to improve the purity and quality of the transcoded audio.
[0093] In one embodiment of the present disclosure, before detecting the audio content type and source stream volume of the original audio stream in response to the acquired original audio stream, it further includes: collecting audio data in different scenarios; annotating the audio data based on the scenario type, content type, and whether there is noise to obtain annotated audio; converting the annotated audio into a Mel spectrogram, where the Mel spectrogram characterizes the time-frequency domain features of the annotated audio; training a deep learning model based on the Mel spectrogram to obtain a feature detection model.
[0094] In some embodiments, specific classification criteria for the scenario type (such as indoor, outdoor, inside a vehicle, etc.), content type (such as human voice, music, animal calls, mechanical sounds, etc.), and whether there is noise (with noise, without noise) are defined to perform annotation based on the above criteria.
[0095] In some embodiments, the audio can be processed by frame segmentation, splitting the continuous audio signal into multiple short frames, performing a fast Fourier transform (FFT) on each frame of audio data to convert it from the time domain to the frequency domain to obtain a spectrum, filtering the spectrum through a Mel filter bank to map the spectrum to the Mel frequency scale to obtain a Mel spectrum, and taking the logarithm of the Mel spectrum to obtain a Mel spectrogram.
[0096] In some embodiments, the deep learning model architecture includes but is not limited to convolutional neural networks (CNNs), recurrent neural networks (RNNs) and their variants (such as LSTMs, GRUs), etc.
[0097] In some embodiments, a validation set is used to evaluate the model during the training process, monitoring the performance metrics of the model (such as accuracy, recall rate, F1 value, etc.). According to the evaluation results, the model is adjusted and optimized, such as adjusting the model architecture, increasing or decreasing the training data, adjusting the training parameters, etc. When the performance of the model on the validation set reaches a good level and no longer improves significantly, it is considered that the model has converged.
[0098] In this embodiment, a training structure for the feature detection model is constructed through multi-source audio data collection, precise annotation, Mel spectrogram conversion, and deep learning training. This structure can utilize a large amount of diverse audio data, through annotation and feature conversion, combined with the learning ability of deep learning, enabling the model to have the ability to identify the audio content type and noise, providing a reliable model basis for the information detection of the audio stream, and improving the intelligent level and accuracy of audio processing.
[0099] In one embodiment of the present disclosure, before detecting the audio content type and source stream volume of the original audio stream in response to the acquired original audio stream, it further includes: constructing a volume detection model based on a volume detection formula, and the volume detection formula is: L p =20*log 10(Prms / Pref) dB, where Prms is the sound amplitude value at any moment in the original audio stream, and Pref is the maximum reference value of the sound amplitude.
[0100] In this embodiment, based on the VAD volume detection formula, through the implementation of the model architecture design and formula calculation logic, a volume detection model structure is constructed. This structure can calculate the volume values of the original audio stream at different moments, generate a continuous volume curve, provide intuitive and quantitative volume information for audio processing, make the audio volume processing process more standardized and precise, and meet the requirements of volume detection for applications such as audio transcoding.
[0101] In an embodiment of the present disclosure, it further includes: reporting the transcoded audio stream including the target volume to the big data module based on the timestamp dimension to update the transcoded audio file in the big data module.
[0102] In this embodiment, through operations such as timestamp addition, data reporting, and big data module update, key information such as the target volume of the transcoded audio stream is timely transmitted back to the big data module, ensuring the continuous update of the audio data in the big data module, providing a data basis for audio analysis, model optimization, and transcoding strategy adjustment based on big data, and forming a closed loop for audio processing from transcoding to data feedback optimization.
[0103] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0104] Those skilled in the art of the relevant technical field can understand that various aspects of the present invention can be implemented as a system, method, or program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0105] Next, refer to Figure 5 to describe the audio transcoding device 500 according to this embodiment of the present invention. Figure 5 The illustrated audio transcoding device 500 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0106] The audio transcoding device 500 is embodied in the form of a hardware module. The components of the audio transcoding device 500 may include, but are not limited to: a detection module 502, configured to detect the audio content type and the source stream volume of the original audio stream in response to the acquired original audio stream; an addition module 504, configured to add the audio content type and the source stream volume to the original audio stream based on an alignment operation with the time stamp of the original audio stream to obtain an enhanced audio stream; a determination module 506, configured to determine the target volume of the transcoded audio stream based on the relationship between the source stream volume and a reference volume, wherein when a switch of the audio content type is detected during the transcoding process, the second source stream volume after the switch is adjusted based on the first source stream volume before the switch, so as to determine the target volume based on the relationship between the adjusted second source stream volume and the reference volume.
[0107] The following will describe the electronic device 600 according to this embodiment of the present invention with reference to Figure 6 ... Figure 6 The illustrated electronic device 600 is merely an example and should not impose any limitation on the functions and the scope of use of the embodiments of the present invention.
[0108] As Figure 6 shown, the electronic device 600 is embodied in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one of the above-mentioned processing units 610, at least one of the above-mentioned storage units 620, and a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610).
[0109] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present invention described in the above "exemplary method" section of this specification. For example, the processing unit 610 may execute the steps S202 and S208 as shown in Figure 1 ..., and other steps defined in the audio transcoding method of the present disclosure.
[0110] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.
[0111] The storage unit 620 may further include a program / utilities 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0112] The bus 630 can represent one or more of several types of bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of the various bus structures.
[0113] The electronic device 600 can also communicate with one or more external devices 660 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 640. Moreover, the electronic device 600 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 650. As shown in the figure, the network adapter 650 communicates with other modules of the electronic device 600 through the bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0114] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0115] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0116] A program product for implementing the above method according to an embodiment of the present invention may be a portable compact disc read-only memory (CD-ROM) and includes program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0117] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0118] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0119] The program code contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0120] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0121] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0122] In addition, although the various steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution, etc.
[0123] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to cause a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0124] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. An audio transcoding method, characterized in that: include: In response to the acquired original audio stream, detecting the audio content type and source stream volume of the original audio stream; Based on an alignment operation with the timestamp of the original audio stream, the audio content type and the source stream volume are added to the original audio stream to obtain an enhanced audio stream; The target volume of the transcoded audio stream is determined based on the relationship between the source stream volume and the reference volume, wherein, when a switch of the audio content type is detected during the transcoding process, the volume of the second source stream after the switch is adjusted based on the volume of the first source stream before the switch, so as to determine the target volume based on the relationship between the adjusted second source stream volume and the reference volume.
2. The audio transcoding method according to claim 1, characterized in that: In response to the acquired original audio stream, detecting the audio content type and source stream volume of the original audio stream includes: Dividing the original audio stream based on a specified time window, collecting audio features at audio sampling points in each of the time windows, and inputting the audio features into a feature detection model, wherein the feature detection model outputs the audio content type, the noise detection result, and the corresponding timestamp information, wherein the feature detection model is generated by deep learning; and The original audio stream is input into a volume detection model, and the corresponding output volume curve is used as the source stream volume.
3. The audio transcoding method according to claim 2, characterized in that: Based on an alignment operation with the timestamp of the original audio stream, the audio content type and the source stream volume are added to the original audio stream to obtain an enhanced audio stream, including: The audio content type, the noise detection result and the volume curve are aligned in units of the time window to obtain alignment information; Based on the format of supplemental enhancement information SEI, encapsulating the aligned information into SEI data units; Based on the corresponding timestamp information, the SEI data unit is inserted into the original audio stream to obtain the enhanced audio stream.
4. The audio transcoding method according to claim 1, characterized in that: Determining a target volume of the transcoded audio stream based on a relationship between the source stream volume and the reference volume includes: Determine a reference time period corresponding to the transcoding process, so as to determine the reference volume based on the volume attribute of the transcoded audio file belonging to the reference time period recorded in the big data module; Calculate the absolute difference between the source stream volume and the reference volume, and the ratio between the absolute difference and the reference volume; If the ratio is less than or equal to the reference ratio, maintaining the target volume as the source stream volume; If the ratio is greater than the reference ratio, the target volume is adjusted to the reference volume.
5. The audio transcoding method according to claim 2, characterized in that: When the switching of the audio content type is detected during the transcoding process, adjusting the volume of the second source stream after the switching based on the volume of the first source stream before the switching includes: When the switching of the audio content type occurs, determining a median value of the first source stream volume and the second source stream volume; The median value is used as the adjustment result of the volume of the first source stream.
6. The audio transcoding method according to claim 3, characterized in that: Also includes: parsing the noise detection result in the SEI data unit; If it is determined based on the noise detection result that the original audio stream contains noise information, a noise removal function is performed during the transcoding process.
7. The audio transcoding method according to claim 2, characterized in that: Before detecting the audio content type and source stream volume of the original audio stream in response to the acquired original audio stream, the method further includes: Collect audio data in different scenarios; Based on the scene type, the content type, and whether there is noise, the audio data is labeled to obtain labeled audio; Converting the annotated audio into a Mel-spectrogram, wherein the Mel-spectrogram represents the time-frequency domain features of the annotated audio; The deep learning model is trained based on the mel-spectrogram to obtain the feature detection model.
8. The audio transcoding method according to claim 2, characterized in that: Before detecting the audio content type and source stream volume of the original audio stream in response to the acquired original audio stream, the method further includes: The volume detection model is constructed based on the volume detection formula, and the volume detection formula is: p =20*log 10 (Prms / Pref) dB, where Prms is the sound amplitude value at any time in the original audio stream, and Pref is the maximum reference value of the sound amplitude.
9. The audio transcoding method according to claim 1, characterized in that: Also includes: The transcoded audio stream including the target volume is reported to the big data module based on the timestamp dimension to update the transcoded audio file in the big data module.
10. An audio transcoding device, characterized in that: include: A detection module, configured to detect the audio content type and source stream volume of the original audio stream in response to the acquired original audio stream; An adding module, configured to add the audio content type and the source stream volume to the original audio stream based on an alignment operation with the timestamp of the original audio stream to obtain an enhanced audio stream; A determination module is used to determine a target volume of a transcoded audio stream based on the relationship between the source stream volume and a reference volume, wherein when a switch of the audio content type is detected during the transcoding process, the volume of the second source stream after the switch is adjusted based on the volume of the first source stream before the switch, so as to determine the target volume based on the relationship between the adjusted second source stream volume and the reference volume.
11. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the audio transcoding method according to any one of claims 1 to 9 by executing the executable instructions.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the audio transcoding method according to any one of claims 1 to 9 is implemented.
13. A computer program product having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the audio transcoding method according to any one of claims 1 to 9 is implemented.