Audio Processing Method, Apparatus, Device, and Storage Medium
By determining the decoding start frame identifier and end frame identifier in the preset frame sequence, the on-demand decoding of audio files is achieved, and the problem of time-consuming decoding of large or long-term audio files in the prior art is solved, and the efficiency and performance of audio processing are improved.
Patent Information
- Application Number
- CN202111654061.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-12-30
AI Technical Summary
The prior art is complex and time-consuming when processing large or long audio files, resulting in a decoding process that degrades the audio processing performance, which may cause browser crashes and affect machine performance.
By determining the decoding start frame identification and the decoding end frame identification in the preset frame sequence, the fragment data to be decoded in the audio resource is accurately obtained and decoded, so as to realize on-demand decoding and avoid full decoding.
This method improves the efficiency and flexibility of audio processing, reduces memory usage, avoids browser crashes and performance degradation, and ensures timeliness of audio processing.
Smart Images

Figure CN114299972B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of audio processing, and in particular, to an audio processing method, apparatus, device, and storage medium. Background Art
[0002] With the development of audio technology, there are more and more application scenarios involving processing such as playing or editing audio. In these application scenarios, the sources of audio files to be processed can be very diverse, but in the processing process, the audio files usually need to be decoded.
[0003] Currently, when the audio file is large or has a long duration, the decoding process is complex and time-consuming, affecting the audio processing performance. Taking an application scenario involving the front end of a web page (Web) as an example, when the front end of a web page processes an audio file, it usually decodes the entire audio file. When the audio file is large or has a long duration, it is very easy to occupy a large amount of memory during the decoding process, resulting in the browser crashing. Moreover, operating a large amount of memory will seriously affect the machine performance. At the same time, the full decoding process is time-consuming and it is difficult to ensure the timeliness of audio processing. Summary of the Invention
[0004] Embodiments of the present disclosure provide an audio processing method, apparatus, storage medium, and device, which can optimize the existing audio processing solution.
[0005] In a first aspect, embodiments of the present disclosure provide an audio processing method, including:
[0006] Determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence, where the preset frame sequence includes frame information of each audio frame in at least one audio resource, the frame information includes a frame identifier, the frame identifier includes an audio resource identifier and a frame index, the audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs;
[0007] Obtain the data of the to-be-decoded segment in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier;
[0008] Decode the to-be-decoded segment data to obtain corresponding target decoded data.
[0009] In a second aspect, embodiments of the present disclosure provide an audio processing apparatus, including:
[0010] A frame identifier determination module, configured to determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence, where the preset frame sequence includes frame information of each audio frame in at least one audio resource, the frame information includes a frame identifier, and the frame identifier includes an audio resource identifier and a frame index. The audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs;
[0011] A data to be decoded acquisition module, configured to acquire the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier;
[0012] A decoding module, configured to decode the data of the segment to be decoded to obtain corresponding target decoded data.
[0013] In a third aspect, an embodiment of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the audio processing method provided by the embodiment of the present disclosure is implemented.
[0014] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the audio processing method provided by the embodiment of the present disclosure is implemented.
[0015] In the audio processing solution provided by the embodiment of the present disclosure, a decoding start frame identifier and a decoding end frame identifier are determined in a preset frame sequence. The preset frame sequence includes frame information of each audio frame in at least one audio resource. The frame information includes a frame identifier, and the frame identifier includes an audio resource identifier and a frame index. The audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs. According to the decoding start frame identifier and the decoding end frame identifier, the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier is acquired, and the data of the segment to be decoded is decoded to obtain corresponding target decoded data. By adopting the above technical solution, the frame information of each audio frame in the audio resource is stored in a sequence form in advance. When decoding is required, the data range to be decoded is accurately located according to the decoding start frame identifier and the decoding end frame identifier, and the segment data is acquired from the corresponding audio resource and decoded. There is no need to decode the entire audio file, and on-demand decoding is realized, making the decoding more flexible and improving the audio processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic flowchart of an audio processing method provided by an embodiment of the present disclosure;
[0017] Figure 2 Schematic diagram of an audio processing solution provided by an embodiment of the present disclosure;
[0018] Figure 3 Flowchart of another audio processing method provided by an embodiment of the present disclosure;
[0019] Figure 4 Schematic diagram of an audio playback control process provided by an embodiment of the present disclosure;
[0020] Figure 5 Block diagram of an audio processing device provided by an embodiment of the present disclosure;
[0021] Figure 6 Block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0022] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0023] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0024] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0025] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0026] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0027] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.
[0028] In each of the following embodiments, optional features and examples are provided simultaneously. The features described in the embodiments can be combined to form multiple alternative solutions. Each numbered embodiment should not be regarded as only one technical solution.
[0029] Figure 1 It is a schematic flowchart of an audio processing method provided by an embodiment of the present disclosure. This method can be executed by an audio processing device and is applicable to the application scenario of decoding audio. The device can be implemented by software and / or hardware and is generally integrated in an electronic device. The electronic device can be a mobile device such as a mobile phone, a smart watch, a tablet computer, and a personal digital assistant; it can also be other devices such as a desktop computer. As Figure 1 shown, the method includes:
[0030] Step 101: Determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence. Among them, the preset frame sequence contains frame information of each audio frame in at least one audio resource. The frame information contains a frame identifier, and the frame identifier includes an audio resource identifier and a frame index. The audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs.
[0031] In the embodiments of the present disclosure, an audio resource can be understood as an original audio file, and its specific source is not limited. It can be an audio file stored locally on an electronic device, an audio file stored on a server (such as in the cloud), or an audio file from other sources. The audio resource stored on the server can be an audio file uploaded by a user to the server, or an audio file obtained by converting (such as format conversion, etc.) the audio file uploaded by the user. The audio resource is associated with an audio resource identifier, which is used to represent the identity of the audio resource and can be denoted as a resource identity identifier (Identification, ID).
[0032] Generally, an audio file is composed of a series of encoded audio frames. An audio frame can be understood as the smallest unit of an audio segment that can be independently decoded. The frame structures of audio frames in different formats of audio files may be different. Based on acoustic principles, the duration of each frame is generally between 20 ms (milliseconds) and 50 ms. Information related to the audio frame can be maintained in the audio frame, such as the resource ID associated with the audio frame (i.e., the audio resource identifier associated with the audio resource to which the audio frame belongs), the order of the audio frame among all the audio frames of the audio resource to which it belongs, the position of the audio frame in the audio resource to which it belongs, the data volume size of the audio frame, and the meta-information of the audio resource to which the audio frame belongs, etc.
[0033] In the embodiments of the present disclosure, before this step, the audio resource may be frame-divided to obtain the frame information of each audio frame in the audio resource, and a preset frame sequence is constructed according to the frame information. Among them, the frame division process can be understood as determining the corresponding frame information for each audio frame in the audio resource respectively. Exemplarily, the required information can be obtained in advance from the information maintained by each audio frame in one or more audio resources, and the corresponding frame information of each audio frame can be obtained by means of direct extraction and / or secondary calculation, etc. Among them, the frame information may include a frame identifier or other information, and no specific limitation is made. The frame index is used to represent the order of the corresponding audio frame among all the audio frames of the audio resource to which it belongs. For example, the frame index of the first audio frame in the audio resource can be recorded as 0, the frame index of the second audio frame can be recorded as 1, and so on. After obtaining the frame information of each audio frame, the frame information can be arranged in a preset order to obtain a preset frame sequence, that is, the objects in the preset frame sequence are sorted in units of frame information. Among them, the preset order can be set according to actual needs, and no specific limitation is made, and it can also be dynamically adjusted according to actual needs during the application process. Optionally, the preset order can be sorted according to the audio resource identifier, that is, the frame information associated with the same audio resource identifier is arranged together. For the frame information associated with the same audio resource identifier, it can be sorted in order according to the frame index, that is, the sorting of the frame information is consistent with the original order of each audio frame in the audio resource to which it belongs, or it can be sorted in other orders. For example, there may be other frame information arranged between the frame information of adjacent frame indices. For example, there may be frame information with a frame index of 3 between the frame information with a frame index of 1 and the frame information with a frame index of 2, etc. Optionally, the preset order can be an interleaved sorting of the frame information corresponding to different audio resources. For example, there may be frame information with a resource ID of 2 between the two frame information with a resource ID of 1, etc.
[0034] In the embodiments of the present disclosure, the decoding start frame identifier can be understood as the frame identifier in the frame information corresponding to the first audio frame to be decoded this time, and the decoding end frame identifier can be understood as the frame identifier in the frame information corresponding to the last audio frame to be decoded this time. Exemplarily, when it is detected that a preset decoding event is triggered, the decoding start frame identifier and the decoding end frame identifier can be determined in a preset frame sequence. The triggering condition of the preset decoding event is not limited and can be set according to actual decoding requirements. The decoding requirements can include, for example, playing requirements, decoding data buffering requirements, audio-to-text requirements, and audio waveform drawing requirements, etc. The decoding requirements can be automatically determined according to the current usage scenario or determined according to the operations input by the user. The preset decoding event can indicate the requirement parameters of the current decoding requirements. The requirement parameters can include, for example, the decoding start frame identifier, the decoding end frame identifier, or the target decoding duration, etc. Optionally, the decoding start frame identifier and the decoding end frame identifier are determined in the preset frame sequence according to the requirement parameters. For example, when the requirement parameters include the decoding start frame identifier and the decoding end frame identifier, the decoding start frame identifier and the decoding end frame identifier can be directly found in the preset frame sequence according to the decoding start frame identifier and the decoding end frame identifier; another example is that when the requirement parameters include the decoding start frame identifier and the target decoding duration, the decoding start frame identifier can be first found in the preset frame sequence according to the decoding start frame identifier, and the durations of the audio frames corresponding to the frame information behind in the preset frame sequence are sequentially accumulated starting from the audio frame corresponding to the decoding start frame identifier until the target decoding duration is reached, and the decoding end frame identifier is determined according to the frame identifier in the frame information at this time.
[0035] Step 102: Obtain the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier.
[0036] Exemplarily, the corresponding audio resource identifier can be understood as the audio resource identifier included in the decoding start frame identifier and / or the decoding end frame identifier.
[0037] Exemplarily, assume that the frame information to which the decoding start frame identifier belongs can be recorded as the start frame information, and the frame information to which the decoding end frame identifier belongs can be recorded as the end frame information. If the audio resource identifiers corresponding to the start frame information, the end frame information, and the frame information between the start frame information and the end frame information (which can be recorded as the intermediate frame information) in the preset frame sequence are all the same, it means that the audio frames to be decoded come from the same audio resource. The corresponding-order audio frames in the audio resource can be obtained according to the frame indexes included in the frame identifiers (which can be recorded as the intermediate frame identifiers) included in the decoding start frame identifier, the decoding end frame identifier, and the intermediate frame information, and the data of the segment to be decoded is obtained.
[0038] Exemplarily, if there are at least two different audio resource identifiers among the audio resource identifiers corresponding to the start frame information, end frame information, and intermediate frame information, it indicates that the audio frames to be decoded come from at least two audio resources. The audio frames in the corresponding order can be obtained from the audio resources associated with the corresponding audio resource identifiers respectively according to the frame indexes included in the decoding start frame identifier, decoding end frame identifier, and intermediate frame identifier, so as to obtain the data of the segment to be decoded.
[0039] Step 103: Decode the data of the segment to be decoded to obtain the corresponding target decoded data.
[0040] Exemplarily, after obtaining the data of the segment to be decoded, a preset decoding algorithm can be used or a preset decoding interface can be called to decode the data of the segment to be decoded, and the target decoded data required for this decoding can be determined according to the decoding result.
[0041] In the audio processing method provided in the embodiments of the present disclosure, a decoding start frame identifier and a decoding end frame identifier are determined in a preset frame sequence. The preset frame sequence includes the frame information of each audio frame in at least one audio resource, the frame information includes a frame identifier, and the frame identifier includes an audio resource identifier and a frame index. The audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all the audio frames in the audio resource to which it belongs. According to the decoding start frame identifier and the decoding end frame identifier, the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier is obtained, and the data of the segment to be decoded is decoded to obtain the corresponding target decoded data. By adopting the above technical solution, the frame information of each audio frame in the audio resource is stored in sequence in advance. When decoding is required, the data range to be decoded is accurately located according to the decoding start frame identifier and the decoding end frame identifier, the segment data is obtained from the corresponding audio resource and decoded, and there is no need to decode the entire audio file, realizing on-demand decoding, making the decoding more flexible, and improving the audio processing efficiency.
[0042] In some embodiments, determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence includes: determining the target decoding duration and the decoding start frame identifier; starting traversing from the frame information corresponding to the decoding start frame identifier in the preset frame sequence, and when the preset traversal termination condition is satisfied, determining the decoding end frame identifier according to the corresponding frame information; wherein, the preset traversal termination condition includes: the cumulative duration of the audio frames corresponding to the traversed frame information reaches the target decoding duration. The advantage of this setting is that by traversing frame by frame, the duration of the audio frames corresponding to the traversed frame information is accumulated. When the target decoding duration is reached, the traversal ends, and the decoding of the audio frame data at the specified start position and the specified duration can be achieved according to the decoding start frame identifier and the target decoding duration. Among them, the duration of each audio frame is generally related to the sampling rate of the audio resource to which it belongs. The corresponding sampling rate can be obtained according to the audio resource identifier in the currently traversed frame information, and then the duration of the audio frame corresponding to the current frame information can be determined.
[0043] Exemplarily, assuming that after accumulating the duration of the audio frame corresponding to the current frame information, the obtained cumulative duration is greater than or equal to the target decoding duration, the frame identifier in the current frame information can be determined as the decoding end frame identifier.
[0044] In some embodiments, the preset traversal termination condition further includes at least one of the following: the audio resource identifier in the current frame information is inconsistent with the audio resource identifier in the previous frame information; the frame index in the current frame information is not continuous with the frame index in the previous frame information; the frame index in the current frame information is the last one in the audio resource to which it belongs. Among them, when the preset traversal termination condition is satisfied, determining the decoding end frame identifier according to the corresponding frame information includes: when any one of the preset traversal termination conditions is satisfied, determining the decoding end frame identifier according to the corresponding frame information. The advantage of this setting is that by enriching the items in the preset traversal termination condition, the data of the segment to be decoded can come from the same audio resource, and the audio frames in the data of the segment to be decoded can also be continuous. When any one is satisfied, the traversal is terminated, ensuring that the data of the segment to be decoded is obtained from the same audio resource each time, reducing the difficulty of obtaining the data of the segment to be decoded, and improving the data acquisition efficiency.
[0045] Exemplarily, assuming that the audio resource identifier in the current frame information is inconsistent with the audio resource identifier in the previous frame information, the frame identifier in the previous frame information can be determined as the decoding end frame identifier; assuming that the frame index in the current frame information is not continuous with the frame index in the previous frame information, the frame identifier in the previous frame information can be determined as the decoding end frame identifier; assuming that the frame index in the current frame information is the last one in the audio resource to which it belongs, the frame identifier in the current frame information can be determined as the decoding end frame identifier.
[0046] In some embodiments, the frame information further includes a frame offset and a frame data amount; the obtaining of the data segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier includes: determining the audio resource associated with the audio resource identifier corresponding to the decoding start frame identifier as the target audio resource; determining the data start position according to the first frame offset corresponding to the decoding start frame identifier, determining the data end position according to the second frame offset and the frame data amount corresponding to the decoding end frame identifier, determining the target data range according to the data start position and the data end position; and obtaining the audio data within the target data range in the target audio resource to obtain the data segment to be decoded. The advantage of such a setting is that the data segment to be decoded can be obtained more quickly and accurately.
[0047] Exemplarily, the frame offset can be understood as the starting position of the audio frame in the audio resource to which it belongs, and the unit can be bytes. The frame data amount can be understood as the size of the audio frame in the audio resource to which it belongs, and the unit is generally the same as the frame offset and can be bytes. The frame offset corresponding to the frame identifier can be understood as the frame offset included in the frame information where the frame identifier is located, that is, the corresponding frame identifier and frame offset are in the same frame information, and the same applies to the frame data amount. When the preset traversal termination condition includes the above four items at the same time, it can be ensured that the decoding start frame identifier and the decoding end frame identifier correspond to the same audio resource, and the corresponding audio resource identifier can be determined according to any one of them, and then the associated audio resource is determined as the target audio resource. According to the first frame offset, the start position of the data to be obtained in the target audio resource can be determined, and according to the second frame offset and the frame data amount, the end position of the data to be obtained in the target audio resource can be determined (for example, the end position can be expressed as the second frame offset + frame data amount - 1), so as to obtain the target data range, and the corresponding audio data can be extracted from the target audio resource according to this target data range.
[0048] In some embodiments, traversing starting from the frame information corresponding to the decoding start frame identifier in the preset frame sequence includes: determining the format of the audio frame corresponding to the decoding start frame identifier; in the case where the format is a preset format, traversing starting from the frame information corresponding to the target frame index in the preset frame sequence, where the target frame index is obtained by tracing back a preset frame index difference forward based on the start frame index in the decoding start frame identifier. Among them, obtaining the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier includes: obtaining the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the target frame identifier corresponding to the target frame index and the decoding end frame identifier. The advantage of such a setting is that for some formats of audio resources, the audio frames may not be completely independent, and a certain number of audio frames (which can be called preframes) can be traced back forward and added to the data of the segment to be decoded to ensure the integrity and accuracy of the decoded data. The preset frame index difference can be set according to the preset format. Exemplarily, the preset format may include the Moving Picture Experts Group Audio Layer III (MP3) format, and the corresponding preset frame index difference can be 1. It should be noted that for some special cases, such as the decoding start frame identifier being 0, it means that the first audio frame in the target audio resource needs to be decoded. At this time, the decoding start frame identifier can be regarded as the target frame identifier.
[0049] In some embodiments, in the case where the format is a preset format, decoding the data of the segment to be decoded to obtain the corresponding target decoded data includes: decoding the data of the segment to be decoded to obtain the corresponding initial decoded data; removing the redundant decoded data from the initial decoded data to obtain the corresponding target decoded data, where the redundant decoded data includes the decoded data of the audio frames corresponding to the frame indexes before the start frame index. The advantage of such a setting is that for the preset format, the data of the segment to be decoded determined in the above steps contains the data of the redundant preframes. Therefore, the initial decoded data obtained by decoding also contains the decoded data of the preframes. In order to avoid the repeated use of the decoded data, such as repeated playback, etc., the decoded data of the preframes can be removed.
[0050] In some embodiments, after decoding the data of the segment to be decoded to obtain the corresponding target decoded data, the method further includes: recording the decoding end frame identifier and the decoding duration corresponding to the target decoded data. The advantage of this setting is that after setting the above preset traversal termination condition, there may be a situation where the actual decoding duration is not equal to the target decoding duration. Recording the current decoding position and the actual decoding duration in a timely manner facilitates subsequent continuous decoding based on this.
[0051] In an application scenario related to the web front - end, when the web front - end processes an audio file, it usually performs full - volume decoding on the audio file. When the audio file is large or has a long duration (such as dozens of minutes or even more than 1 hour), it is very easy to occupy a large amount of memory during the decoding process, resulting in the browser crashing. Moreover, operating a large amount of memory will seriously affect the machine performance. At the same time, the full - volume decoding process is time - consuming and it is difficult to ensure the timeliness of audio processing. The audio processing solution in the embodiments of the present invention can be applied to the application scenario of the web front - end.
[0052] In some embodiments, when the method is applied to the web front - end, before determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence, the method further includes: performing frame splitting on the audio resource to obtain the frame information of each audio frame in the audio resource; storing the obtained frame information in the preset frame sequence of the web front - end. The advantage of this setting is that maintaining the preset frame sequence in the web front - end does not require storing full - volume decoded data, reducing the large occupation of memory and improving the performance of the browser and the device. Among them, the audio resource for which frame splitting is performed may include all or part of the audio resources involved in the current session.
[0053] In some embodiments, before determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence, the method further includes: obtaining the meta - information of the audio resource, where the meta - information includes the storage information of the audio resource, and the storage information includes the storage location and / or resource data of the audio resource; storing the meta - information in the resource table of the web front - end, where the resource table includes the association relationship between the audio resource identifier involved in the current session and the storage information. Correspondingly, the step of obtaining the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier includes: obtaining the target storage information associated with the corresponding audio resource identifier from the resource table according to the decoding start frame identifier and the decoding end frame identifier, and obtaining the data of the segment to be decoded based on the target storage information. The advantage of this setting is that the storage information corresponding to the audio resource can be stored in the front - end in the form of a storage resource table, which is convenient for quickly obtaining the data of the segment to be decoded through the resource table.
[0054] Exemplarily, the meta information may include global information of the audio resource, and may include storage information of the audio resource. Among them, the storage information includes the storage location and / or resource data of the audio resource. The storage location may include a Uniform Resource Locator (URL) address or a local storage path, etc. The resource data may be understood as the complete data of the audio resource. Generally, in order to save storage resources, the storage location and the resource data may exist alternatively. In addition, the meta information may further include, for example, the format of the audio resource (which may be an enumerated type), the total file size of the audio resource (the unit may be bytes), the total duration of the audio resource (the unit may be seconds), the sampling rate of the audio resource (the unit may be Hertz), the number of channels of the audio file, and other information (such as custom information), etc.
[0055] In some embodiments, it may further include: receiving a preset audio editing operation; performing corresponding editing operations on the corresponding frame information in the preset frame sequence according to the frame identifier to be adjusted indicated by the preset audio editing operation, so as to implement audio editing, where the editing operation includes deleting and / or adjusting the order of the frame information. The advantage of such a setting is that by editing the frame information in the frame sequence, audio editing at the audio frame granularity is realized, and there is no need to operate on the original resource data, which can greatly improve the audio editing efficiency and accuracy.
[0056] Exemplarily, the preset audio editing operation may include insertion, deletion, and sorting, etc. The number of frame identifiers to be adjusted indicated by different preset audio editing operations may be different. When the number of frame identifiers to be adjusted is multiple, the included audio resource identifiers may be the same or different. For example, for insertion, the frame identifier to be adjusted may include the frame identifier of the audio frame to be inserted (which may be denoted as the first frame identifier, and the number may be one or more), and may further include the frame identifier of the audio frame for indicating the insertion position (which may be denoted as the second frame identifier). For example, the frame information corresponding to the first frame identifier is inserted after the frame information corresponding to the second frame identifier. Another example is that for deletion, the frame identifier to be adjusted may include the frame identifier of the audio frame to be deleted. Still another example is that for sorting, the frame identifier to be adjusted may include multiple frame identifiers of the audio frames to be sorted (which may be denoted as the third frame identifier), and the preset audio editing operation may further indicate the target sorting, and re-sort the frame information corresponding to the multiple third frame identifiers according to the target sorting to achieve more accurate audio editing.
[0057] In some embodiments, the preset frame sequence further includes waveform summary information corresponding to each audio frame; the method further includes: in response to receiving a preset waveform drawing instruction, obtaining target waveform summary information corresponding to the frame information corresponding to the frame identifier to be drawn indicated by the preset waveform drawing instruction; and drawing a corresponding waveform diagram according to the target waveform summary information. The advantage of such a setting is that by storing the waveform summary information corresponding to each audio frame in the preset frame sequence, when it is necessary to draw a waveform diagram, there is no need to decode the audio data, and the waveform diagram can be directly drawn according to the waveform summary information of the audio frame to be drawn, which can effectively improve the drawing efficiency of the waveform diagram.
[0058] Exemplarily, the waveform summary information may include a plurality of amplitude values, and may further include the time interval between every two adjacent amplitude values. The plurality of amplitude values may be uniformly distributed or non-uniformly distributed in the time dimension, and no specific limitation is made. The preset waveform drawing instruction may be automatically generated according to the current scenario, or may be generated according to an operation input by the user, etc.
[0059] In some embodiments, before obtaining the target waveform summary information corresponding to the frame information corresponding to the preset frame sequence, it further includes: decoding the audio resource corresponding to the preset frame sequence; for the decoded frame data of each audio frame, dividing the current decoded frame data into sub-interval data of a first preset number, determining the interval amplitudes corresponding to the respective sub-interval data, and determining the waveform summary information corresponding to the current audio frame according to the interval amplitudes; storing the waveform summary information corresponding to each audio frame in the preset frame sequence, and establishing an association with the corresponding frame information. The advantage of such a setting is that the decoded frame data is divided into intervals, and the interval amplitudes are determined in units of sub-intervals, so that the waveform summary information corresponding to each audio frame can be quickly and accurately obtained and stored in the preset frame sequence, facilitating subsequent waveform diagram drawing.
[0060] Exemplarily, when performing interval division, it can be carried out in an equal-interval manner, that is, the size of each interval can be the same, ensuring the uniformity of the interval amplitude distribution and being able to more accurately reflect the variation law of the audio signal. Optionally, the first preset quantity can be determined according to the duration of the audio frame and the preset amplitude interval. For example, the preset amplitude interval can be used to represent calculating the amplitude every preset duration. For example, the amplitude is calculated every 20 ms, and the duration of the audio frame is 40 ms, then the first preset quantity can be 2, and the current decoded frame data is divided into 2 sub-interval data. When determining the interval amplitude corresponding to the sub-interval data, the maximum amplitude in the sub-interval data can be determined as the interval amplitude. After obtaining all the interval amplitudes corresponding to an audio frame, the interval amplitudes can be summarized in the order of the corresponding sub-interval data to form the waveform summary information corresponding to the audio frame, and stored in the position of the frame information corresponding to the audio frame, or added to the frame information, so as to establish an association with the corresponding frame information.
[0061] Exemplarily, when decoding the audio resource corresponding to the preset frame sequence, a target duration (which can be the above-mentioned target decoding duration) can be set, and the audio resource is decoded in batches with the target duration as the unit, that is, the waveform summary information of the audio frames within this batch is determined batch by batch. After the single decoding is completed and the waveform summary information has been determined, the decoded data can be deleted to reduce the occupation of storage resources.
[0062] In some embodiments, it further includes: dividing the preset frame sequence into a second preset quantity of sub-sequences; for each sub-sequence, performing partial decoding on the current sub-sequence, and determining the sub-sequence amplitude corresponding to the current sub-sequence according to the decoding result; drawing a waveform sketch according to the sub-sequence amplitudes corresponding to each sub-sequence respectively. The advantage of such a setting is that through the partial decoding method, partial amplitude information can be quickly and selectively obtained, so as to timely obtain the overall variation law of the audio signal.
[0063] Exemplarily, when performing sub-sequence division, it can be carried out in an equal-interval manner, that is, the size of each sub-sequence can be the same, ensuring the uniformity of the sub-sequence amplitude distribution and being able to more accurately reflect the overall variation law of the audio signal. The second preset quantity can be set according to actual needs, that is, the number of amplitudes to be output. For example, if it is desired to output the amplitude values at the preset numerical equal division points of the entire preset frame sequence, then the second preset quantity can be equal to the preset value.
[0064] In some embodiments, partially decoding the current subsequence and determining the subsequence amplitude corresponding to the current subsequence according to the decoding result includes: dividing the current subsequence into a third preset number of decoding units; for each decoding unit, obtaining data to be decoded according to the starting frame index and the preset number of decoded frames corresponding to the current decoding unit, and after decoding the data to be decoded, determining the maximum amplitude in the obtained decoded data as the unit amplitude of the current decoding unit; and determining the maximum unit amplitude among the unit amplitudes as the subsequence amplitude corresponding to the current subsequence. The advantage of this setting is that when partially decoding a subsequence, further division is performed, and partial decoding within the decoding unit is carried out with the decoding unit as the unit, so that the data subjected to partial decoding is more evenly distributed, and the overall variation law of the audio signal is more accurately reflected.
[0065] Exemplarily, the maximum number and the minimum number of audio frames included in a single decoding unit can be preset in advance, and the number of audio frames included in a single decoding unit is estimated according to the total number of frames of the current subsequence, so that the number of audio frames is between the maximum number and the minimum number, and then the third preset number is determined according to the total number of frames and the number of audio frames.
[0066] In some embodiments, the target decoded data is used to be stored in a playback buffer, and the method further includes: determining whether to determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence according to the amount of undecoded data in the playback buffer. The advantage of this setting is that for an audio playback scenario, for an on-demand decoding audio decoding method, a playback buffer is set, and a fill-in playback mode can be realized. Whether more audio data needs to be decoded is dynamically determined according to the amount of remaining undecoded data in the buffer, ensuring the smoothness of playback.
[0067] Exemplarily, if the amount of undecoded data is less than a preset data amount threshold, a decoding start frame identifier and a decoding end frame identifier are determined in the preset frame sequence. Among them, the decoding start frame identifier can be determined according to the decoding end frame identifier recorded after the last decoding is completed. For example, the frame identifier of the next frame information of the frame information to which the decoding end frame identifier in the preset frame sequence belongs is determined as the current decoding start frame identifier.
[0068] In some embodiments, when applied to the web front end, it further includes: synchronizing the preset frame sequence corresponding to the current session to the server. The advantage of this setting is that the preset frame sequence can be synchronized to the server to ensure that the preset frame sequence is not lost in cases such as web page refreshing.
[0069] Exemplarily, other relevant data of the current session, such as a resource table, can also be synchronized to the server. For the data that needs to be synchronized to the server, it can be determined whether compression is required according to the data volume. Generally, when the duration of the audio resource is long or the number of audio resources is large, etc., the data volume of the preset frame sequence may be relatively large. At this time, the preset frame sequence can be compressed before synchronization.
[0070] Taking the web front-end application scenario as an example below, the embodiments of the present disclosure will be further described. Figure 2 It is a schematic diagram of the architecture of an audio processing solution provided by an embodiment of the present disclosure. This architecture mainly includes the cloud, a Software Development Kit (SDK), and a Web container. The audio processing method provided by the embodiments of the present disclosure can be implemented through the SDK. The SDK can be understood as the packaging and interface exposure of the functions implemented by the audio processing method. The SDK may include a frame splitter, a decoder, a player, a waveform drawer, a serializer, a compressor, etc. Among them, the frame splitter is responsible for analyzing the source file and extracting meta information and frame information; the decoder encapsulates the function of decoding audio by segments and serves the player and the waveform drawer; the player encapsulates the capabilities of real-time loading, decoding, and playing audio; the waveform drawer is responsible for drawing waveforms and constructing waveform summaries of frames; the serializer is responsible for serializing frame information into binary data and corresponding deserialization for convenient persistent storage; the compressor is responsible for compressing and decompressing the serialized data; the Web container contains the data that needs to be maintained when the web front-end applies the SDK, which may include a resource table, a frame sequence (i.e., the preset frame sequence), and a waveform (waveform diagram).
[0071] Figure 3 It is a schematic flowchart of another audio processing method provided by an embodiment of the present disclosure. The embodiments of the present disclosure are optimized based on the above-mentioned optional solutions in the above embodiments and can be combined with Figure 3 for understanding.
[0072] Specifically, the method includes the following steps:
[0073] Step 301: Perform frame splitting on the audio resource to obtain the frame information of each audio frame in the audio resource and the meta information of the audio resource, store the obtained frame information in the preset frame sequence of the web front-end, and store the meta information in the resource table of the web front-end.
[0074] Exemplarily, a framer can be used to perform frame processing on the source file corresponding to the audio resource. For audio files in different formats, the frame processing methods may be different. Before frame processing, the format of the audio resource can be analyzed first, and then the corresponding frame processing method can be matched, that is, the framer of the corresponding format is used for frame processing. For example, the estimated file format can be determined first according to the file name suffix, and the source file can be detected to determine whether the source file is the estimated file format (that is, to determine whether the file name suffix matches the actual format). If so, the framer corresponding to the estimated file format is selected. If it is not the estimated file format, the preset file format set can be traversed, and the format that matches the source file is determined as the target format, and then the framer corresponding to the target format is selected. The preset file format set can include all audio file formats supported by the embodiments of the present disclosure, such as MP3, MP4, Windows Wave (WAV), and Advanced Audio Coding (AAC), etc., which are not specifically limited.
[0075] After frame processing, the meta-information of the obtained audio resource may include the audio format enumeration type (type), the total size (size) of the audio file, the total duration (duration) of the audio, the storage address (url) of the audio file, the complete data (data, generally not present at the same time as url) of the audio file, the sampling rate (sampleRate) of the audio file, and the number of channels (channelCount) of the audio file, etc. The frame information may include the resource id (uri) associated with the frame, the original order (index, generally starting from 0) of the frame in all frames of the original audio file, the starting position (offset) of the frame in the original audio file, the size (size) of the frame in the original audio file, and the number of sampling points stored per channel (sampleSize) of the frame, etc. A preset frame sequence is constructed according to the above frame information, and the frame identifier includes uri and index. The preset frame sequence may also include waveform summary information (wave), which is subsequently constructed by the waveform drawer. When constructing the preset frame sequence, the storage space for wave can be reserved, and it will be filled after the waveform drawer obtains the waveform summary information.
[0076] Step 302: Divide the preset frame sequence into multiple subsequences, respectively determine the sub-sequence amplitudes corresponding to each sub-sequence, and draw a waveform sketch according to the sub-sequence amplitudes corresponding to each sub-sequence.
[0077] Specifically, divide the preset frame sequence into a second preset number of subsequences. For each subsequence, divide the current subsequence into a third preset number of decoding units. For each decoding unit, obtain the data to be decoded according to the start frame identifier corresponding to the current decoding unit and the preset number of decoding frames. After decoding the data to be decoded, determine the maximum amplitude in the obtained decoded data as the unit amplitude of the current decoding unit. Determine the maximum unit amplitude among all unit amplitudes as the subsequence amplitude corresponding to the current subsequence. Draw a waveform sketch according to the subsequence amplitudes corresponding to each subsequence.
[0078] Exemplarily, waveform drawing is divided into initial drawing and drawing according to the waveform summary. After frame division is completed, the two processes of initial drawing and constructing the waveform summary can be performed in parallel. In the initial drawing, partial decoding of the audio is performed, and a rough waveform sketch can be quickly drawn. In this step, the waveform sketch can be drawn by a waveform drawer.
[0079] Specifically, a preset frame sequence (frames), a resource table (resourceMap), and the number of amplitudes to be output (i.e., the second preset number, denoted as ampCount) can be input into the waveform drawer. The waveform drawer outputs the amplitude values at the ampCount equal division points of the entire preset frame sequence, that is, outputs ampCount amplitude values (the range of each amplitude value can be between 0 and 1), constituting the waveform sketch.
[0080] Exemplarily, the following parameters can be set: the minimum number of frames for each decoding unit (e.g., minSegLen = 6), the maximum number of frames for each decoding unit (e.g., maxSegLen = 60), and the number of frames decoded each time (i.e., the preset number of decoding frames, e.g., decodeFrameCount = 3).
[0081] For the current preset frame sequence, calculate that if it is divided into ampCount intervals, the average number of frames per interval avgRangeLen = frames.length / ampCount, where frames.length represents the number of frame information in the preset frame sequence. If avgRangeLen is less than minSegLen, it means that ampCount is too high, resulting in an excessive amount of data to be decoded, approaching full decoding. Then, the initial drawing process can be terminated, and the waveform is drawn according to the summary after the waveform summary is constructed. If avgRangeLen is not less than minSegLen, then frames can be equally divided into ampCount segments (subsequences) in time, and the following operations are performed for each segment:
[0082] a. Denote the start and end frame numbers of the current segment (i.e., the serial numbers of the frame information in the preset frame sequence) as begin to end;
[0083] b. Calculate the decoding unit length segLen according to the current segment length end - begin:
[0084] Specifically, (end - begin) / n can be calculated, rounded and adjusted to be between minSegLen and maxSegLen, where n can be preset and its specific value is not limited. For example, it can be 10.
[0085] c. Calculate the number of decoding units segCount (the third preset quantity) contained in the current segment according to segLen;
[0086] d. For each decoding unit in the current segment, perform the following operations:
[0087] Record the starting frame position of the current decoding unit as beginIndex (equivalent to the decoding start frame identifier), call the decoder, start decoding decodeFrameCount frames of data from beginIndex, and find the maximum amplitude in the decoding result as the unit amplitude of the current decoding unit;
[0088] e. After obtaining the unit amplitude corresponding to each decoding unit in the current segment, use the maximum unit amplitude as the segment amplitude (sub - sequence amplitude) of the current segment.
[0089] When drawing the waveform for the first time, on - demand sampling and decoding of the audio resource is performed, which greatly reduces the first - draw time. The longer the audio, the more obvious the improvement effect. It has been verified that for a 90 - minute MP3 - format audio, the performance is improved by more than 10 times compared with full - volume decoding.
[0090] Step 303. Decode the audio resource corresponding to the preset frame sequence, determine the waveform summary information corresponding to each audio frame, store the waveform summary information in the preset frame sequence, and establish an association with the corresponding frame information.
[0091] Specifically, decode the audio resource corresponding to the preset frame sequence. For the decoded frame data of each audio frame, divide the current decoded frame data into sub - interval data of the first preset quantity, determine the interval amplitude corresponding to each sub - interval data, and determine the waveform summary information corresponding to the current audio frame according to the interval amplitudes. Store the waveform summary information corresponding to each audio frame in the preset frame sequence and establish an association with the corresponding frame information.
[0092] Exemplarily, in the process of constructing the waveform summary, a preset frame sequence (frames) and a resource table (resourceMap) can be input into the waveform plotter. The wave attribute of each audio frame, that is, the waveform summary information, can be output through the waveform plotter. The format can be Uin8Array, and each amplitude can be a value between 0 and 255. The following parameters can be set: the preset amplitude interval (msPerAmp) and the target duration (decodeTime, that is, the duration of each decoding).
[0093] Specifically, call the decoder to perform full decoding with decodeTime as the single decoding target duration. During each decoding process, record the frame range beginIndex (equivalent to the decoding start frame identifier) to endIndex (equivalent to the decoding end frame identifier) and the decoded data Data of this decoding. Traverse each frame included therein and perform the following operations on each frame:
[0094] a. Calculate the range of the data corresponding to this frame in Data based on the start and end times of this frame: from frameBeginSampleIndex to frameEndSampleIndex, and cut it out from Data, denoted as frameData (decoded frame data);
[0095] b. Calculate the number of amplitude values Count (the first preset quantity) that need to be generated for this frame based on the duration of this frame and the msPerAmp parameter, that is, the duration of this frame / msPerAmp;
[0096] c. Divide frameData into Count intervals (sub-interval data). For each interval, find the maximum amplitude as the amplitude of this interval (interval amplitude). Finally, obtain a Uint8Array containing Count amplitude values, denoted as the waveform summary information, and add the waveform summary information as the wave attribute to the frame information in the preset frame sequence.
[0097] By establishing the waveform summary at one time and saving it in the preset frame sequence, subsequent waveform plotting does not require decoding operations, and the performance is very high.
[0098] Step 304: Receive a preset audio editing operation, and perform corresponding editing operations on the corresponding frame information in the preset frame sequence according to the frame identifier to be adjusted indicated by the preset audio editing operation to implement audio editing.
[0099] Exemplarily, when initially constructing the preset frame sequence, the frame information can be arranged in the order of the original audio frames in the audio resource. During use, there may be various editing requirements. For example, if some audio frames in Audio 1 are to be inserted between two certain audio frames in Audio 2, in this case, the embodiments of the present disclosure do not need to operate on the decoded data, and the editing can be quickly completed by operating on the preset frame sequence to adjust the order of the frame information.
[0100] Step 305: Determine the target decoding duration and the decoding start frame identifier, start traversing from the frame information corresponding to the decoding start frame identifier in the preset frame sequence, and when any one of the preset traversal termination conditions is met, determine the decoding end frame identifier according to the corresponding frame information.
[0101] Among them, the preset traversal termination conditions include: the cumulative duration of the audio frames corresponding to the traversed frame information reaches the target decoding duration, the audio resource identifier in the current frame information is inconsistent with the audio resource identifier of the previous frame information, the frame index in the current frame information is not continuous with the frame index in the previous frame information, and the frame index in the current frame information is the last one in the corresponding audio resource.
[0102] Exemplarily, after editing on the basis of the initial preset frame sequence, there may be a situation where the frame information of audio frames of other audio resources is interspersed between the frame information of two audio frames of the same audio resource. At this time, in order to ensure that the data to be decoded participating in decoding comes from the same audio resource and is continuous in the audio resource, the above preset traversal termination conditions are set to dynamically determine the decoding end frame identifier.
[0103] Optionally, it can be determined whether to determine the target decoding duration and the decoding start frame identifier according to the amount of undecoded data in the playback buffer. When the web page is first opened, the playback buffer is usually empty. At this time, this step can be executed after the frame splitting process. At this time, the target decoding duration can be determined according to the settings of the player, and the decoding start frame identifier can be the frame identifier in the first frame information in the preset frame sequence. During the session, it can be determined whether to execute this step according to the actual situation.
[0104] Figure 4 This is a schematic diagram of an audio playback control process provided by the embodiments of the present disclosure. As Figure 4 shown, a circular playback buffer can be set, and it is determined whether a new segment needs to be decoded through a data loading scheduling strategy. Figure 4An audio playback context (AudioContext) may include one or more audio processing nodes, such as a ScriptProcessor, which can process audio data through a script. By filling the decoded audio data into the audio processing node, the content to be played can be controlled. A gain control node (GainNode) can be used for playback control. During playback, the ScriptProcessor is connected to it, and the two are disconnected during pausing. The playback buffer, also known as the data buffer, can be a ring buffer (RingBuffer). The loaded audio data is written into the playback buffer, and during playback, the audio data is read from it and filled into the ScriptProcessor. The data loading scheduling strategy continuously or periodically determines whether new data needs to be loaded as the playback progresses. If so, the decoder is called to load new data and write it into the playback buffer. In the embodiments of the present disclosure, a fill-type playback design can be implemented based on the ScriptProcessor, making it better fit the scenario of real-time loading and playback, ensuring the smoothness of playback, and strengthening the perception and control of the playback progress and status.
[0105] Exemplarily, a preset frame sequence, a resource table, a decoding start frame identifier, a target decoding duration, and a decoding sampling rate can be input to the decoder, and the actually decoded frame identifier, the duration of the actually decoded segment, the decoded sampling data, and whether the end of the audio resource file is reached can be output through the decoder.
[0106] Exemplarily, if the frame type of the audio frame corresponding to the decoding start frame identifier is MP3, the previous frame index of the frame index in the decoding start frame identifier can be determined first. If the frame index exists, the corresponding frame identifier is used as the new decoding start frame identifier, and the traversal of the frame information is started, that is, the traversal starts from the frame information of the previous frame of the start frame to be decoded in the original audio.
[0107] Step 306: Determine the audio resource associated with the audio resource identifier corresponding to the decoding start frame identifier as the target audio resource, determine the data start position according to the first frame offset corresponding to the decoding start frame identifier, determine the data end position according to the second frame offset and the frame data amount corresponding to the decoding end frame identifier, determine the target data range according to the data start position and the data end position, obtain the target storage information associated with the target audio resource from the resource table, and obtain the audio data within the target data range in the target audio resource based on the target storage information to obtain the data segment to be decoded.
[0108] Exemplarily, after the traversal is completed, the first frame (beginFrame) and the last frame (endFrame) to be decoded are obtained. Since the preset traversal termination condition can ensure that these two frames and the intermediate frames belong to the same audio resource and are continuous in position in the source file, a HyperText Transfer Protocol (HTTP) data request can be made. The request address is resourceMap[beginFrame.uri].url, and the requested data range is: beginFrame.offset~endFrame.offset+endFrame.size-1. After the data request is successful, the data (AudioClipData) of the segment to be decoded can be obtained.
[0109] Step 307: Decode the data of the segment to be decoded to obtain the corresponding target decoded data, and record the decoded end frame identifier and the decoding duration corresponding to the target decoded data.
[0110] Exemplarily, after obtaining the AudioClipData, the audio decoding interface of the web front end (such as BaseAudioContext.decodeAudioDate) can be called for decoding to obtain the decoded audio sampling data.
[0111] Optionally, for the case where the above audio frame is in MP3 format, the initial decoded data obtained by calling the audio decoding interface is cropped to remove redundant decoded data to obtain the target decoded data.
[0112] Step 308: Play the target decoded data.
[0113] Exemplarily, as described above, the target decoded data obtained after decoding can be first pushed to the playback buffer. When playback is required, it is filled into the audio processing node for playback.
[0114] Step 309: In response to receiving a preset waveform drawing instruction, according to the frame identifier to be drawn indicated by the preset waveform drawing instruction, obtain the target waveform summary information corresponding to the corresponding frame information in the preset frame sequence, and draw the corresponding waveform diagram according to the target waveform summary information.
[0115] Exemplarily, the waveform drawing tool is also responsible for drawing the waveform diagram according to the waveform summary information. When drawing the waveform diagram according to the waveform summary information, the waveform diagram corresponding to the entire preset frame sequence can be drawn. At this time, the frame identifier to be drawn can be all, or the waveform diagram corresponding to some frame information in the preset frame sequence can be drawn. At this time, the identifier to be drawn can include the starting frame identifier to be drawn and the ending frame identifier to be drawn.
[0116] Exemplarily, taking the drawing of the waveform diagram corresponding to the entire preset frame sequence as an example, the frame sequence, the resource table, and the number of output amplitudes (which can be denoted as the preset amplitude number) can be input to the waveform drawer. The preset frame sequence is equally divided into the preset amplitude number of subsequences according to time. For each subsequence, the start frame identifier and the end frame identifier corresponding to the current subsequence are determined. The waveform summary information corresponding to all frame information between the frame information to which the start frame identifier belongs and the frame information to which the end frame identifier belongs is traversed, and the maximum amplitude value is determined as the amplitude value corresponding to the current subsequence, obtaining the amplitude values of the preset amplitude number, so as to quickly obtain the waveform diagram.
[0117] Step 310: Synchronize the preset frame sequence and the resource table corresponding to the current session to the server.
[0118] Exemplarily, after the preset frame sequence and the resource table are obtained for the first time, they can be synchronized to the cloud, and during the continuous process of the session, synchronization can also continue. It should be noted that this synchronization process can be real-time, or triggered at every preset time interval, or triggered when the resource table or the preset frame sequence changes, and no specific limitation is made.
[0119] Exemplarily, the data volume of the resource table is generally small and can be stored in JSON format without serialization and compression processing. The data volume of the preset frame sequence is generally large. The preset frame sequence can be serialized into a binary format and then compressed, such as gzip compression, to meet the needs of network transmission, and generally can achieve the effect that the frame information only accounts for about 1.2M of data volume per hour.
[0120] Exemplarily, a frame field enumeration (FrameField) and the value type (FrameType) of each field can be defined. In this way, each field name of the frame can be stored with a unit8, and the field value is read and written in a specific format. The waveform summary information can adopt a custom data format, and its structure is that the first byte stores the number of amplitudes, and each subsequent byte stores the amplitude value of each amplitude. Each field in the frame is traversed, the field id is written in the format of unit8, and then the specific value is written according to the field value type. After that, the next field is processed in the same way. After all fields are serialized, the total length is written in the format of unit8 at the beginning of the serialization result. The serialization of multiple frames can splice the serialization results of each frame to obtain the serialization result of the preset frame sequence.
[0121] The audio processing method provided by the embodiments of the present disclosure performs frame splitting on an audio resource and outputs a frame sequence and a resource table. When audio decoding is required, on-demand decoding can be achieved, and different audio resources and frames of different formats can be supported for mixed storage in a preset frame sequence. The required audio segments are automatically calculated and loaded based on the characteristics of the input and the frames, making the decoding more flexible and improving the audio processing efficiency. By adopting a parallel first waveform drawing process and a waveform summary construction process, the first drawing time can be greatly reduced by performing partial decoding for the first drawing. After constructing the waveform summary information at one time, the performance of subsequent waveform drawing can be greatly improved. Moreover, the resource table and the frame sequence are timely synchronized to the cloud, and the data transmission volume is reduced through serialization and compression processing to ensure that session information is not lost.
[0122] Figure 5 The following is a structural block diagram of an audio processing device provided by the embodiments of the present disclosure. This device can be implemented by software and / or hardware and is generally integrated in an electronic device. It can perform audio processing by executing the audio processing method. As Figure 5 shown, the device includes:
[0123] A frame identifier determination module 501, configured to determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence. Among them, the preset frame sequence contains frame information of each audio frame in at least one audio resource, the frame information contains a frame identifier, the frame identifier includes an audio resource identifier and a frame index, the audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs;
[0124] A data to be decoded acquisition module 502, configured to obtain the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier;
[0125] A decoding module 503, configured to decode the data of the segment to be decoded to obtain corresponding target decoded data.
[0126] In the audio processing device provided by the embodiments of the present disclosure, the frame information of each audio frame in the audio resource is stored in sequence in advance. When decoding is required, the data range to be decoded is accurately located according to the decoding start frame identifier and the decoding end frame identifier, and the segment data is obtained from the corresponding audio resource and decoded. There is no need to perform full decoding of the audio file, and on-demand decoding is achieved, making the decoding more flexible and improving the audio processing efficiency.
[0127] Optionally, the frame identification determination module specifically includes: a first determination unit, configured to determine a target decoding duration and a decoding start frame identification; a second determination unit, configured to start traversing from the frame information corresponding to the decoding start frame identification in a preset frame sequence, and when a preset traversal termination condition is satisfied, determine a decoding end frame identification according to the corresponding frame information. Wherein, the preset traversal termination condition includes: the cumulative duration of the audio frames corresponding to the traversed frame information reaches the target decoding duration.
[0128] Optionally, the preset traversal termination condition further includes at least one of the following: the audio resource identification in the current frame information is inconsistent with the audio resource identification in the previous frame information; the frame index in the current frame information is not continuous with the frame index in the previous frame information; the frame index in the current frame information is the last one in the corresponding audio resource. Wherein, the step of determining a decoding end frame identification according to the corresponding frame information when the preset traversal termination condition is satisfied includes: when any one of the preset traversal termination conditions is satisfied, determining a decoding end frame identification according to the corresponding frame information.
[0129] Optionally, the frame information further includes a frame offset and a frame data amount. The module for obtaining data to be decoded specifically includes: a target audio resource determination unit, configured to determine the audio resource associated with the audio resource identification corresponding to the decoding start frame identification as the target audio resource; a target data range determination unit, configured to determine a data start position according to the first frame offset corresponding to the decoding start frame identification, determine a data end position according to the second frame offset and the frame data amount corresponding to the decoding end frame identification, and determine a target data range according to the data start position and the data end position; a data acquisition unit, configured to acquire audio data within the target data range in the target audio resource to obtain the data segment to be decoded.
[0130] Optionally, when the second determination unit performs the traversal starting from the frame information corresponding to the decoding start frame identification in the preset frame sequence, it is specifically configured to: determine the format of the audio frame corresponding to the decoding start frame identification; in the case where the format is a preset format, start traversing from the frame information corresponding to a target frame index in the preset frame sequence, where the target frame index is obtained by tracing back a preset frame index difference based on the start frame index in the decoding start frame identification; wherein, the module for obtaining data to be decoded is specifically configured to: obtain the data segment to be decoded in the audio resource associated with the corresponding audio resource identification according to the target frame identification corresponding to the target frame index and the decoding end frame identification.
[0131] Optionally, the decoding module is specifically configured to: when the format is a preset format, decode the to-be-decoded segment data to obtain corresponding initial decoded data; remove redundant decoded data from the initial decoded data to obtain corresponding target decoded data, where the redundant decoded data includes decoded data of audio frames corresponding to frame indexes before the start frame index.
[0132] Optionally, the apparatus further includes: a recording module, configured to record the decoding end frame identifier and the decoding duration corresponding to the target decoded data after decoding the to-be-decoded segment data to obtain the corresponding target decoded data.
[0133] Optionally, the apparatus is integrated into the web front end, and further includes: a frame information acquisition module, configured to perform frame splitting on the audio resource to obtain frame information of each audio frame in the audio resource before determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence; a frame information storage module, configured to store the obtained frame information into the preset frame sequence of the web front end.
[0134] Optionally, the apparatus is integrated into the web front end, and further includes: a meta information acquisition module, configured to acquire meta information of the audio resource before determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence, where the meta information includes storage information of the audio resource, and the storage information includes the storage location and / or resource data of the audio resource; a meta information storage module, configured to store the meta information into the resource table of the web front end, where the resource table includes an association relationship between the audio resource identifier involved in this session and the storage information; correspondingly, the to-be-decoded data acquisition module is specifically configured to: according to the decoding start frame identifier and the decoding end frame identifier, acquire the target storage information associated with the corresponding audio resource identifier from the resource table, and acquire the to-be-decoded segment data based on the target storage information.
[0135] Optionally, the apparatus further includes: an editing operation receiving module, configured to receive a preset audio editing operation; an audio editing module, configured to perform a corresponding editing operation on the corresponding frame information in the preset frame sequence according to the to-be-adjusted frame identifier indicated by the preset audio editing operation to implement audio editing, where the editing operation includes deleting and / or adjusting the order of the frame information.
[0136] Optionally, the preset frame sequence further includes waveform summary information corresponding to each audio frame; the apparatus further includes: a waveform summary acquisition module, configured to, in response to receiving a preset waveform drawing instruction, acquire target waveform summary information corresponding to the corresponding frame information in the preset frame sequence according to the to-be-drawn frame identifier indicated by the preset waveform drawing instruction; a waveform graph drawing module, configured to draw a corresponding waveform graph according to the target waveform summary information.
[0137] Optionally, the device includes: an audio resource decoding module, configured to decode the audio resource corresponding to the preset frame sequence before obtaining the target waveform summary information corresponding to the frame information in the preset frame sequence; a waveform summary determination module, configured to divide the decoded frame data of each decoded audio frame into sub-interval data of a first preset quantity, determine the interval amplitudes corresponding to the respective sub-interval data, and determine the waveform summary information corresponding to the current audio frame according to the respective interval amplitudes; and a waveform summary storage module, configured to store the waveform summary information corresponding to each audio frame into the preset frame sequence and establish an association with the corresponding frame information.
[0138] Optionally, the device includes: a first division module, configured to divide the preset frame sequence into sub-sequences of a second preset quantity; a sub-sequence amplitude determination module, configured to perform partial decoding on the current sub-sequence for each sub-sequence and determine the sub-sequence amplitude corresponding to the current sub-sequence according to the decoding result; and a waveform sketch drawing module, configured to draw a waveform sketch according to the sub-sequence amplitudes corresponding to the respective sub-sequences.
[0139] Optionally, the sub-sequence amplitude determination module includes: a first division unit, configured to divide the current sub-sequence into decoding units of a third preset quantity; a unit amplitude determination unit, configured to obtain data to be decoded according to the start frame identifier and the preset number of decoded frames corresponding to the current decoding unit for each decoding unit, and after decoding the data to be decoded, determine the maximum amplitude in the obtained decoded data as the unit amplitude of the current decoding unit; and a sub-sequence amplitude determination unit, configured to determine the maximum unit amplitude among the respective unit amplitudes as the sub-sequence amplitude corresponding to the current sub-sequence.
[0140] Optionally, the target decoded data is used to be stored in a playback buffer, and the device further includes: a data volume determination module, configured to determine whether to determine a decoding start frame identifier and a decoding end frame identifier in the preset frame sequence according to the data volume of the undecoded decoded data in the playback buffer.
[0141] Optionally, the device is applied to a web front end and further includes: a synchronization module, configured to synchronize the preset frame sequence corresponding to the current session to a server.
[0142] Reference is made below to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6The electronic device shown is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0143] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0144] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wirelessly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be implemented or included alternatively.
[0145] Specifically, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the methods of the embodiments of the present disclosure are executed.
[0146] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0147] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.
[0148] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to: determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence, where the preset frame sequence contains frame information of each audio frame in at least one audio resource, the frame information contains a frame identifier, the frame identifier includes an audio resource identifier and a frame index, the audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs; according to the decoding start frame identifier and the decoding end frame identifier, obtain the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier; decode the data of the segment to be decoded to obtain the corresponding target decoded data.
[0149] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0151] The modules described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the decoding module may also be described as "a module that decodes the to-be-decoded segment data to obtain the corresponding target decoded data".
[0152] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0153] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0154] According to one or more embodiments of the present disclosure, an audio processing method is provided, including:
[0155] Determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence, where the preset frame sequence contains frame information of each audio frame in at least one audio resource, the frame information contains a frame identifier, the frame identifier includes an audio resource identifier and a frame index, the audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs;
[0156] Obtain the data of the to-be-decoded segment in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier;
[0157] Decode the to-be-decoded segment data to obtain the corresponding target decoded data.
[0158] Further, the determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence includes:
[0159] Determine a target decoding duration and a decoding start frame identifier;
[0160] Start traversing from the frame information corresponding to the decoding start frame identifier in the preset frame sequence, and when a preset traversal termination condition is met, determine the decoding end frame identifier according to the corresponding frame information;
[0161] Wherein, the preset traversal termination condition includes:
[0162] The cumulative duration of the audio frames corresponding to the traversed frame information reaches the target decoding duration.
[0163] Further, the preset traversal termination condition further includes at least one of the following:
[0164] The audio resource identifier in the current frame information is inconsistent with the audio resource identifier in the previous frame information;
[0165] The frame index in the current frame information is not continuous with the frame index in the previous frame information;
[0166] The frame index in the current frame information is the last one in the affiliated audio resource;
[0167] Among them, when the preset traversal termination condition is satisfied, determining the decoding end frame identifier according to the corresponding frame information includes:
[0168] When any one of the preset traversal termination conditions is satisfied, determining the decoding end frame identifier according to the corresponding frame information.
[0169] Furthermore, the frame information further includes a frame offset and a frame data volume; obtaining the data segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier includes:
[0170] Determining the audio resource associated with the audio resource identifier corresponding to the decoding start frame identifier as the target audio resource;
[0171] Determining the data start position according to the first frame offset corresponding to the decoding start frame identifier, determining the data end position according to the second frame offset and the frame data volume corresponding to the decoding end frame identifier, and determining the target data range according to the data start position and the data end position;
[0172] Obtaining the audio data within the target data range in the target audio resource to obtain the data segment to be decoded.
[0173] Furthermore, starting to traverse from the frame information corresponding to the decoding start frame identifier in the preset frame sequence includes:
[0174] Determining the format of the audio frame corresponding to the decoding start frame identifier;
[0175] When the format is the preset format, starting to traverse from the frame information corresponding to the target frame index in the preset frame sequence, where the target frame index is the frame index obtained by tracing back a preset frame index difference based on the start frame index in the decoding start frame identifier;
[0176] Among them, obtaining the data segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier includes:
[0177] Obtain the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the target frame identifier corresponding to the target frame index and the decoding end frame identifier.
[0178] Further, when the format is a preset format, the decoding of the segment data to be decoded to obtain the corresponding target decoded data includes:
[0179] Decode the segment data to be decoded to obtain the corresponding initial decoded data;
[0180] Remove redundant decoded data from the initial decoded data to obtain the corresponding target decoded data, where the redundant decoded data includes the decoded data of the audio frames corresponding to the frame indices before the start frame index.
[0181] Further, after decoding the segment data to be decoded to obtain the corresponding target decoded data, it further includes:
[0182] Record the decoding end frame identifier and the decoding duration corresponding to the target decoded data.
[0183] Further, when applied to the web front end, before determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence, it further includes:
[0184] Perform frame splitting on the audio resource to obtain the frame information of each audio frame in the audio resource;
[0185] Store the obtained frame information into the preset frame sequence of the web front end.
[0186] Further, before determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence, it further includes:
[0187] Obtain the meta information of the audio resource, where the meta information includes the storage information of the audio resource, and the storage information includes the storage location and / or resource data of the audio resource;
[0188] Store the meta information into the resource table of the web front end, where the resource table includes the association relationship between the audio resource identifiers involved in this session and the storage information;
[0189] Correspondingly, the obtaining of the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier includes:
[0190] According to the decoding start frame identifier and the decoding end frame identifier, obtain the target storage information associated with the corresponding audio resource identifier from the resource table, and obtain the data of the segment to be decoded based on the target storage information.
[0191] Further, it further includes:
[0192] Receiving a preset audio editing operation;
[0193] According to the frame identification to be adjusted indicated by the preset audio editing operation, performing corresponding editing operations on the corresponding frame information in the preset frame sequence to implement audio editing, wherein the editing operations include deleting and / or adjusting the order of the frame information.
[0194] Further, the preset frame sequence further includes waveform summary information corresponding to each audio frame; the method further includes:
[0195] In response to receiving a preset waveform drawing instruction, according to the frame identification to be drawn indicated by the preset waveform drawing instruction, obtaining the target waveform summary information corresponding to the corresponding frame information in the preset frame sequence;
[0196] Drawing a corresponding waveform diagram according to the target waveform summary information.
[0197] Further, before obtaining the target waveform summary information corresponding to the corresponding frame information in the preset frame sequence, it further includes:
[0198] Decoding the audio resource corresponding to the preset frame sequence;
[0199] For the decoded frame data of each audio frame after decoding, dividing the current decoded frame data into sub-interval data of a first preset quantity, determining the interval amplitudes corresponding to the respective sub-interval data, and determining the waveform summary information corresponding to the current audio frame according to the respective interval amplitudes;
[0200] Storing the waveform summary information corresponding to each audio frame into the preset frame sequence and establishing an association with the corresponding frame information.
[0201] Further, it further includes:
[0202] Dividing the preset frame sequence into sub-sequences of a second preset quantity;
[0203] For each sub-sequence, performing partial decoding on the current sub-sequence and determining the sub-sequence amplitude corresponding to the current sub-sequence according to the decoding result;
[0204] Drawing a waveform sketch according to the sub-sequence amplitudes corresponding to the respective sub-sequences.
[0205] Further, the performing partial decoding on the current sub-sequence and determining the sub-sequence amplitude corresponding to the current sub-sequence includes:
[0206] Dividing the current sub-sequence into decoding units of a third preset quantity;
[0207] For each decoding unit, obtain the data to be decoded according to the start frame identifier corresponding to the current decoding unit and the preset number of decoding frames. After decoding the data to be decoded, determine the maximum amplitude in the obtained decoded data as the unit amplitude of the current decoding unit;
[0208] Determine the maximum unit amplitude among the unit amplitudes as the subsequence amplitude corresponding to the current subsequence.
[0209] Further, the target decoded data is used to be stored in a playback buffer, and the method further includes:
[0210] Determine whether to determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence according to the amount of undecoded data in the playback buffer.
[0211] Further, when applied to the web front end, it further includes: synchronizing the preset frame sequence corresponding to the current session to the server.
[0212] According to one or more embodiments of the present disclosure, an audio processing device is provided, including:
[0213] A frame identifier determination module, configured to determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence, where the preset frame sequence includes frame information of each audio frame in at least one audio resource, the frame information includes a frame identifier, the frame identifier includes an audio resource identifier and a frame index, the audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs;
[0214] A data to be decoded acquisition module, configured to acquire the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier;
[0215] A decoding module, configured to decode the data of the segment to be decoded to obtain corresponding target decoded data.
[0216] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the present disclosure.
[0217] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable subcombination in multiple embodiments.
[0218] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An audio processing method, characterized in that, Including: Determine a decoding start frame identifier and a decoding end frame identifier in a preset frame sequence, where the preset frame sequence includes frame information of each audio frame in at least one audio resource, the frame information includes a frame identifier, the frame identifier includes an audio resource identifier and a frame index, the audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs; Obtain the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier; Decode the data of the segment to be decoded to obtain corresponding target decoded data.
2. The method according to claim 1, characterized in that, The determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence includes: Determine a target decoding duration and a decoding start frame identifier; Start traversing from the frame information corresponding to the decoding start frame identifier in the preset frame sequence, and when a preset traversal termination condition is met, determine the decoding end frame identifier according to the corresponding frame information; Wherein, the preset traversal termination condition includes: The cumulative duration of the audio frames corresponding to the traversed frame information reaches the target decoding duration.
3. The method according to claim 2, characterized in that, The preset traversal termination condition further includes at least one of the following: The audio resource identifier in the current frame information is inconsistent with the audio resource identifier in the previous frame information; The frame index in the current frame information is not continuous with the frame index in the previous frame information; The frame index in the current frame information is the last one in the audio resource to which it belongs; Wherein, the determining the decoding end frame identifier according to the corresponding frame information when the preset traversal termination condition is met includes: When any one of the preset traversal termination conditions is met, determine the decoding end frame identifier according to the corresponding frame information.
4. The method according to claim 3, characterized in that, The frame information further includes a frame offset and a frame data volume; the obtaining the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier includes: Determine the audio resource associated with the audio resource identifier corresponding to the decoding start frame identifier as the target audio resource; Determine the data start position according to the first frame offset corresponding to the decoding start frame identifier, determine the data end position according to the second frame offset and the frame data volume corresponding to the decoding end frame identifier, and determine the target data range according to the data start position and the data end position; Obtain the audio data within the target data range in the target audio resource to obtain the data of the segment to be decoded.
5. The method according to claim 2, characterized in that, The starting to traverse from the frame information corresponding to the decoding start frame identifier in the preset frame sequence includes: Determine the format of the audio frame corresponding to the decoding start frame identifier; In the case where the format is a preset format, start traversing from the frame information corresponding to the target frame index in the preset frame sequence, where the target frame index is the frame index obtained by tracing back a preset frame index difference based on the start frame index in the decoding start frame identifier; Among them, obtaining the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier includes: Obtaining the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the target frame identifier corresponding to the target frame index and the decoding end frame identifier.
6. The method according to claim 5, characterized in that, In the case where the format is a preset format, decoding the data of the segment to be decoded to obtain the corresponding target decoded data includes: Decoding the data of the segment to be decoded to obtain the corresponding initial decoded data; Removing redundant decoded data from the initial decoded data to obtain the corresponding target decoded data, where the redundant decoded data includes the decoded data of the audio frames corresponding to the frame indices before the start frame index.
7. The method according to claim 3, characterized in that, After decoding the data of the segment to be decoded to obtain the corresponding target decoded data, it further includes: Recording the decoding end frame identifier and the decoding duration corresponding to the target decoded data.
8. The method according to claim 1, characterized in that, Applied to the web front end, before determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence, it further includes: Performing frame splitting on the audio resource to obtain the frame information of each audio frame in the audio resource; Storing the obtained frame information into the preset frame sequence of the web front end.
9. The method according to claim 8, wherein, Before determining the decoding start frame identifier and the decoding end frame identifier in the preset frame sequence, it further includes: Obtaining the meta information of the audio resource, where the meta information includes the storage information of the audio resource, and the storage information includes the storage location and / or resource data of the audio resource; Storing the meta information into the resource table of the web front end, where the resource table includes the association relationship between the audio resource identifier involved in this session and the storage information; Correspondingly, obtaining the data of the segment to be decoded in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier includes: According to the decoding start frame identifier and the decoding end frame identifier, obtaining the target storage information associated with the corresponding audio resource identifier from the resource table, and obtaining the data of the segment to be decoded based on the target storage information.
10. The method according to claim 1, wherein, It further includes: Receiving a preset audio editing operation; Performing corresponding editing operations on the corresponding frame information in the preset frame sequence according to the frame identifier to be adjusted indicated by the preset audio editing operation to implement audio editing, where the editing operation includes deleting and / or adjusting the order of the frame information.
11. The method according to claim 1, wherein, The preset frame sequence further includes waveform summary information corresponding to each audio frame; the method further includes: In response to receiving a preset waveform drawing instruction, obtaining the target waveform summary information corresponding to the frame information corresponding to the frame identifier to be drawn indicated by the preset waveform drawing instruction from the preset frame sequence; Drawing a corresponding waveform diagram according to the target waveform summary information.
12. The method according to claim 11, wherein, Before obtaining the target waveform summary information corresponding to the frame information corresponding to the preset frame sequence, it further includes: Decoding the audio resource corresponding to the preset frame sequence; For the decoded frame data of each decoded audio frame, divide the current decoded frame data into sub-interval data of a first preset quantity, determine the interval amplitudes corresponding to the respective sub-interval data, and determine the waveform summary information corresponding to the current audio frame according to the respective interval amplitudes; Store the waveform summary information corresponding to each audio frame into the preset frame sequence and establish an association with the corresponding frame information.
13. The method according to claim 1, wherein, It further includes: Divide the preset frame sequence into sub-sequences of a second preset quantity; For each sub-sequence, perform partial decoding on the current sub-sequence and determine the sub-sequence amplitude corresponding to the current sub-sequence according to the decoding result; Draw a waveform sketch according to the sub-sequence amplitudes corresponding to the respective sub-sequences.
14. The method according to claim 13, wherein, The performing partial decoding on the current sub-sequence and determining the sub-sequence amplitude corresponding to the current sub-sequence according to the decoding result includes: Divide the current sub-sequence into decoding units of a third preset quantity; For each decoding unit, obtain the data to be decoded according to the start frame identifier and the preset number of decoded frames corresponding to the current decoding unit, and after decoding the data to be decoded, determine the maximum amplitude in the obtained decoded data as the unit amplitude of the current decoding unit; Determine the maximum unit amplitude among the respective unit amplitudes as the sub-sequence amplitude corresponding to the current sub-sequence.
15. The method according to claim 1, wherein, The target decoded data is used to be stored into the playback buffer, and the method further includes: Determine whether to determine a decoding start frame identifier and a decoding end frame identifier in the preset frame sequence according to the data volume of the undecoded data in the playback buffer.
16. The method according to claim 1, wherein, When applied to the web front end, it further includes: Synchronize the preset frame sequence corresponding to the current session to the server.
17. An audio processing apparatus, wherein, It includes: A frame identifier determination module, configured to determine a decoding start frame identifier and a decoding end frame identifier in the preset frame sequence, wherein the preset frame sequence includes the frame information of each audio frame in at least one audio resource, the frame information includes a frame identifier, the frame identifier includes an audio resource identifier and a frame index, the audio resource identifier is used to represent the identity of the audio resource to which the corresponding audio frame belongs, and the frame index is used to represent the order of the corresponding audio frame among all audio frames of the audio resource to which it belongs; An undecoded data acquisition module, configured to obtain the undecoded segment data in the audio resource associated with the corresponding audio resource identifier according to the decoding start frame identifier and the decoding end frame identifier; A decoding module, configured to decode the undecoded segment data to obtain the corresponding target decoded data.
18. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1-16.
19. A computer-readable storage medium, having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-16.
Citation Information
Patent Citations
Voice coded data editing device and method
JP2009216915A