Audio fixed-point playing method, related device, equipment, system and storage medium
By determining the target frame number and predicted data amount based on audio metadata and timestamps in audio fixed-point playback, and updating the predicted data amount, the problem of low accuracy of fixed-point playback in the prior art is solved, and efficient and accurate audio fixed-point playback is achieved.
Patent Information
- Application Number
- CN202411885586.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the positioning points recorded in the positioning table in the audio file are too sparse, resulting in poor audio accuracy for fixed-point playback.
By determining the target frame number and predicted data amount based on the metadata of the target audio and the first time stamp indicating the fixed-point playback of the target audio, the target frame number and the predicted data amount are updated until the candidate frame number is consistent with the target frame number, precise fixed-point playback is achieved.
It improves the accuracy and efficiency of fixed-point playback, and can accurately predict and locate audio data when the audio data is unevenly distributed.
Smart Images

Figure CN120032672A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to an audio fixed-point playback method and related devices, equipment, systems and storage media. Background Art
[0002] Some audio encodings, such as Free Lossless Audio Codec (Flac), contain a positioning table in the audio file so that the positioning points stored in the positioning table can be used during fixed-point playback.
[0003] In the prior art, during fixed-point playback, the positioning points in the audio stream are usually saved based on the positioning table in the audio file to quickly locate the corresponding position in the audio stream. However, since the positioning points recorded in the positioning table in the audio file are too sparse, the audio accuracy of the actual fixed-point playback is poor. In view of this, how to improve the accuracy of the fixed-point playback of the target audio while improving the efficiency of the fixed-point playback as much as possible has become an urgent problem to be solved. Summary of the invention
[0004] The main technical problem solved by the present application is to provide an audio fixed-point playback method and related devices, equipment, system and storage medium, which can improve the accuracy of target audio fixed-point playback while improving the fixed-point playback efficiency as much as possible. In order to solve the above technical problems, the first aspect of the present application provides an audio fixed-point playback method, comprising determining a target frame number and a predicted data amount based on metadata of a target audio and a first timestamp indicating fixed-point playback of the target audio; wherein the target audio at least also includes audio data, the metadata is used to define audio parameters, the audio data contains a number of audio frames, the target frame number represents the frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data amount represents the amount of audio data predicted in the audio data as of the first timestamp; selecting a frame number corresponding to the predicted data amount at a positioning point of the audio data as a candidate frame number; in response to inconsistency between the candidate frame number and the target frame number, updating the predicted data amount based on the difference between the candidate frame number and the target frame number, and returning to execute the step of selecting a frame number corresponding to the predicted data amount at the positioning point of the audio data as the candidate frame number, until the candidate frame number is consistent with the target frame number; based on the latest predicted data amount, playing the target audio at a fixed point.
[0005] In order to solve the above technical problems, the second aspect of the present application provides an audio fixed-point playback device, including a determination module, a positioning module, an update module and a playback module, wherein the determination module is used to determine the target frame number and the predicted data volume based on the metadata of the target audio and the first timestamp indicating the fixed-point playback of the target audio; wherein the target audio at least also includes audio data, the metadata is used to define audio parameters, the audio data contains a number of audio frames, the target frame number represents the frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data volume represents the audio data volume predicted in the audio data as of the first timestamp; the positioning module is used to select the frame number corresponding to the predicted data volume at the positioning point of the audio data as the candidate frame number; the update module is used to update the predicted data volume based on the difference between the candidate frame number and the target frame number in response to the inconsistency between the candidate frame number and the target frame number, and return to execute the step of selecting the frame number corresponding to the predicted data volume at the positioning point of the audio data as the candidate frame number until the candidate frame number is consistent with the target frame number; the playback module is used to play the target audio at a fixed point based on the latest predicted data volume.
[0006] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other, the memory storing program instructions, and the processor being used to execute the program instructions to implement the audio fixed-point playback method of the above first aspect.
[0007] In order to solve the above-mentioned technical problems, the fourth aspect of the present application provides an audio playback system, including an audio playback application layer and an audio playback processing layer, the audio playback application layer is used to obtain fixed-point playback instructions for the target audio and display the timestamp of the fixed-point playback of the target audio, and the audio playback processing layer is the electronic device of the third aspect mentioned above, which is used for fixed-point playback of the target audio.
[0008] In order to solve the above-mentioned technical problems, the fifth aspect of the present application provides a computer-readable storage medium / computer program product, wherein the computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the audio fixed-point playback method of the first aspect is implemented; the computer program product includes computer program instructions, and when the computer program instructions are executed by a processor, the audio fixed-point playback method of the first aspect is implemented.
[0009] In the above scheme, the target audio includes at least metadata and audio data, the metadata is used to define audio parameters, and the audio data includes a number of audio frames. The target frame number and the predicted data volume are determined based on the metadata and the first timestamp indicating the fixed-point playback of the target audio. The target frame number represents the frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data volume represents the amount of audio data predicted to be in the audio data as of the first timestamp. The predicted data volume is taken at a positioning point in the audio data as the predicted positioning point, and the frame number of the audio frame corresponding to the predicted positioning point in the audio data is obtained as the candidate frame number. In response to the inconsistency between the candidate frame number and the target frame number, the predicted data volume is updated based on the difference between the candidate frame number and the target frame number, and the step of taking the predicted data volume at the positioning point in the audio data as the predicted positioning point is returned to be executed until the candidate frame number is consistent with the target frame number, and the target audio is played at a fixed point based on the latest predicted data volume. On the one hand, since the target frame number and the initial predicted data volume are both generated based on the first timestamp and the target audio metadata, the difference between the candidate frame number and the target frame number is as small as possible, which can improve the efficiency of the adjustment and update process. On the other hand, based on the audio format of the target audio, the target frame number corresponding to the first timestamp is obtained, and then the target frame number is used as reference information to adjust and update the predicted data volume until the candidate frame number corresponding to the audio data positioning point is consistent with the target frame number. In the case of uneven distribution of the audio data of the target audio, the accuracy of the predicted data volume for fixed-point playback of the target audio can be improved. Therefore, the accuracy of the fixed-point playback of the target audio can be improved while improving the efficiency of fixed-point playback as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a flowchart of an embodiment of the method for fixed-point audio playback of the present application; Figure 2 It is a schematic diagram of the framework of an embodiment of the audio fixed-point playback device of the present application; Figure 3 It is a schematic diagram of the framework of an embodiment of the electronic device of the present application; Figure 4 It is a schematic diagram of the framework of an embodiment of the audio playback system of the present application; Figure 5 It is a schematic diagram of the framework of an embodiment of the computer-readable storage medium / computer program product of the present application. DETAILED DESCRIPTION
[0011] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0012] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.
[0013] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the method for fixed-point audio playback of the present application. Specifically, it may include the following steps: Step S10: Determine a target frame number and a predicted data amount based on metadata of the target audio and a first timestamp indicating fixed-point playback of the target audio.
[0014] In the disclosed embodiments, the target audio is in a lossless audio compression encoding format, such as a Flac audio file, etc. The target audio includes at least metadata and audio data. The metadata is used to define audio parameters, to describe the audio stream information of the target audio and some ancillary information, such as stream information, padding blocks, application data, positioning tables, tag information, index tables and pictures, etc., wherein the stream information contains information about the entire audio stream, such as sampling rate, number of channels, total number of samples, etc., and the audio data contains a number of audio frames, which are the actual content data of the target audio.
[0015] In one implementation scenario, based on a fixed-point playback instruction for the target audio, the first timestamp of the fixed-point playback is determined. It should be noted that the specific form of the fixed-point playback instruction for the target audio is not limited in this application, such as dragging the playback progress bar of the target audio, entering the desired playback time point, etc.
[0016] In one implementation scenario, the target frame number represents the frame number of the audio frame corresponding to the first timestamp in the audio data, and the target frame number is related to the audio format of the target audio.
[0017] In a specific implementation scenario, based on parameters such as a sampling rate and a number of channels of the audio format in the metadata, a frame number of an audio frame corresponding to the first timestamp in the audio data is determined.
[0018] In another specific implementation scenario, based on the conversion rules of the Flac standard specification, the target frame number is obtained based on the first timestamp conversion.
[0019] It can be understood that when the sampling rate, number of channels and other parameters of the audio format are obtained, the timestamp and the frame number have a mapping relationship, so the target frame number is used as reference information for the subsequent adjustment process. Since the audio data is not arranged evenly, the amount of data at different stages may vary according to the characteristics of the song itself at different time periods, such as prelude, chorus, complex sound effects, etc. Therefore, the specific positioning point in the audio data cannot be obtained based only on the target frame number.
[0020] In one implementation scenario, the predicted data amount represents the amount of audio data predicted to be present in the audio data as of the first time stamp, and the predicted data amount is related to audio content of the target audio.
[0021] In a specific implementation scenario, it is detected whether there is a positioning table for the target audio in the metadata, the positioning table includes a mapping relationship between at least one preset sampling number and the amount of audio data, and based on the detection result of the positioning table, a method for obtaining the predicted data amount is determined.
[0022] In a specific implementation scenario, when there is no positioning table for the target audio in the metadata, the predicted data volume is obtained based on the audio parameters and the first timestamp in the metadata that characterize the total duration and total data volume of the target audio.
[0023] In a specific implementation scenario, the predicted data volume is obtained based on the proportion of the first timestamp in the total playback time of the target audio. For example, it is known that the first timestamp is 10s, and the total playback time of the target audio is 100s based on the metadata, and the total data size is 2000KB. The predicted data volume = first timestamp / total playback time*total data size, that is, 10s / 100s*2000KB=200KB, that is, when the first timestamp is predicted to be 10s, the corresponding predicted data volume is at 200KB of the audio data of the target audio.
[0024] In a specific implementation scenario, when there is a positioning table for the target audio, the predicted data volume is obtained based on the audio parameters representing the positioning table and the audio sampling rate in the metadata, and the first timestamp.
[0025] In a specific implementation scenario, the positioning table for the target audio includes at least one mapping relationship between a preset sampling number and an amount of audio data. For example, when the preset sampling number is 1000, the corresponding amount of audio data is 50KB, and when the preset sampling number is 1500, the corresponding amount of audio data is 85KB, and so on. Based on the product between the audio sampling rate and the first timestamp, the target sampling number corresponding to the first timestamp is obtained, and the mapping relationship of the preset sampling number closest to the target sampling number in the positioning table is selected, and the audio data amount corresponding to the mapping relationship of the closest preset sampling number is used as the predicted data amount.
[0026] It is understandable that in some specific implementation scenarios, there is a preset sampling number consistent with the target sampling number in the positioning table, and in other specific implementation scenarios, there is no preset sampling number consistent with the target sampling number in the positioning table, which is not limited in this application.
[0027] In a specific implementation scenario, when the positioning table has a preset sampling number that is consistent with the target sampling number, the amount of audio data corresponding to the mapping relationship of the preset sampling number that is consistent with the target sampling number is used as the predicted data amount. It can be understood that at this time, the frame number corresponding to the audio data positioning point of the predicted data amount is consistent with the target frame number.
[0028] Step S20: Select the frame number corresponding to the positioning point of the audio data at the predicted data volume as the candidate frame number.
[0029] In one implementation scenario, audio data is read based on the predicted data volume, a predicted positioning point about the predicted data volume is obtained, and the frame number of the audio frame corresponding to the predicted positioning point in the audio data is obtained as the candidate frame number. The above scheme can be applied to various situations where there is a positioning table in the metadata and there is no positioning table, and improves the versatility of fixed-point audio playback. Since the target frame number and the initial predicted data volume are both generated based on the first timestamp and the target audio metadata, the difference between the candidate frame number and the target frame number is as small as possible, which can improve the efficiency of the adjustment and update process.
[0030] Step S30: In response to the inconsistency between the candidate frame number and the target frame number, the predicted data amount is updated based on the difference between the candidate frame number and the target frame number, and the step of selecting the frame number corresponding to the positioning point of the audio data with the predicted data amount is returned to be executed as the candidate frame number until the candidate frame number is consistent with the target frame number.
[0031] In one implementation scenario, a data offset is determined based on the difference between the candidate frame number and the target frame number, the data offset represents the numerical adjustment size of the predicted data volume, and the updated predicted data volume is obtained based on the data offset. In the above scheme, since the target frame number and the initial predicted data volume are both generated based on the first timestamp and the target audio metadata, the difference between the candidate frame number and the target frame number is as small as possible, which can improve the efficiency of the adjustment and update process, and obtain the target frame number corresponding to the first timestamp based on the audio format of the target audio, and then use the target frame number as reference information to adjust and update the predicted data volume until the predicted data volume is consistent with the target frame number at the candidate frame number corresponding to the audio data positioning point, which can improve the accuracy of the predicted data volume for fixed-point playback of the target audio when the audio data of the target audio is unevenly distributed.
[0032] In a specific implementation scenario, before determining the data offset based on the difference between the candidate frame number and the target frame number, the maximum data amount of each audio frame in the audio parameter is obtained based on the audio parameter in the metadata used to define the data amount of each audio frame of the target audio, and the data offset is obtained based on half of the product of the target difference and the maximum data amount. The target difference represents the difference between the candidate frame number and the target frame number. That is, it is assumed that the data frames between the target frame number and the candidate frame number are all the largest data frames in the target audio, that is, the maximum value of the theoretical distance of the data amount between the target frame number and the candidate frame number. In fact, for the target audio, since it is usually impossible for all of them to be the largest data frames, the data amount between the target frame number and the candidate frame number will be smaller than this theoretical maximum value. Then, half of the value is taken, that is, the binary method is used for adjustment, and then the target frame number is gradually approached in multiple adjustments, which can improve the efficiency of adjustment and update.
[0033] In a specific implementation scenario, when the candidate frame number is greater than the target frame number, an updated predicted data amount is obtained based on the difference between the predicted data amount and the data offset.
[0034] In a specific implementation scenario, when the candidate frame number is smaller than the target frame number, an updated predicted data amount is obtained based on the sum of the predicted data amount and the data offset.
[0035] In a specific implementation scenario, when there is no positioning table, the amount of candidate data deviates. For example, the predicted data amount based on the first timestamp of 10s is 200KB, but in fact the data amount corresponding to 10s is 150KB. The reason is that the data arrangement of the audio data of the target audio is not evenly arranged. The amount of data at different stages may vary according to the characteristics of the song itself at different time periods (such as prelude, chorus, complex sound effects), and the predicted data amount is calculated based on the assumption that the data amount is evenly distributed. Therefore, the candidate data amount deviates, which leads to inconsistency between the candidate frame number and the target frame number.
[0036] In a specific implementation scenario, when there is a positioning table, the amount of candidate data deviates. This is usually because the table entries recorded in the positioning table in the target audio metadata are not particularly dense, and the two preset sampling numbers corresponding to some adjacent table entries are spaced at a large distance. For example, the calculated target sampling number is 2000, but the preset sampling number recorded in the table entry closest to it in the positioning table is 2200, so the predicted data amount is the audio data amount corresponding to the preset sampling number 2200. Therefore, the candidate data amount deviates, which leads to inconsistency between the candidate frame number and the target frame number.
[0037] Step S40: Based on the latest predicted data volume, play the target audio at a fixed point.
[0038] In one implementation scenario, based on the latest predicted data volume, the target audio is played at a fixed point. On the one hand, since the target frame number and the initial predicted data volume are both generated based on the first timestamp and the target audio metadata, the difference between the candidate frame number and the target frame number is as small as possible, which can improve the efficiency of the adjustment and update process. On the other hand, based on the audio format of the target audio, the target frame number corresponding to the first timestamp is obtained, and then the target frame number is used as reference information to adjust and update the predicted data volume until the candidate frame number corresponding to the audio data positioning point of the predicted data volume is consistent with the target frame number. In the case of uneven distribution of the audio data of the target audio, the accuracy of the predicted data volume used for fixed-point playback of the target audio can be improved. Therefore, the accuracy of the fixed-point playback of the target audio can be improved while the efficiency of fixed-point playback is improved as much as possible.
[0039] In one implementation scenario, a second timestamp corresponding to the latest predicted data volume in the audio data is obtained, and based on the positioning point of the latest predicted data volume in the audio data, the audio data corresponding to the target audio is played, and the current timestamp of the target audio is positioned to the second timestamp.
[0040] In a specific implementation scenario, the current timestamp of the target audio is positioned to the second timestamp, for example, a playback progress bar of the target audio is dragged to the second timestamp, and the audio playback time point is displayed based on the second timestamp.
[0041] In the above scheme, the target audio includes at least metadata and audio data, the metadata is used to define audio parameters, and the audio data includes a number of audio frames. The target frame number and the predicted data volume are determined based on the metadata and the first timestamp indicating the fixed-point playback of the target audio. The target frame number represents the frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data volume represents the amount of audio data predicted to be in the audio data as of the first timestamp. The predicted data volume is taken at a positioning point in the audio data as the predicted positioning point, and the frame number of the audio frame corresponding to the predicted positioning point in the audio data is obtained as the candidate frame number. In response to the inconsistency between the candidate frame number and the target frame number, the predicted data volume is updated based on the difference between the candidate frame number and the target frame number, and the step of taking the predicted data volume at the positioning point in the audio data as the predicted positioning point is returned to be executed until the candidate frame number is consistent with the target frame number, and the target audio is played at a fixed point based on the latest predicted data volume. On the one hand, since the target frame number and the initial predicted data volume are both generated based on the first timestamp and the target audio metadata, the difference between the candidate frame number and the target frame number is as small as possible, which can improve the efficiency of the adjustment and update process. On the other hand, based on the audio format of the target audio, the target frame number corresponding to the first timestamp is obtained, and then the target frame number is used as reference information to adjust and update the predicted data volume until the candidate frame number corresponding to the audio data positioning point is consistent with the target frame number. In the case of uneven distribution of the audio data of the target audio, the accuracy of the predicted data volume for fixed-point playback of the target audio can be improved. Therefore, the accuracy of the fixed-point playback of the target audio can be improved while improving the efficiency of fixed-point playback as much as possible.
[0042] See also Figure 2 , Figure 2 2 is a flow chart of an embodiment of the audio fixed-point playback device 20 of the present application. Figure 2As shown, the audio fixed-point playback device 20 includes a determination module 21, a positioning module 22, an updating module 23 and a playback module 24, wherein the determination module 21 is used to determine the target frame number and the predicted data volume based on the metadata of the target audio and the first timestamp indicating the fixed-point playback of the target audio; wherein the target audio at least includes audio data, the metadata is used to define audio parameters, the audio data includes a number of audio frames, the target frame number represents the frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data volume represents the audio data volume predicted in the audio data as of the first timestamp; the positioning module 22 is used to select the frame number corresponding to the predicted data volume at the positioning point of the audio data as the candidate frame number; the updating module 23 is used to respond to the inconsistency between the candidate frame number and the target frame number, based on the difference between the candidate frame number and the target frame number, update the predicted data volume, and return to execute the step of selecting the frame number corresponding to the predicted data volume at the positioning point of the audio data as the candidate frame number, until the candidate frame number is consistent with the target frame number; the playback module 24 is used to play the target audio at the fixed point based on the latest predicted data volume.
[0043] Therefore, the target audio of the audio fixed-point playback device 20 includes at least metadata and audio data, the metadata is used to define audio parameters, and the audio data includes a number of audio frames. The target frame number and the predicted data volume are determined based on the metadata and the first timestamp indicating the fixed-point playback of the target audio. The target frame number represents the frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data volume represents the amount of audio data predicted to be in the audio data as of the first timestamp. The predicted data volume is used as the prediction positioning point at the positioning point of the audio data, and the frame number of the audio frame corresponding to the prediction positioning point in the audio data is obtained as the candidate frame number. In response to the inconsistency between the candidate frame number and the target frame number, the predicted data volume is updated based on the difference between the candidate frame number and the target frame number, and the step of using the predicted data volume at the positioning point of the audio data as the prediction positioning point is returned to be executed until the candidate frame number is consistent with the target frame number. Based on the latest predicted data volume, the target audio is played at a fixed point. On the one hand, since the target frame number and the initial predicted data volume are both generated based on the first timestamp and the target audio metadata, the difference between the candidate frame number and the target frame number is as small as possible, which can improve the efficiency of the adjustment and update process. On the other hand, based on the audio format of the target audio, the target frame number corresponding to the first timestamp is obtained, and then the target frame number is used as reference information to adjust and update the predicted data volume until the candidate frame number corresponding to the audio data positioning point is consistent with the target frame number. In the case of uneven distribution of the audio data of the target audio, the accuracy of the predicted data volume for fixed-point playback of the target audio can be improved. Therefore, the accuracy of the fixed-point playback of the target audio can be improved while improving the efficiency of fixed-point playback as much as possible.
[0044] In some disclosed embodiments, the update module 23 also includes an offset determination module (not shown) for determining the data offset based on the difference between the candidate frame number and the target frame number; the update module 23 also includes a first update submodule (not shown) for obtaining an updated predicted data amount based on the difference between the predicted data amount and the data offset in response to the candidate frame number being greater than the target frame number; the update module 23 also includes a second update submodule (not shown) for obtaining an updated predicted data amount based on the sum of the predicted data amount and the data offset in response to the candidate frame number being less than the target frame number.
[0045] In some disclosed embodiments, before determining the data offset based on the difference between the candidate frame number and the target frame number, the audio fixed-point playback device 20 also includes a maximum data amount acquisition module (not shown), which is used to obtain the maximum data amount of each audio frame in the audio parameters based on the audio parameters in the metadata used to define the data amount size of each audio frame of the target audio; the offset determination module (not shown) also includes a first calculation module (not shown), which is used to obtain the data offset based on half of the product of the target difference and the maximum data amount; wherein the target difference represents the difference between the candidate frame number and the target frame number.
[0046] In some disclosed embodiments, the determination module 21 also includes a positioning table detection module (not shown) for detecting whether there is a positioning table for the target audio in the metadata; wherein the positioning table includes a mapping relationship between at least one preset sampling number and the amount of audio data; the determination module 21 also includes a first prediction module (not shown) for obtaining a predicted data amount based on audio parameters and a first timestamp representing the total duration of the target audio and the total amount of target audio data in the metadata in response to the absence of a positioning table for the target audio; the determination module 21 also includes a second prediction module (not shown) for obtaining a predicted data amount based on audio parameters and a first timestamp representing the positioning table and the audio sampling rate in the metadata in response to the presence of a positioning table for the target audio.
[0047] In some disclosed embodiments, the second prediction module (not shown) further includes a second calculation module (not shown) for obtaining a target sampling number corresponding to the first timestamp based on the product between the audio sampling rate and the first timestamp; the second prediction module (not shown) further includes a mapping relationship selection module (not shown) for selecting a mapping relationship of a preset sampling number in a positioning table that is closest to the target sampling number; the second prediction module (not shown) further includes a predicted data amount determination module (not shown) for using the audio data amount corresponding to the mapping relationship of the closest preset sampling number as the predicted data amount.
[0048] In some disclosed embodiments, the playback module 24 also includes a second timestamp acquisition module (not shown) for acquiring a second timestamp corresponding to the latest predicted data amount in the audio data; the playback module 24 also includes a fixed-point playback module (not shown) for playing the target audio based on the positioning point of the latest predicted data amount in the audio data, and positioning the current timestamp of the target audio to the second timestamp.
[0049] See also Figure 3 , Figure 3 Schematic diagram of the framework of an embodiment of the electronic device 30 of the present application. Figure 3 As shown, the electronic device 30 includes a memory 31 and a processor 32 coupled to each other, the memory 31 stores program instructions, and the processor 32 is used to execute the program instructions to implement the steps in any of the above-mentioned audio fixed-point playback method embodiments. Specifically, the electronic device 30 may include but is not limited to: a server, a desktop computer, a laptop computer, a tablet computer, a smart phone, etc., which are not limited here. Specifically, the processor 32 is used to control itself and the memory 31 to implement the steps in any of the above-mentioned audio fixed-point playback method embodiments. The processor 32 can also be called a CPU (Central Processing Unit). The processor 32 may be an integrated circuit chip with signal processing capabilities. The processor 32 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 32 can be implemented by an integrated circuit chip.
[0050] Therefore, the target audio of the electronic device 30 includes at least metadata and audio data, the metadata is used to define audio parameters, and the audio data includes a number of audio frames. The target frame number and the predicted data volume are determined based on the metadata and the first timestamp indicating the fixed-point playback of the target audio. The target frame number represents the frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data volume represents the amount of audio data predicted to be present in the audio data as of the first timestamp. The predicted data volume is taken at the positioning point of the audio data as the predicted positioning point, and the frame number of the audio frame corresponding to the predicted positioning point in the audio data is obtained as the candidate frame number. In response to the inconsistency between the candidate frame number and the target frame number, the predicted data volume is updated based on the difference between the candidate frame number and the target frame number, and the step of taking the predicted data volume at the positioning point of the audio data as the predicted positioning point is returned to be executed until the candidate frame number is consistent with the target frame number, and the target audio is played at the fixed point based on the latest predicted data volume. On the one hand, since the target frame number and the initial predicted data volume are both generated based on the first timestamp and the target audio metadata, the difference between the candidate frame number and the target frame number is as small as possible, which can improve the efficiency of the adjustment and update process. On the other hand, based on the audio format of the target audio, the target frame number corresponding to the first timestamp is obtained, and then the target frame number is used as reference information to adjust and update the predicted data volume until the candidate frame number corresponding to the audio data positioning point is consistent with the target frame number. In the case of uneven distribution of the audio data of the target audio, the accuracy of the predicted data volume for fixed-point playback of the target audio can be improved. Therefore, the accuracy of the fixed-point playback of the target audio can be improved while improving the efficiency of fixed-point playback as much as possible.
[0051] See also Figure 4 , Figure 4 4 is a schematic diagram of the framework of an embodiment of the audio playback system 40 of the present application. Figure 4 As shown, the audio playback system 40 includes an audio playback application layer 41 and an audio playback processing layer 42. The audio playback application layer 41 is used to obtain fixed-point playback instructions for the target audio and display the timestamp of the fixed-point playback of the target audio, and the audio playback processing layer 42 is the electronic device 30 in the aforementioned embodiment, which is used for fixed-point playback of the target audio.
[0052] In one implementation scenario, the audio playback processing layer 42 includes an audio playback framework layer 421 and an audio playback parser 422. The audio playback framework layer 421 is used to determine the first timestamp of the fixed-point playback and send it to the audio playback parser 422. The audio playback parser 422 is used to determine the predicted data amount and send it to the audio playback framework layer 421. The audio playback framework layer 421 is also used to play the target video at a fixed point based on the predicted data amount.
[0053] In the above scheme, the target audio of the audio playback system 40 includes at least metadata and audio data, the metadata is used to define audio parameters, and the audio data includes a number of audio frames. The target frame number and the predicted data volume are determined based on the metadata and the first timestamp indicating the fixed-point playback of the target audio. The target frame number represents the frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data volume represents the amount of audio data predicted to be in the audio data as of the first timestamp. The predicted data volume is used at the positioning point of the audio data as the predicted positioning point, and the frame number of the audio frame corresponding to the predicted positioning point in the audio data is obtained as the candidate frame number. In response to the inconsistency between the candidate frame number and the target frame number, the predicted data volume is updated based on the difference between the candidate frame number and the target frame number, and the step of using the predicted data volume at the positioning point of the audio data as the predicted positioning point is returned to be executed until the candidate frame number is consistent with the target frame number, and the target audio is played at the fixed point based on the latest predicted data volume. On the one hand, since the target frame number and the initial predicted data volume are both generated based on the first timestamp and the target audio metadata, the difference between the candidate frame number and the target frame number is as small as possible, which can improve the efficiency of the adjustment and update process. On the other hand, based on the audio format of the target audio, the target frame number corresponding to the first timestamp is obtained, and then the target frame number is used as reference information to adjust and update the predicted data volume until the candidate frame number corresponding to the audio data positioning point is consistent with the target frame number. In the case of uneven distribution of the audio data of the target audio, the accuracy of the predicted data volume for fixed-point playback of the target audio can be improved. Therefore, the accuracy of the fixed-point playback of the target audio can be improved while improving the efficiency of fixed-point playback as much as possible.
[0054] See also Figure 5 , Figure 5 1 is a schematic diagram of a computer readable storage medium 50 / computer program product 50 of the present application. The computer readable storage medium 50 stores computer program instructions 51, and the computer program product 50 includes computer program instructions 51, which are used to implement the steps in any of the above-mentioned audio fixed-point playback method embodiments.
[0055] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0056] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.
[0057] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0058] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed over multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0059] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0060] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0061] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
Claims
1. A method for fixed-point audio playback, characterized in that: include: Based on metadata of target audio and a first timestamp indicating fixed-point playback of the target audio, a target frame number and a predicted data amount are determined; wherein the target audio at least includes audio data, the metadata is used to define audio parameters, the audio data includes a number of audio frames, the target frame number represents a frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data amount represents an amount of audio data predicted to be present in the audio data as of the first timestamp; Selecting the frame number corresponding to the predicted data amount at the positioning point of the audio data as a candidate frame number; In response to the candidate frame number being inconsistent with the target frame number, updating the predicted data amount based on the difference between the candidate frame number and the target frame number, and returning to the step of selecting the frame number corresponding to the predicted data amount at the positioning point of the audio data as the candidate frame number until the candidate frame number is consistent with the target frame number; Based on the latest predicted data amount, the target audio is played at a fixed point.
2. The method according to claim 1, characterized in that The updating of the predicted data amount based on the difference between the candidate frame sequence number and the target frame sequence number comprises: Determining a data offset based on a difference between the candidate frame sequence number and the target frame sequence number; In response to the candidate frame sequence number being greater than the target frame sequence number, obtaining the updated predicted data amount based on a difference between the predicted data amount and the data offset; In response to the candidate frame number being smaller than the target frame number, the updated predicted data amount is obtained based on the sum of the predicted data amount and the data offset.
3. The method according to claim 2, characterized in that Before determining the data offset based on the difference between the candidate frame sequence number and the target frame sequence number, the method further includes: Based on the audio parameters in the metadata used to define the data size of each audio frame of the target audio, obtaining the maximum data size of each audio frame in the audio parameters; The determining of the data offset based on the difference between the candidate frame sequence number and the target frame sequence number comprises: The data offset is obtained based on half of the product of the target difference and the maximum data amount; wherein the target difference represents the difference between the candidate frame sequence number and the target frame sequence number.
4. The method according to claim 1, characterized in that: Determining the predicted data amount based on the metadata includes: Detecting whether there is a positioning table for the target audio in the metadata; wherein the positioning table includes at least one mapping relationship between a preset sampling number and an amount of audio data; In response to the absence of the positioning table for the target audio, obtaining the predicted data volume based on the audio parameters representing the total duration of the target audio and the total data volume of the target audio in the metadata and the first timestamp; In response to the existence of the positioning table for the target audio, the predicted data amount is obtained based on the audio parameters representing the positioning table and the audio sampling rate in the metadata, and the first timestamp.
5. The method according to claim 4, characterized in that The obtaining the predicted data amount based on the audio parameter representing the positioning table and the audio sampling rate in the metadata and the first timestamp includes: Obtaining a target number of samples corresponding to the first timestamp based on a product of the audio sampling rate and the first timestamp; Selecting a mapping relationship of the preset sampling number in the positioning table that is closest to the target sampling number; The audio data amount corresponding to the mapping relationship of the closest preset sampling number is used as the predicted data amount.
6. The method according to claim 1, characterized in that The step of playing the target audio at a fixed point based on the latest predicted data amount includes: Obtaining a second timestamp corresponding to the latest predicted data amount in the audio data; Based on the positioning point of the latest predicted data amount in the audio data, the target audio is played, and the current timestamp of the target audio is positioned to the second timestamp.
7. An audio fixed-point playback device, characterized in that: include: A determination module, configured to determine a target frame number and a predicted data amount based on metadata of a target audio and a first timestamp indicating fixed-point playback of the target audio; wherein the target audio at least includes audio data, the metadata is used to define audio parameters, the audio data includes a number of audio frames, the target frame number represents a frame number of the audio frame corresponding to the first timestamp in the audio data, and the predicted data amount represents an amount of audio data predicted to be present in the audio data as of the first timestamp; A positioning module, used for selecting the frame number corresponding to the positioning point of the audio data with the predicted data amount as a candidate frame number; an updating module, configured to update the predicted data amount based on a difference between the candidate frame number and the target frame number in response to the candidate frame number being inconsistent with the target frame number, and return to the step of selecting the frame number corresponding to the positioning point of the audio data with the predicted data amount as the candidate frame number until the candidate frame number is consistent with the target frame number; A playing module is used to play the target audio at a fixed point based on the latest predicted data volume.
8. An electronic device, characterized in that: It comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the audio fixed-point playback method according to any one of claims 1 to 6.
9. An audio playback system, characterized in that: It includes an audio playback application layer and an audio playback processing layer, wherein the audio playback application layer is used to obtain a fixed-point playback instruction for a target audio and display a timestamp of the fixed-point playback of the target audio, and the audio playback processing layer is the electronic device described in claim 8, and the audio playback processing layer is used to play the target audio at a fixed point.
10. A computer-readable storage medium / computer program product, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the audio fixed-point playback method according to any one of claims 1 to 6 is implemented; The computer program product comprises computer program instructions, and when the computer program instructions are executed by a processor, the audio fixed-point playback method according to any one of claims 1 to 6 is implemented.