A method to extract amr audio files from damaged storage devices
By analyzing the amr file format and data composition, obtaining cluster byte lengths and classifying them, and combining the amr data frame characteristics, the dirty data problem when extracting amr audio files in the prior art is solved, and more accurate file recovery is achieved.
Patent Information
- Application Number
- CN202210740276.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-06-27
AI Technical Summary
The prior art is prone to extract dirty data and miss useful data when extracting amr audio files from damaged storage devices.
By analyzing the format and data composition of the amr file, the byte length of the cluster is obtained, and each cluster is classified. Combined with the characteristics of the amr data frame, the playable amr audio file is extracted.
It realizes more accurate recovery of amr audio files, avoids the extraction of dirty data, and obtains more effective data.
Smart Images

Figure CN115114090B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of data recovery and relates to a method for extracting an AMR audio file from a damaged storage device. Background Art
[0002] AMR is an audio file format mainly used on mobile devices. Since AMR files occupy less resources and have a small capacity, the size of AMR files per second can be controlled at around 1K, which is convenient for sending recordings and MMS, and therefore complies with the technical specification that the size of MMS in China should not exceed 50K. Most of the recordings on mobile phones on the market are in AMR format. For example, AMR was once the ringtone audio file format in Nokia phones, and the recording files in Nokia phones also use the AMR format. Therefore, AMR audio is widely used in life.
[0003] If the storage device storing the AMR is damaged, the recovery method in the prior art is to directly obtain the file through the AMR header and the length of the header record (that is, what we usually call extraction based on the file signature). The AMR audio file extracted by this method will extract a lot of dirty data, and at the same time, it will also miss a lot of useful data.
[0004] This method studies the format of AMR, the composition of data, and the characteristics of each data frame. It first confirms the correct cluster size, then classifies each cluster, and extracts a method for playing AMR based on the composition of AMR files and the characteristics of AMR frames. Summary of the invention
[0005] In view of the technical problems of the prior art, the present invention provides a method for extracting an amr audio file from a damaged storage device: by analyzing the format of the amr and the composition of the data, the byte length of the cluster is first obtained, and then each cluster is classified, and the amr audio file is extracted in combination with the characteristics of the amr data frame.
[0006] The present invention comprises the following steps:
[0007] S100: Obtaining the byte length of the cluster, including the following steps:
[0008] S101: traverse N clusters starting with the amr header identifier and record the starting address of each cluster, where N is a natural number not less than 1;
[0009] S102: Calculate and record the difference between the start addresses of two adjacent clusters in each cluster;
[0010] S103: Obtain the minimum value of the difference between the N-1 starting addresses in step S102 as the byte length of the cluster;
[0011] S200: Obtaining the address of the first cluster, including the following steps:
[0012] S201: Get the address of the first cluster starting with the amr header identifier;
[0013] S202: using the address of the first cluster marked with the amr header as the starting address and the byte length of the cluster as the offset, addressing each cluster forward until the address of the first cluster is addressed;
[0014] S300: Classify the amr data in the storage device and store them into various map data structures according to the classification, traverse each map data structure and extract the amr audio file, wherein the amr data includes an amr header and an amr frame, and accordingly, the map data structure includes a header map storage and a data map storage, and the step S300 includes the following steps:
[0015] S3000: Addressing the first cluster;
[0016] S3001: Determine whether the extraction of all clusters is completed, if yes, execute step S3008, otherwise, execute step S3002;
[0017] S3002: Read the content of the current cluster according to the byte length of the cluster and the address of the current cluster;
[0018] S3003: Determine whether the current cluster has an AMR header, if yes, execute step S3004, otherwise, execute step S3005;
[0019] S3004: store the AMR header in a map data structure, recorded as header map storage;
[0020] Using the offset address of the current cluster in the storage device as the keyword and the AMR header information as the key value, add the current keyword and the current key value to the header map storage, address the next cluster and execute step S3001, wherein the AMR header information includes the value of the frame header in the current cluster and the byte length missing from the last frame in the current cluster;
[0021] S3005: Determine whether the current cluster has an AMR frame, if yes, execute step S3006, otherwise, execute step S3007;
[0022] S3006: storing the amr frame in a map data structure, recorded as data map storage;
[0023] Using the offset address of the current cluster in the storage device as the keyword and the amr frame information as the key value, add the current keyword and the current key value to the data map storage, address the next cluster and execute step S3001, wherein the amr frame information includes the value of the frame header, the starting address of the first frame, the byte length of the last frame of the current cluster missing in the current cluster, and the value of the file end mark;
[0024] S3007: address the next cluster and execute step S3001;
[0025] S3008: traverse the map data structure and extract the amr audio file.
[0026] Preferably, N is 20.
[0027] Preferably, determining whether the current cluster has an AMR header in step S3003 includes the following steps:
[0028] A: Whether the current cluster starts with the amr header identifier;
[0029] B: Whether to store AMR frames continuously after the AMR header mark.
[0030] Preferably, the AMR header identifier is #! AMR.
[0031] Preferably, the file end identifier in step S3006 is bIsEnd, and the value of bIsEnd is true, indicating the end of the file, and vice versa.
[0032] Preferably, step S3008 includes the following steps:
[0033] S30081: Determine whether the header map storage has been traversed. If yes, end the process. Otherwise, execute S30082.
[0034] S30082: read a cluster from the head map storage and delete the currently read cluster from the head map storage, and store the currently read cluster to the output temporary file;
[0035] S30083: Determine whether the data map storage has been traversed. If yes, execute step S30087; otherwise, execute step S30084.
[0036] S30084: Read a cluster from the data map storage;
[0037] S30085: Determine whether the next cluster is found, if so, execute step S30086, otherwise, address the next cluster and execute step S30083;
[0038] S30086: Determine whether the file end flag in the current cluster read from the data map storage is true, if so, execute step S30087, otherwise, execute step S30088;
[0039] S30087: Add the tail of the current output temporary file to the amr audio file, and execute step S30081;
[0040] S30088: Store the current cluster to the output temporary file, delete the current cluster from the data map storage, and execute step S30083.
[0041] Preferably, determining whether the next cluster is found in step S30085 includes the following steps:
[0042] 1. The current cluster and the next cluster of the current cluster have the same frame header;
[0043] 2. The byte length of the last frame of the current cluster missing in the current cluster is equal to the byte length of the first frame of the next cluster of the current cluster.
[0044] The present invention has the following beneficial effects:
[0045] 1. Recovering the amr audio file through the amr data storage format is more accurate than recovering it through the signature method.
[0046] 2. Based on the characteristics of files stored in the file system, the cluster size is recalculated instead of the default value, making the recovered amr audio file more accurate.
[0047] 3. By analyzing the characteristics of the AMR frame, the cluster storing the AMR data is found, and more valid AMR audio files can be obtained compared with the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A flow chart of the method provided by the present invention;
[0049] Figure 2 This is an example diagram of the data structure of the AMR header identifier in an embodiment provided by the present invention;
[0050] Figure 3 A specific flow chart of classifying amr data and storing them in various map data structures according to the classification and extracting amr audio files in the method provided by the present invention;
[0051] Figure 4 This is a specific flow chart of traversing the map data structure and extracting the amr audio file in the method provided by the present invention. DETAILED DESCRIPTION
[0052] Figure 1 The flow chart of the method provided by the present invention is shown. Figure 1 As shown, the method of the present invention comprises the following steps:
[0053] S100: Obtaining the byte length of the cluster, including the following steps:
[0054] S101: traverse N clusters starting with the amr header identifier and record the starting address of each cluster, where N is a natural number not less than 1; in this embodiment, we set N to 20.
[0055] In addition, the AMR header identifier is #! AMR. Specifically, the AMR header identifier is divided into two types: narrowband AMR-NB and wideband AMR-WB. Among them, the AMR header identifier of narrowband AMR-NB is #! AMR; the AMR header identifier of wideband AMR-WB is #! AMR_MC1.0. Both start with #! AMR.
[0056] Figure 2 FIG. 2 shows an example of a data structure of an AMR header identifier in an embodiment provided by the present invention. Figure 2 As shown in FIG. 1 , the byte content of the address 0x00000000 to 0x00000005 is 0x2321414D520A, indicating the hexadecimal representation of the ASCII code of the AMR header identifier #! AMR of the narrowband AMR-NB.
[0057] S102: Calculate and record the difference between the start addresses of two adjacent clusters in each cluster;
[0058] S103: Obtain the minimum value of the difference between the N-1 (19 in this embodiment) starting addresses in step S102 as the byte length of the cluster;
[0059] S200: Obtaining the address of the first cluster, including the following steps:
[0060] S201: Get the address of the first cluster starting with the amr header identifier;
[0061] S202: using the address of the first cluster marked with the amr header as the starting address and the byte length of the cluster as the offset, addressing each cluster forward until the address of the first cluster is addressed;
[0062] S300: Classify the amr data in the storage device and store them into various map data structures according to the classification, traverse each map data structure and extract the amr audio file, wherein the amr data includes an amr header and an amr frame, and correspondingly, the map data structure includes a header map storage and a data map storage.
[0063] Figure 3The specific flow chart of the method provided by the present invention for classifying the amr data and storing them in various map data structures according to the classification and extracting the amr audio files is shown. Figure 3 As shown, step S300 includes the following steps:
[0064] S3000: Addressing the first cluster;
[0065] S3001: Determine whether the extraction of all clusters is completed, if yes, execute step S3008, otherwise, execute step S3002;
[0066] S3002: Read the content of the current cluster according to the byte length of the cluster and the address of the current cluster;
[0067] S3003: Determine whether the current cluster has an AMR header, if yes, execute step S3004, otherwise, execute step S3005;
[0068] Specifically, determining whether the current cluster has an amr header includes the following steps:
[0069] A: Whether the current cluster starts with the amr header identifier;
[0070] B: Whether to store AMR frames continuously after the AMR header mark.
[0071] S3004: store the AMR header in a map data structure, recorded as header map storage;
[0072] Using the offset address of the current cluster in the storage device as the keyword and the AMR header information as the key value, add the current keyword and the current key value to the header map storage, address the next cluster and execute step S3001, wherein the AMR header information includes the value of the frame header in the current cluster and the byte length missing from the last frame in the current cluster;
[0073] S3005: Determine whether the current cluster has an AMR frame, if yes, execute step S3006, otherwise, execute step S3007;
[0074] S3006: storing the amr frame in a map data structure, recorded as data map storage;
[0075] Using the offset address of the current cluster in the storage device as the keyword and the amr frame information as the key value, the current keyword and the current key value are added to the data map storage, addressing the next cluster and executing step S3001, wherein the amr frame information includes the value of the frame header, the starting address of the first frame, the byte length of the last frame of the current cluster missing in the current cluster, and the value of the file end identifier; wherein the file end identifier is bIsEnd, and the value of bIsEnd is true, indicating the end of the file, and vice versa.
[0076] S3007: address the next cluster and execute step S3001;
[0077] S3008: traverse the map data structure and extract the amr audio file.
[0078] Figure 4 The specific flow chart of traversing the map data structure and extracting the amr audio file in the method provided by the present invention is shown. Figure 4 As shown, step S3008 includes the following steps:
[0079] S30081: Determine whether the header map storage has been traversed. If yes, end the process. Otherwise, execute S30082.
[0080] S30082: read a cluster from the head map storage and delete the currently read cluster from the head map storage, and store the currently read cluster to the output temporary file;
[0081] S30083: Determine whether the data map storage has been traversed. If yes, execute step S30087; otherwise, execute step S30084.
[0082] S30084: Read a cluster from the data map storage;
[0083] S30085: Determine whether the next cluster is found. If so, execute step S30086. Otherwise, address the next cluster and execute step S30083. Specifically, determining whether the next cluster is found in step S30085 includes the following steps:
[0084] 1. The current cluster and the next cluster of the current cluster have the same frame header;
[0085] 2. The byte length of the last frame of the current cluster missing in the current cluster is equal to the byte length of the first frame of the next cluster of the current cluster.
[0086] S30086: Determine whether the file end flag in the current cluster read from the data map storage is true, if so, execute step S30087, otherwise, execute step S30088;
[0087] S30087: Add the tail of the current output temporary file to the amr audio file, and execute step S30081;
[0088] S30088: Store the current cluster to the output temporary file, delete the current cluster from the data map storage, and execute step S30083.
[0089] The method provided by the present invention solves the technical problem that there is no method for extracting amr audio files from a damaged storage device in the prior art.
[0090] It should be understood that the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A method for extracting amr audio files from a damaged storage device, characterized in that The following steps are involved: S100: Obtaining the byte length of the cluster, including the following steps: S101: traverse N clusters starting with the amr header identifier and record the starting address of each cluster, where N is a natural number not less than 1; S102: Calculate and record the difference between the start addresses of two adjacent clusters in each cluster; S103: Obtain the minimum value of the difference between the N-1 starting addresses in step S102 as the byte length of the cluster; S200: Obtaining the address of the first cluster, including the following steps: S201: Get the address of the first cluster starting with the amr header identifier; S202: using the address of the first cluster marked with the amr header as the starting address and the byte length of the cluster as the offset, addressing each cluster forward until the address of the first cluster is addressed; S300: Classify the amr data in the storage device and store them into various map data structures according to the classification, traverse each map data structure and extract the amr audio file, wherein the amr data includes an amr header and an amr frame, and accordingly, the map data structure includes a header map storage and a data map storage, and the step S300 includes the following steps: S3000: Addressing the first cluster; S3001: Determine whether the extraction of all clusters is completed, if yes, execute step S3008, otherwise, execute step S3002; S3002: Read the content of the current cluster according to the byte length of the cluster and the address of the current cluster; S3003: Determine whether the current cluster has an AMR header, if yes, execute step S3004, otherwise, execute step S3005; S3004: store the AMR header in a map data structure, recorded as header map storage; Using the offset address of the current cluster in the storage device as the keyword and the AMR header information as the key value, add the current keyword and the current key value to the header map storage, address the next cluster and execute step S3001, wherein the AMR header information includes the value of the frame header in the current cluster and the byte length missing from the last frame in the current cluster; S3005: Determine whether the current cluster has an AMR frame, if yes, execute step S3006, otherwise, execute step S3007; S3006: storing the amr frame in a map data structure, recorded as data map storage; Using the offset address of the current cluster in the storage device as the keyword and the amr frame information as the key value, add the current keyword and the current key value to the data map storage, address the next cluster and execute step S3001, wherein the amr frame information includes the value of the frame header, the starting address of the first frame, the byte length of the last frame of the current cluster missing in the current cluster, and the value of the file end mark; S3007: address the next cluster and execute step S3001; S3008: traverse the map data structure and extract the amr audio file.
2. A method for extracting an amr audio file from a damaged storage device according to claim 1, characterized in that: N is 20.
3. A method for extracting an amr audio file from a damaged storage device according to claim 1, characterized in that: In step S3003, determining whether the current cluster has an AMR header includes the following steps: A: Whether the current cluster starts with the amr header identifier; B: Whether to store AMR frames continuously after the AMR header mark.
4. A method for extracting an amr audio file from a damaged storage device according to claim 3, characterized in that: The AMR header is identified by #! AMR.
5. A method for extracting an amr audio file from a damaged storage device according to claim 1, characterized in that: The file end identifier in step S3006 is bIsEnd. The value of bIsEnd is true, indicating the end of the file, and vice versa.
6. A method for extracting an amr audio file from a damaged storage device according to claim 1, characterized in that: Step S3008 includes the following steps: S30081: Determine whether the header map storage has been traversed. If yes, end the process. Otherwise, execute S30082. S30082: read a cluster from the head map storage and delete the currently read cluster from the head map storage, and store the currently read cluster to the output temporary file; S30083: Determine whether the data map storage has been traversed. If yes, execute step S30087; otherwise, execute step S30084. S30084: Read a cluster from the data map storage; S30085: Determine whether the next cluster is found, if so, execute step S30086, otherwise, address the next cluster and execute step S30083; S30086: Determine whether the file end flag in the current cluster read from the data map storage is true, if so, execute step S30087, otherwise, execute step S30088; S30087: Add the tail of the current output temporary file to the amr audio file, and execute step S30081; S30088: Store the current cluster to the output temporary file, delete the current cluster from the data map storage, and execute step S30083.
7. A method for extracting an amr audio file from a damaged storage device according to claim 6, characterized in that: In step S30085, determining whether the next cluster is found includes the following steps:
1. The current cluster and the next cluster of the current cluster have the same frame header; 2. The byte length of the last frame of the current cluster missing in the current cluster is equal to the byte length of the first frame of the next cluster of the current cluster.
Citation Information
Patent Citations
Indexing system and method supporting MOV (Movie Digital Video Technology) / 3GP (3G Player) / MP4 (Mobile Pentium 4) files
CN102682016A
Method for restoring audio files of mobile phone
CN105630633A