Chat record recovery method based on WeChat synchronous cache file

By analyzing the data structure of WeChat synchronous cache files and nesting and recursively parsing data blocks, efficient recovery of WeChat chat records is achieved, solving the problem of incomplete recovery in the existing technology and avoiding system corruption.

CN120256201APending Publication Date: 2025-07-04XLY SALVATIONDATA TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510391298.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art cannot efficiently restore WeChat chat records, especially in the event of no backup, and third-party software may cause system damage or incomplete recovery.

Method used

Analyze the data structure of WeChat synchronous cache files, parse data blocks through nesting and recursive methods, extract the type and content of chat records, and restore chat records.

Benefits of technology

It provides an efficient method without additional investment, which can restore WeChat chat history and avoid the risk of system damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256201A_ABST
    Figure CN120256201A_ABST
Patent Text Reader

Abstract

The invention discloses a chat record recovery method based on a WeChat synchronous cache file. The chat record recovery method is characterized by comprising the following steps: S100, acquiring the WeChat synchronous cache file; s200, judging whether a data block exists or not, if yes, executing the step S300, and if not, executing the step S600; s300, acquiring an initial offset address of the current data block, and calculating a first byte length and a second byte length; s400, according to the first byte length and the second byte length, reading the current data block and counting basic information of the current data block into a first set; s500, updating the initial offset address, and executing the step S200; s600, analyzing the data structure of the data block in a nested and recursive manner until each subset does not contain any subset; s700, storing the type value and the content of the chat record in a sixth layer of the first subset; and S800, according to the type of the chat record, analyzing and recovering the chat record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data recovery and forensics, and relates to a method for recovering chat records based on WeChat synchronization cache files. Background Art

[0002] With the rise of smart phones, WeChat is a masterpiece of instant messaging and is closely related to daily life.

[0003] During the use of WeChat, it is often encountered that chat records are deleted and need to be recovered. Therefore, the recovery of chat records is crucial.

[0004] If the mobile phone has not been backed up, or the "fault repair" of WeChat cannot recover chat records, several third-party data recovery software for recovering WeChat chat records in the prior art can be used. However, due to the frequent updates of WeChat and multiple upgrades and modifications with the changes in the market, the data recovery solution should also be updated accordingly. In addition, these software may cause damage to the mobile phone system, and the data recovery ability is affected by various factors, such as the time of data deletion, the degree of data coverage, the condition of the storage medium, etc. Therefore, even if third-party data recovery software is used, it cannot guarantee 100% recovery of deleted chat records. Summary of the Invention

[0005] In view of the deficiencies of the prior art, starting from the synchronization data file of WeChat, the present invention analyzes the data structure of the WeChat synchronization cache file and provides a method for recovering chat records based on the WeChat synchronization cache file, including the following steps:

[0006] A method for recovering chat records based on the WeChat synchronization cache file, characterized by including the following steps:

[0007] S100: Extract the mirror data of WeChat and obtain the WeChat synchronization cache file;

[0008] S200: Determine whether there is a data block. If so, execute step S300; otherwise, execute step S600;

[0009] S300: Obtain the starting offset address of the current data block, and calculate the first byte length and the second byte length;

[0010] S400: Read the current data block according to the first byte length and the second byte length, and record the basic information of the current data block in the first set, where the basic information includes the starting offset address and the second byte length;

[0011] S500: Update the starting offset address, and execute step S200;

[0012] S600: Parse the data structure of the data block in a nested and recursive manner until each subset no longer contains any subsets;

[0013] S700: Store the type value and content of the chat record in the sixth layer of the first subset, where the fourth item in the sixth layer represents the type of the chat record, the fifth item represents the content of the chat record, and the eighth item represents the attachment of the chat record;

[0014] S800: Parse and restore the chat record according to the type of the chat record.

[0015] Preferably, step S100 includes the following steps:

[0016] Obtain the WeChat synchronization cache file under the mirror data directory com.tencent.mm\files\mmkv of WeChat, and its file name starts with SyncMMKV.

[0017] Preferably, step S200 includes the following steps:

[0018] The method for determining whether there is a data block includes: querying whether the WeChat synchronization cache file contains the keyword

[0019] 0x1170726F636573735F646174615F6C69737401001170726F636573735F646174615F6C697374.

[0020] Preferably, step S300 includes the following steps:

[0021] S301: Record the starting address of the current keyword, and the value obtained by adding 0x26 to the starting address of the current keyword is used as the starting offset address of the current data block;

[0022] S302: Address the starting offset address and set the initial value of the first value to 0;

[0023] S303: Read the content of one byte at the current address;

[0024] S304: Determine whether the highest bit of the current byte is 1. If it is, execute step S305; otherwise, execute step 306;

[0025] S305: Increment the first value by 1 and address the next byte, then execute step S304;

[0026] S306: Increment the first value by 1, remove the highest bits of all the read bytes, store them in little-endian format, and pad with zeros if the high bits are insufficient. The obtained value is used as the length of the first byte;

[0027] S307: Update the starting offset address to the sum of the starting offset address and the current first value, and set the initial value of the second value to 0;

[0028] S308: Address the starting offset address;

[0029] S309: Read the content of one byte at the current address;

[0030] S310: Determine whether the highest bit of the current byte is 1. If it is, execute step S311; otherwise, execute step 312;

[0031] S311: Increment the second value by 1, address the next byte, and execute step S310;

[0032] S312: Increment the second value by 1. After removing the highest bits of all the read bytes, store them in little-endian format, padding with zeros if the high bits are insufficient. The resulting value is used as the second byte length;

[0033] S313: Determine whether the first byte length is equal to the sum of the second byte length and the second value. If it is, execute step S400; otherwise, execute step S314:

[0034] S314: Increment the starting offset address by 1 and execute step S200.

[0035] Preferably, step S500 includes the following steps:

[0036] Update the starting offset address to the sum of the current starting offset address and the first byte length.

[0037] Preferably, step S800 includes the following steps:

[0038] S801: Read the value of the fourth item in the sixth layer of the first subset;

[0039] S802: Determine whether the fourth item in the sixth layer of the first subset is one of the following values:

[0040] If the value is 1, it indicates text, and execute step S803;

[0041] If the value is 3, it indicates a picture, and execute step S804;

[0042] If the value is 34, it indicates voice, and execute step S804;

[0043] If the value is 43 or 62, it indicates video, and execute step S804;

[0044] S803: Parse the text content and output the chat record of the text type. Among them, the text content is stored in the first entry of the fifth item in the sixth layer of the first subset, and execute step S805;

[0045] S804: Parse the attachment and output the chat record of non - text type. Among them, the attachment content is stored in the eighth item of the sixth layer of the first subset, which contains two entries. The first entry stores the length of the media file, and the second entry stores the underlying binary data of the media file;

[0046] S805: Output the restored chat record and end the process.

[0047] The beneficial effects of the present invention are as follows: When the chat record is deleted, based on the WeChat synchronization cache file, a method is provided that can efficiently restore the chat record without any additional investment. Description of the Drawings

[0048] Figure 1 It is the overall flow chart of the present invention;

[0049] Figure 2A 、 Figure 2B It is the instance diagram of the data structure of the database in the embodiment of the present invention;

[0050] Figure 3 It is the flow chart of calculating the length of the first byte and the length of the second byte in the embodiment of the present invention;

[0051] Figure 4 It is the instance diagram of the first set and the first subset in the embodiment of the present invention. Detailed Embodiments

[0052] The present invention will be further described below in conjunction with the drawings and embodiments.

[0053] As Figure 1 shown, the method of the present invention includes the following steps:

[0054] S100: Extract the mirror data of WeChat and obtain the WeChat synchronization cache file.

[0055] Obtain the WeChat synchronization cache file in the mirror data directory com.tencent.mm\files\mmkv of WeChat. Its file name starts with SyncMMKV, for example: SyncMMKV_uin, where uin is the custom id of WeChat. In this embodiment, the WeChat synchronization cache file is SyncMMKV_3408357621.

[0056] S200: Determine whether there is a data block process_data_list. If so, execute step S300; otherwise, execute step S600.

[0057] The method for determining whether there is a data block process_data_list is as follows: Query whether the WeChat synchronization cache file contains the keyword 0x1170726F636573735F646174615F6C69737401001170726F636573735F646174615F6C697374.

[0058] Figure 2A The figure shows an example of the data structure of the database in the embodiment of the present invention (only part of the data is shown). As Figure 2A shown in the shaded part,

[0059] 0x1170726F636573735F646174615F6C69737401001170726F636573735F646174615F6C697374 is the keyword to be queried;

[0060] Among them, 0x70726F636573735F646174615F6C697374 is the ASCII code corresponding to process_data_list.

[0061] S300: Obtain the starting offset address offset of the current data block, and calculate the first byte length Length and the second byte length length.

[0062] Figure 3 The figure shows the flowchart for calculating the first byte length and the second byte length in the embodiment of the present invention.

[0063] As Figure 3 shown, step S300 includes the following steps:

[0064] S301: Record the starting address 0x08 of the current keyword. The value obtained by adding 0x26 to the starting address of the current keyword is 0x2E, which is used as the starting offset address offset of the current data block, as Figure 2A shown.

[0065] S302: Address the starting offset address offset and set the initial value of the first value Size to 0;

[0066] S303: Read the content of one byte at the current address 0x2E, which is 0xB6, as Figure 2A shown by the underlined part in the figure;

[0067] S304: Determine whether the highest bit of the current byte is 1. If it is, execute step S305; otherwise, execute step 306;

[0068] Taking this embodiment as an example, as Figure 2AAs shown, the current byte is 0xB6, and its binary representation is 10110110. Since its most significant bit is 1, step 305 needs to be executed;

[0069] S305: Increment the first value Size by 1 and address the next byte, then execute step S304;

[0070] In this embodiment, after incrementing the first value Size by 1, its value becomes 1, and the address of the next byte is 0x2F.

[0071] As Figure 2A shown, at this time, the content of the byte at address 0x2F is 0xFE (as shown by the underlined part in Figure 2A ). Its binary representation is 11111110, and its most significant bit is 1. Therefore, step 305 is executed again. After incrementing the first value Size by 1, its value becomes 2, and the address of the next byte is 0x030;

[0072] The content of the current byte at address 0x030 is 0x01 (as shown by the underlined part in Figure 2A ). Its binary representation is 00000001, and its most significant bit is 0. Step 306 is executed.

[0073] S306: Increment the first value Size by 1. After removing the most significant bit of all the read bytes, store them in little - endian format, padding with zeros if the high - order bits are insufficient. The resulting value is used as the first byte length Length;

[0074] After removing the most significant bit of the three read bytes, their values are 0110110, 1111110, and 0000001 respectively. Stored in little - endian format, it is 000000111111100110110. After padding three zeros for the insufficient high - order bits, the value is 00000000011111100110110, which is 0x7F36. This value is the first byte length Length.

[0075] S307: Update the starting offset address offset to the sum of the starting offset address offset and the current first value Size, and set the initial value of the second value size to 0;

[0076] As Figure 2A shown, in this embodiment, the starting offset address offset is 0x2E, and the current Size is 0x3. The updated starting offset address offset is 0x2E + 0x3 = 0x31;

[0077] S308: As Figure 2A shown, address the starting offset address offset, that is, address 0x31;

[0078] S309: As Figure 2AAs shown, the content of one byte read from the current address 0x31 is 0xB3, as Figure 2A shown in the rectangular box part of

[0079] S310: Determine whether the highest bit of the current byte is 1. If it is, execute step S311; otherwise, execute step 312;

[0080] Taking this embodiment as an example, as Figure 2A shown, the current byte is 0xB3, and its binary representation is 10110011. Its highest bit is 1. Therefore, step 311 needs to be executed;

[0081] S311: Increment the second value size by 1 and address the next byte, then execute step S310;

[0082] In this embodiment, after the first value size is incremented by 1, its value is 1, and the address of the next byte is 0x32, as Figure 2A shown.

[0083] As Figure 2A shown, at this time, the content of the byte at address 0x32 is 0xFE (as Figure 2A shown in the rectangular box part), and its binary representation is 11111110. Its highest bit is 1. Therefore, step 311 is executed again. After the first value size is incremented by 1, its value is 2, and the address of the next byte is 0x033;

[0084] The content of the current byte at address 0x033 is 0x01 in binary (as Figure 2A shown in the rectangular box part), which is represented as 00000001. Its highest bit is 0, so execute step 312.

[0085] S312: Increment the second value size by 1, remove the highest bits of all the read bytes, store them in little - endian format, and pad with zeros if the high - order bits are insufficient. The resulting value is used as the second byte length length;

[0086] After removing the highest bits of the three read bytes, their values are 0110011, 1111110, and 0000001 respectively. Stored in little - endian format, it is 000000111111100110110. After padding three zeros for the insufficient high - order bits, the value is 00000000011111100110011, which is 0x7F33. This value is the second byte length length.

[0087] S313: Determine whether the first byte length Length is equal to the sum of the second byte length length and the second value size. If it is, execute step S400; otherwise, execute step S314:

[0088] In this embodiment, the length of the first byte Length = the length of the second byte length + the second value size, that is, 0x7F36 = 0x7F33 + 0x3. Since the judgment condition is satisfied, step S400 is executed.

[0089] S314: As a processing method when the judgment condition is false: the starting offset address offset is incremented by 1, and step S200 is executed.

[0090] S400: According to the length of the first byte Length and the length of the second byte length, the current data block is read and the basic information of the current data block is recorded in the first set Set, where the basic information includes the starting offset address offset and the length of the second byte length.

[0091] In this embodiment, the updated starting offset address offset (i.e., 0x31) is used as the starting address to read data, and the byte length of the read data is the length of the first byte Length (i.e., 0x7F36); in other words, with the updated starting offset address offset (i.e., 0x31) as the starting address and the length of the first byte Length (i.e., 0x7F36) as the offset, the data between address 0x31 and address 0x7F36 is read; the starting offset address offset, the length of the second byte length, and the read data are recorded in the first set Set.

[0092] Figure 2B Another example diagram of the data structure of the database in the embodiment of the present invention is shown (only part of the data is shown). That is, a diagram of the data structure of the data read by using the updated starting offset address offset (i.e., 0x31) as the starting address, with the byte length as the length of the first byte Length (i.e., 0x7F36) as the offset, addressing to the last byte of the read data (partially shown) and the data structure of the next keyword, as Figure 2B The part shown in the rounded rectangle box in is the next keyword.

[0093] It should be understood that data can also be read using the length of the second byte length, which will not be elaborated here.

[0094] S500: Update the starting offset address offset, and execute step S200;

[0095] The starting offset address offset is updated to the sum of the current starting offset address offset and the length of the first byte Length.

[0096] S600: Parse the data structure of the data block process_data_list in a nested and recursive manner until each subset no longer contains any subsets;

[0097] Parse the data structure of the data block process_data_list through open-source tools, so no more details about the structure parsing will be given.

[0098] S700: Store the type value and content of the chat record in the sixth layer of the first subset. Among them, the fourth item in the sixth layer represents the type of the chat record, the fifth item represents the content of the chat record, and the eighth item represents the attachment of the chat record;

[0099] Figure 4 Shows an example diagram of the first set and the first subset in the embodiments of the present invention. As Figure 4 shown, among them,

[0100] The value of the fourth item in the sixth layer is 34, indicating that the type of the chat record is voice;

[0101] The fifth item represents the content of the chat record, which includes an end mark, a cancel mark, a forward mark, a voice format, a byte length, a voice URL, a key of the Advanced Encryption Standard, etc.;

[0102] The eighth item represents the attachment of the chat record.

[0103] S800: Parse and restore the chat record according to the type of the chat record, including the following steps:

[0104] S801: Read the value of the fourth item in the sixth layer of the first subset;

[0105] S802: Determine whether the fourth item in the sixth layer of the first subset is one of the following values:

[0106] If the value is 1, it represents text, and execute step S803;

[0107] If the value is 3, it represents a picture, and execute step S804;

[0108] If the value is 34, it represents voice, and execute step S804;

[0109] If the value is 43 or 62, it represents a video, and execute step S804;

[0110] S803: Parse the text content and output the chat record of the text type. Among them, the text content is stored in the first entry of the fifth item in the sixth layer of the first subset, and execute step S805;

[0111] S804: Parse the attachment and output the chat record of the non-text type. Among them, the attachment content is stored in the eighth item in the sixth layer of the first subset, which contains two entries. The first entry stores the length of the media file, and the second entry stores the underlying binary data of the media file;

[0112] AsFigure 4 As shown, in this embodiment, the chat record is voice, and the attachment content is stored in the eighth item of the sixth layer of the first subset, which contains two entries. The first entry stores the length of the media file, and its value is 3953 bytes. The second entry stores the data of the voice file in binary representation.

[0113] S805: Output the restored chat record and end the process.

[0114] Through the method provided by the present invention, the technical problem that there is no chat record recovery method based on WeChat synchronization cache files in the prior art after the WeChat chat record is deleted is solved.

[0115] It should be understood that the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for recovering chat records based on WeChat synchronized cache files, characterized in that It includes the following steps: S100: Extract the mirror data of WeChat and obtain the WeChat synchronization cache file; S200: Determine whether there is a data block. If so, execute step S300; otherwise, execute step S600; S300: Obtain the starting offset address of the current data block, and calculate the first byte length and the second byte length; S400: According to the first byte length and the second byte length, read the current data block and record the basic information of the current data block in the first set. Among them, the basic information includes the starting offset address and the second byte length; S500: Update the starting offset address and execute step S200; S600: Parse the data structure of the data block in a nested and recursive manner until each subset no longer contains any subsets; S700: Store the type value and content of the chat record in the sixth layer of the first subset. Among them, the fourth item in the sixth layer represents the type of the chat record, the fifth item represents the content of the chat record, and the eighth item represents the attachment of the chat record; S800: Parse and restore the chat record according to the type of the chat record.

2. The chat record recovery method based on WeChat synchronized cache files according to claim 1, characterized in that Step S100 includes the following steps: Obtain the WeChat synchronization cache file in the mirror data directory com.tencent.mm\files\mmkv of WeChat, and its file name starts with SyncMMKV.

3. A chat record recovery method based on WeChat synchronized cache files according to claim 1, characterized in that Step S200 includes the following steps: The method for determining whether there is a data block includes: querying whether the keyword 0x1170726F636573735F646174615F6C69737401001170726F636573735F646174615F6C697374 is included in the WeChat synchronization cache file.

4. A chat record recovery method based on WeChat synchronous cache files according to claim 1, characterized in that, It is characterized in that Step S300 includes the following steps: S301: Record the starting address of the current keyword, and the value obtained by adding 0x26 to the starting address of the current keyword is used as the starting offset address of the current data block; S302: Address the starting offset address and set the initial value of the first value to 0; S303: Read the content of one byte at the current address; S304: Determine whether the highest bit of the current byte is 1. If so, execute step S305; otherwise, execute step 306; S305: Increment the first value by 1 and address the next byte, then execute step S304; S306: Increment the first value by 1, remove the highest bit of all the read bytes, store them in little-endian format, and fill zeros for the insufficient high bits. The obtained value is used as the first byte length; S307: Update the starting offset address to the sum of the starting offset address and the current first value, and set the initial value of the second value to 0; S308: Address the starting offset address; S309: Read the content of one byte at the current address; S310: Determine whether the highest bit of the current byte is 1. If so, execute step S311; otherwise, execute step 312; S311: Increment the second value by 1 and address the next byte, then execute step S310; S312: Increment the second value by 1. After removing the highest bits of all the read bytes, store them in little - endian format, padding with zeros if the high bits are insufficient. The resulting value is used as the second byte length. S313: Determine whether the first byte length is equal to the sum of the second byte length and the second value. If so, execute step S400; otherwise, execute step S314: S314: Increment the starting offset address by 1 and execute step S200.

5. A chat record recovery method based on WeChat synchronized cache files according to claim 1, characterized in that, It is characterized in that Step S500 includes the following steps: Update the starting offset address to the sum of the current starting offset address and the first byte length.

6. A chat record recovery method based on WeChat synchronous cache files according to claim 1, characterized in that It is characterized in that Step S800 includes the following steps: S801: Read the value of the fourth item in the sixth layer of the first subset. S802: Determine whether the fourth item in the sixth layer of the first subset is one of the following values: If the value is 1, it represents text, and execute step S803; If the value is 3, it represents a picture, and execute step S804; If the value is 34, it represents voice, and execute step S804; If the value is 43 or 62, it represents video, and execute step S804; S803: Parse the text content and output the chat record of the text type. The text content is stored in the first entry of the fifth item in the sixth layer of the first subset, and then execute step S805; S804: Parse the attachment and output the chat record of non - text type. The attachment content is stored in the eighth item of the sixth layer of the first subset, which contains two entries. The first entry stores the length of the media file, and the second entry stores the underlying binary data of the media file. S805: Output the restored chat record and end the process.