Evidence obtaining method and system for multiple different input methods
By positioning, analyzing and analyzing the commonly used vocabulary files of Sogou input method and QQ input method, the problem of difficulty in extracting key vocabulary information in the existing technology is solved, and more reliable and accurate electronic data evidence collection is achieved.
Patent Information
- Application Number
- CN202510362185.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
It is difficult for the prior art to efficiently and accurately extract the key information of the commonly used vocabulary of Sogou input method and QQ input method for users, which affects the reliability of electronic data evidence collection.
By positioning the user's commonly used word library files of the input method, analyzing the file header data, reading multiple parameter sets, calculating the file validity, and converting the binary data into the user's commonly used word string and its word frequency, which is displayed in the user interface.
It realizes efficient and accurate extraction of Sogou input method and QQ input method vocabulary database, and improves the reliability and accuracy of electronic data evidence collection.
Smart Images

Figure CN120295875A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of evidence collection, and particularly to a method and system for collecting evidence for multiple different input methods. Background Art
[0002] In today's digital information age, Chinese input methods have become key tools for people's daily text input and are widely used on various electronic devices. Sogou Input Method and QQ Input Method have stood out among numerous input methods with their powerful functions and excellent user experiences, and have a large user base. According to relevant data, as of 2024, the monthly active user number of Sogou Input Method is as high as 480 million, and the monthly active user number of QQ Input Method has also reached 120 million. The two occupy a considerable share in the input method market.
[0003] Sogou Input Method and QQ Input Method have powerful self-learning capabilities. They can deeply analyze users' input habits, automatically construct and continuously update users' common chat phrases, and save these phrases in a carefully constructed common word library. With their advanced algorithms, Sogou Input Method and QQ Input Method can accurately capture the high-frequency words and unique expression methods input by users, and the number of common words in their word libraries exceeds 5 million. The word library significantly improves the convenience of users' input. During the user input process, the input method can quickly associate the words or phrases that the user may need to input based on the content of the word library, greatly reducing the number of user keystrokes and improving the input efficiency.
[0004] In the field of electronic data evidence collection, users' common word libraries contain extremely rich and key information, which is of inestimable value for deeply understanding users' behavior patterns and activity characteristics. By analyzing the high-frequency words, specific domain terms, and unique expression methods in the word library, it is possible to accurately understand users' interests, hobbies, professional characteristics, and daily focus.
[0005] For example, in a cybercrime investigation, if it is found that words related to specific online fraud methods, such as "brush orders for rebates" and "online loan thawing funds", frequently appear in the common word library of a criminal suspect, combined with other evidence, it can provide strong support for determining the suspect's modus operandi and crime type. Therefore, collecting evidence for Sogou Input Method and QQ Input Method has important value in the fields of electronic data evidence collection and forensic appraisal. Summary of the Invention
[0006] Based on this, it is necessary to provide a method for collecting evidence for multiple different input methods that can improve the reliability of evidence collection.
[0007] At the same time, a system for collecting evidence for multiple different input methods that can improve the reliability of evidence collection is provided.
[0008] A forensics method for multiple different input methods, including:
[0009] Location: Locate the user's common word library file of the input method;
[0010] Reading and analysis: Analyze the data at the beginning of the file header, read and calculate the parameter set, determine the validity of the file, and obtain the key parameters;
[0011] Parsing: Perform loop reading and parsing, convert the binary data into the user's common word strings and their word frequencies, and display them.
[0012] In a preferred embodiment, in the reading and analysis step, analyze the first 20 bytes of data at the beginning of the file header, and read the file format according to the parameter set. The input method is Sogou Input Method or QQ Input Method.
[0013] In a preferred embodiment, in the reading and analysis step, read the first 20 bytes of data at the beginning of the file header, and assign them to the 4-byte data header data A at the beginning of the file header, the next 4-byte data header data B, the next 4-byte data header data C, the next 4-byte data header data D, and the next 4-byte data header data E. Perform a conditional judgment. If it is determined that the file data is abnormal and does not meet the parsing conditions, then exit the subsequent operations.
[0014] In a preferred embodiment, for the conditional judgment in the reading and analysis step, if the header data A, header data B, header data C, header data D, and header data E <= 0, or the header data B >= the file size, then it is determined that the file data is abnormal and does not meet the parsing conditions, and the subsequent operations are exited.
[0015] In a preferred embodiment, the parameter set includes: parameter set 1. The reading and analysis includes reading parameter set 1. The reading of parameter set 1 includes: copy the data from the 20th byte after the current position of the file to the end of the file, and name it file block A.
[0016] Define and assign values to the keyItems array: Loop through the header data C times. For each loop, obtain 2-byte data from after the header of file block A and denote it as keyItem.tpdef. Then obtain the next 2-byte data and denote it as the key type quantity. If the key type quantity > 0, then loop through the key type quantity times. Each time, obtain the next 2-byte data from file block A and denote it as datype, and add it to the datype array of keyItem. Continue to obtain the next 4-byte data from file block A and denote it as keyItem.at_idx, the next 4-byte data as keyItem.key_da_idx, the next 4-byte data as keyItem.da_idx, and the next 4-byte data as keyItem.v6. Add the constructed keyItem to the keyItems array set;
[0017] Define and assign values to the atItems array: Loop through the header data D times. Each time, obtain the next 4-byte data from file block A and denote it as atItem.count, then the next 4-byte data as atItem.a2, the next 4-byte data as atItem.da_id, and the next 4-byte data as atItem.b2. Add the constructed atItem to the atItems array set;
[0018] Define and assign values to the intItems array: Loop through the header data E times. Each time, obtain the next 4-byte data from file block A and denote it as intItem.count. Add the constructed intItem to the intItems array set.
[0019] In a preferred embodiment, the parameter set includes: Parameter set 2, and the reading and analysis includes reading Parameter set 2 and determining the file validity,
[0020] The reading Parameter set 2 and determining the file validity includes:
[0021] Define the constant block header size = 12 * (atItems.size() + intItem.size() + keyItem.size()) + 24, and the block length = file length - 80;
[0022] Move the file pointer to the 80th byte after the file header, obtain the next 4-byte data and denote it as tmp1, then the next 4-byte data as tmp2, and the next 4-byte data as the total block size;
[0023] Perform a conditional judgment. If it is determined that the file does not conform to the valid format and cannot be parsed, then exit the subsequent operations.
[0024] In a preferred embodiment, in the step of reading the parameter set 2 and determining the validity of the file, a conditional judgment is made: if (total block size <= 0 || block length < 0 || (total block size + block header size + header data B + 8) != block length), it is determined that the file does not conform to the valid format and cannot be parsed, and subsequent operations are exited.
[0025] In a preferred embodiment, the parameter set includes: parameter set 3, parameter set 4. The reading and analysis includes reading parameter set 3 and reading parameter set 4. The reading of parameter set 3 includes:
[0026] Define a constant vector array:
[0027] Define a constant vector array datypeHSize[] = {0, 27, 414, 512, -1, -1, 512, 0};
[0028] Define a constant vector array typeSize[] = {4, 1, 1, 2, 1, 2, 2, 4, 4, 8, 4, 4, 4, 0, 0, 0};
[0029] Define an integer vector array keyHSize(500, 0, 0, 0, 0, 0, 0, 0, 0, 0);
[0030] Calculate and define relevant arrays:
[0031] baseHSize array: For each element in the keyItems array, calculate size = keyItems[i].tpdef >> 2, and then perform a bitwise AND operation with 4; mask_td = keyItems[i].tpdef & 0xFFFFFF8F. If keyHSize.size() == 0, add datypeHSize[mask_td] to baseHSize; if keyHSize[idx_key] > 0, add keyHSize[idx_key] to baseHSize; otherwise, add datypeHSize[mask_td] to baseHSize;
[0032] datypeSize array: If keyItems[i].at_idx < 0, loop keyItems[i].datype.size() times, and record the loop index as isize. When isize > 0 || mask_td != 4, size += typeSize[key[i].datype[isize]]. If keyItems[i].at_idx == -1, then size += 4. Finally, add size to the datypeSize array; if keyItems[i].at_idx >= 0, first calculate num_non_at = keyItems[i].datype.size() - at[keyItems[i].at_idx].count, then loop num_non_at times, and record the loop index as isize. When isize > 0 || mask_td != 4, size += typeSize[keyItems[i].datype[isize]]. If (keyItems[i].tpdef & 0x60) > 0, then size += 4, and add it to the datypeSize array after adding 4;
[0033] atSize array: Initialize the atSize array to zero. Loop from num_non_at to keyItems[i].datype.size() - 1. Each time, size += typeSize[keyItems[i].datype[isize]]. If (keyItems[i].tpdef & 0x40) == 0, then size += 4. Finally, assign size to atSize[keyItems[i].at_idx];
[0034] Define other parameter vector arrays:
[0035] heIitemIndex vector: According to the next 4-byte number sizeB2 in the file, loop sizeB2 times. Each time, obtain the next 4-byte data from the file and record it as head.offset, then the next 4-byte data as head.dasize, and the next 4-byte data as head.usedDasize. Add the constructed head to the heIitemIndex vector;
[0036] heItemAt vector: According to the next 4-byte number size4_b2 in the file, loop size4_b2 times. Each time, obtain the next 4-byte data from the file and record it as heAt.offset, the next 4-byte data as heAt.dasize, and the next 4-byte data as heAt.usedDasize. Add the constructed head to the heItemAt vector;
[0037] dastoItems vector: According to the next 4-byte number size5_b2 in the file, loop size5_b2 times. Each time, obtain the next 4-byte data from the file and record it as dastore.offset, the next 4-byte data as dastore.dasize, and the next 4-byte data as dastore.usedDasize. Add the constructed dastore to the dastoItems vector;
[0038] usrHead parameter: Position the file pointer to the 76th byte position of the file, named fr. Create a vectorins array. Loop 19 times. Each time, read the next 4-byte number of fr and add it to the ins array, and at the same time offset fr by 4 bytes. Finally, assign ins
[14] to usrHead.p2 and ins
[15] to usrHead.p3;
[0039] The said read parameter set 4 includes:
[0040] haStoreBase: Copy the data from the start after offsetting the current position of the file by heItemAt[0].offset - 8 * baseHSize[0] to the end of the file to haStoreBase;
[0041] num_at: If heAt[keyItems[0].attr_idx].used_datasize == 0, then num_at = heAt[keyItems[0].attr_idx].dasize; otherwise, num_at = heAt[keyItems[0].attr_idx].usedDasize.
[0042] In the preferred embodiment, in the said parsing, parse the user's common words:
[0043] Start looping to read the file data until the end of the file. The loop count idx_HasStore ranges from 0 to baseHSize[0] - 1;
[0044] Obtain the first 4 - byte number HasStore.offset of the haStoreBase header and the next 4 - byte number HasStore.count of the haStoreBase;
[0045] Loop: The number of loop iterations at_id ranges from 0 to HasStore.count - 1. Read at_id_offset: Read the 4 - byte data starting from heItemAt[index_id].offset + HasStore.offset+datypeSize[index_id]*at_id + datypeSize[key_id] - 4 after the current file position;
[0046] For each at2_id ranging from 0 to num_at - 1, obtain the data from the position heAt[keyItems[0].at_idx].offset+at_id_offset to the end of the file at the current file position, and denote it as at2_base;
[0047] Wd_inf reading: Denote the 4 - byte number after the data header of at2_base as Wd_inf.offset, the next 2 - byte number as Wd_inf.freq, skip the following 6 - byte value, then obtain the next 2 - byte number and denote it as Wd_inf.p1, continue to offset by 8 bytes, and obtain the next 2 - byte number and denote it as Wd_inf.pos;
[0048] Wd_ba reading: Obtain da_id = atItems[keyItems[0].at_idx].da_id, and denote the data from the position dastoItems[da_id].offset + Wd_inf.offset bytes after the current file position to the end of the file as Wd_ba;
[0049] Wd_ba parsing: Calculate k1 = ((Wd_inf.p1 + usrHead.p2) << 2), k2 = ((Wd_inf.p1 + usrHead.p3) << 2), xk = (k1 + k2) & 0xffff. Denote the 2 - byte data after the data header of Wd_ba as n, and offset by 2 bytes. Loop n times. Each time, calculate shift = p2 % 8, read the 2 - byte data after the current position of Wd_ba and denote it as ch, and offset by 2 bytes. Perform bit operations on ch. First, shift it left by 16-(shift % 8) bits, then OR it with the value after shifting it right by shift bits, then AND it with 0xffff, and finally XOR it with xk, and denote it as dch. Convert the lower 8 bits of dch to char type and add it to the byte array decwords, and convert the higher 8 bits of dch after shifting it right by 8 bits to char type and append it to decwords;
[0050] Convert the byte array decwords to Unicode encoding to obtain a string of commonly used words by users. Save the string and its word frequency Wd_inf.freq to a queue named decodedWordStr;
[0051] Skip backward from the current position by a data block of size atSize[keyItems[0].at_idx]. If the end of the file has been reached, end the loop; otherwise, return to the beginning of the loop to read the file data until the end of the file;
[0052] Show the result: Output decodedWordStr, and display the commonly used words by users using the input method and their usage frequencies on the UI to visually present the evidence collection result.
[0053] An evidence collection system for multiple different input methods, including:
[0054] Location module: Locate the commonly used word library file of the input method;
[0055] Reading and analysis module: Analyze the file header data, read and calculate the parameter sets to determine the validity of the file and obtain key parameters;
[0056] Parsing module: Perform loop reading and parsing, convert binary data into a string of commonly used words by users and their word frequencies, and display them.
[0057] The above-mentioned evidence collection method and system for multiple different input methods locate the commonly used word library file under the data storage path of input methods such as Sogou Pinyin Input Method and QQ Pinyin Input Method. Then, by analyzing the file header data and reading and calculating multiple parameter sets, determine the validity of the file and obtain key parameters. Finally, through specific loop reading and parsing operations, convert binary data into a string of commonly used words by users and their word frequencies, and display the result on the UI. The present invention solves the problem in the prior art that it is difficult to efficiently and accurately extract key information of the input method word library, and provides a more reliable and accurate means for electronic data evidence collection. Description of the Drawings
[0058] Figure 1 It is a partial flowchart of the evidence collection method for multiple different input methods according to an embodiment of the present invention;
[0059] Figure 2 It is a partial flowchart of the evidence collection method for multiple different input methods according to a preferred embodiment of the present invention;
[0060] Figure 3 It is a partial result display of the evidence collection method for multiple different input methods according to an embodiment of the present invention when used in Sogou Input Method;
[0061] Figure 4 The evidence collection method for multiple different input methods according to an embodiment of the present invention is used for partial result display in QQ Input Method. Specific implementation manner
[0062] As Figure 1 shown, the evidence collection method for multiple different input methods according to an embodiment of the present invention includes:
[0063] Step S101, positioning: Locate the user's common word library file of the input method;
[0064] Step S103, reading and analyzing: Analyze the file header data, read and calculate the parameter set, determine the validity of the file, and obtain key parameters;
[0065] Step S105, parsing: Perform loop reading and parsing, convert the binary data into the user's common word string and its word frequency, and display it.
[0066] As Figure 2 shown, further, for the positioning in this embodiment, input the user's common word library file of Sogou Input Method or QQ Input Method, etc., and read the user's common word library file.
[0067] Further, in this embodiment, the input method is preferably Sogou Input Method or QQ Input Method. Of course, it can also be applied to similar input methods and other input methods.
[0068] As Figure 2 shown, in the reading and analyzing step of this embodiment, analyze the first 20-byte data of the file header, and read the file format according to the parameter set.
[0069] Further, in the reading and analyzing step of this embodiment, read the first 20-byte data of the file header, and assign them to the 4-byte data header data A at the beginning of the file header, the 4-byte data header data B in the next position, the 4-byte data header data C in the next position, the 4-byte data header data D in the next position, and the 4-byte data header data E in the next position. Then perform a conditional judgment. If it is determined that the file data is abnormal and does not meet the parsing conditions, then exit the subsequent operations.
[0070] In the conditional judgment in the reading and analyzing step of this embodiment, if the header data A, header data B, header data C, header data D, and header data E <= 0, or the header data B >= the file size, then it is determined that the file data is abnormal and does not meet the parsing conditions, and then exit the subsequent operations.
[0071] Specific reading and analyzing steps:
[0072] Read the first 20-byte data of the file header:
[0073] Header data A = 4-byte data at the start of the file header
[0074] Header data B = 4-byte data at the next position
[0075] Header data C = 4-byte data at the next position
[0076] Header data D = 4-byte data at the next position
[0077] Header data E = 4-byte data at the next position
[0078] Condition 1: If header data A, header data B, header data C, header data D, and header data E <= 0, exit
[0079] If header data B >= file size, exit.
[0080] Furthermore, the parameter set of this embodiment includes: parameter set 1, parameter set 2, parameter set 3, parameter set 4.
[0081] As Figure 2 shown, furthermore, the reading and analysis of this embodiment include reading parameter set 1, reading parameter set 2 and judging the file validity, reading parameter set 3, reading parameter set 4.
[0082] Furthermore, the reading of parameter set 1 in this embodiment includes:
[0083] Copy the data from 20 bytes after the current position of the file to the end of the file, and name it file block A;
[0084] Define and assign values to the keyItems array: Loop header data C times. For each loop, obtain 2-byte data from after the header of file block A and denote it as keyItem.tpdef, then obtain the next 2-byte data and denote it as the key type quantity. If the key type quantity > 0, then loop key type quantity times. Each time, obtain the next 2-byte data from file block A and denote it as datype, and add it to the datype array of keyItem. Continue to obtain the next 4-byte data from file block A and denote it as keyItem.at_idx, the next 4-byte data as keyItem.key_da_idx, the next 4-byte data as keyItem.da_idx, the next 4-byte data as keyItem.v6, and add the constructed keyItem to the keyItems array set;
[0085] Define and assign values to the atItems array: Loop through the data in the loop header D times. Each time, obtain the next 4-byte data from file block A and record it as atItem.count, then the next 4-byte data as atItem.a2, then the next 4-byte data as atItem.da_id, and then the next 4-byte data as atItem.b2. Add the constructed atItem to the atItems array collection;
[0086] Define and assign values to the intItems array: Loop through the data in the loop header E times. Each time, obtain the next 4-byte data from file block A and record it as intItem.count. Add the constructed intItem to the intItems array collection.
[0087] Specifically, the read parameter set 1 of this embodiment:
[0088] Read parameter set 1 (integer arrays keyItems, atItems, and intItems):
[0089] First, copy the data from 20 bytes after the current position of the file to the end of the file and name it file block A.
[0090]
[0091]
[0092] Furthermore, the read parameter set 2 of this embodiment and the judgment of file validity include:
[0093] Define the constant block header size = 12 * (atItems.size() + intItem.size() + keyItem.size()) + 24, and the block length = file length - 80;
[0094] Move the file pointer to the 80th byte after the file header, read the next 4-byte data and record it as tmp1, then the next 4-byte data as tmp2, and then the next 4-byte data as the total block size;
[0095] Perform a conditional judgment. If it is determined that the file does not conform to the valid format and cannot be parsed, then exit the subsequent operations.
[0096] Furthermore, in the read parameter set 2 of this embodiment and the judgment of file validity, perform a conditional judgment:
[0097] If (the total block size <= 0 || the block length < 0 || (the total block size + the block header size + the header data B + 8) != the block length), then judge that the file does not conform to the valid format and cannot be parsed, and then exit the subsequent operations.
[0098] Read the parameter set 2 of the specific embodiment and determine the validity of the file:
[0099] Read the parameter set 2 (block header size, tmp1, tmp2, total block size), and determine whether the file is valid
[0100] First, define constants:
[0101] Block header size = 12 * (atItems.size() + intItem.size() + keyItem.size()) + 24;
[0102] Block length = file length - 80;
[0103] Move the file pointer to the 80th byte after the file header, and continue to read the file data:
[0104] tmp1 = the next 4 - byte number in the file;
[0105] tmp2 = the next 4 - byte number after that;
[0106] Total block size = the next 4 - byte number after that;
[0107] Condition 2: if (total block size <= 0 || block length < 0 || (total block size + block header size + header data B + 8) != block length)
[0108] Exit
[0109] Furthermore, the parameter set 3 read in this embodiment includes:
[0110] Define a constant vector array:
[0111] Define a constant vector array datypeHSize[] = {0, 27, 414, 512, -1, -1, 512, 0},
[0112] Define a constant vector array typeSize[] = {4, 1, 1, 2, 1, 2, 2, 4, 4, 8, 4, 4, 4, 0, 0, 0},
[0113] Define an integer vector array keyHSize(500, 0, 0, 0, 0, 0, 0, 0, 0, 0);
[0114] Calculate and define relevant arrays:
[0115] baseHSize array: For each element in the keyItems array, calculate size = keyItems[i].tpdef shifted right by 2 bits, and then perform a bitwise AND operation with 4; mask_td = keyItems[i].tpdef ANDed with the hexadecimal number 0xFFFFFF8F. If keyHSize.size() == 0, add datypeHSize[mask_td] to baseHSize; if keyHSize[idx_key] > 0, add keyHSize[idx_key] to baseHSize; otherwise, add datypeHSize[mask_td] to baseHSize;
[0116] datypeSize array: If keyItems[i].at_idx < 0, loop keyItems[i].datype.size() times, with the loop index denoted as isize. When isize > 0 || mask_td != 4, size += typeSize[key[i].datype[isize]]. If keyItems[i].at_idx == -1, then size += 4. Finally, add size to the datypeSize array; if keyItems[i].at_idx >= 0, first calculate num_non_at = keyItems[i].datype.size() - at[keyItems[i].at_idx].count, then loop num_non_at times, with the loop index denoted as isize. When isize > 0 || mask_td != 4, size += typeSize[keyItems[i].datype[isize]]. If (keyItems[i].tpdef ANDed with 0x60 > 0), then size += 4, and add it to the datypeSize array after adding 4;
[0117] atSize array: Initialize the atSize array to zero. Loop from num_non_at to keyItems[i].datype.size() - 1. Each time, size += typeSize[keyItems[i].datype[isize]]. If keyItems[i].tpdef ANDed with 0x40 == 0, then size += 4. Finally, assign size to atSize[keyItems[i].at_idx];
[0118] Define other parameter vector arrays:
[0119] heIitemIndex vector: According to the next 4-byte number sizeB2 in the file, loop sizeB2 times. Each time, obtain the next 4-byte data from the file and record it as head.offset, the next 4-byte data as head.dasize, and the next 4-byte data as head.usedDasize. Add the constructed head to the heIitemIndex vector;
[0120] heItemAt vector: According to the next 4-byte number size4_b2 in the file, loop size4_b2 times. Each time, obtain the next 4-byte data from the file and record it as heAt.offset, the next 4-byte data as heAt.dasize, and the next 4-byte data as heAt.usedDasize. Add the constructed head to the heItemAt vector;
[0121] dastoItems vector: According to the next 4-byte number size5_b2 in the file, loop size5_b2 times. Each time, obtain the next 4-byte data from the file and record it as dastore.offset, the next 4-byte data as dastore.dasize, and the next 4-byte data as dastore.usedDasize. Add the constructed dastore to the dastoItems vector;
[0122] usrHead parameter: Position the file pointer to the 76th byte position of the file, named fr. Create a vectorins array. Loop 19 times. Each time, read the next 4-byte number of fr and add it to the ins array, and at the same time offset fr by 4 bytes. Finally, assign ins
[14] to usrHead.p2 and ins
[15] to usrHead.p3.
[0123] Specifically, the read parameter set 3 of this embodiment: Read parameter set 3 (datatypeHSize, typeSize, keyHSize, baseHSize, datypeSize, atSize, heIitemIndex, heItemAt, dastoItems, usrHead, at_head):
[0124] Define a constant vector array:
[0125] Constant vector array datypeHSize[] = {0, 27, 414, 512, -1, -1, 512, 0};
[0126] The constant vector array typeSize[] = {4, 1, 1, 2, 1, 2, 2, 4, 4, 8, 4, 4, 4, 0, 0, 0};
[0127] The integer vector array keyHSize(500, 0, 0, 0, 0, 0, 0, 0, 0, 0);
[0128] 2) The integer vector arrays baseHSize, datypeSize, and atSize are defined as follows:
[0129]
[0130]
[0131] 3) Define multiple parameters such as the integer vector arrays heIitemIndex, heItemAt, dastoItems, usrHead, at_head, etc.
[0132]
[0133]
[0134] Furthermore, the read parameter set 4 of this embodiment includes:
[0135] haStoreBase: Copy the data from the starting position after offsetting the current position of the file by heItemAt[0].offset - 8 * baseHSize[0] to the end of the file to haStoreBase;
[0136] num_at: If heAt[keyItems[0].attr_idx].used_datasize == 0, then num_at = heAt[keyItems[0].attr_idx].dasize; otherwise, num_at = heAt[keyItems[0].attr_idx].usedDasize.
[0137] Specifically, the read parameter set 4 of this embodiment: Define the parameter set 4 (haStoreBase, num_at)
[0138] Block header size = 12 * (atItems.size() + atItems.size() + keyItems.size()) + 24;
[0139]
[0140]
[0141] As Figure 2 shown, further, in the analysis of this embodiment, analyze the user's common words:
[0142] Start looping to read file data until the end of the file. The loop count idx_HasStore ranges from 0 to baseHSize[0] – 1;
[0143] Obtain the first 4-byte number HasStore.offset of the haStoreBase header and the next 4-byte number HasStore.count following the haStoreBase;
[0144] Loop: The loop count at_id ranges from 0 to HasStore.count - 1. Read at_id_offset: The first 4-byte data starting from heItemAt[index_id].offset + HasStore.offset + datypeSize[index_id] * at_id + datypeSize[key_id] - 4 after the current file position;
[0145] For each at2_id ranging from 0 to num_at - 1, obtain the data from the position heAt[keyItems[0].at_idx].offset + at_id_offset at the current file position to the end of the file, denoted as at2_base;
[0146] Wd_inf reading: The first 4-byte number after the data header of at2_base is denoted as Wd_inf.offset. The next 2-byte number is denoted as Wd_inf.freq. Skip the following 6-byte value, and then obtain the next 2-byte number denoted as Wd_inf.p1. Continue to offset by 8 bytes and obtain the next 2-byte number denoted as Wd_inf.pos;
[0147] Wd_ba reading: Obtain da_id = atItems[keyItems[0].at_idx].da_id. The data from the position dastoItems[da_id].offset + Wd_inf.offset bytes after the current file position to the end of the file is denoted as Wd_ba;
[0148] Wd_ba analysis: Calculate the sum of k1 = Wd_inf.p1 and usrHead.p2, shift it left by 2 bits, calculate the sum of k2 = Wd_inf.p1 and usrHead.p3, shift it left by 2 bits, xk = perform a bitwise AND operation on the sum of k1 and k2 with 0xffff. Denote the 2-byte data after the Wd_ba data header as n, and offset by 2 bytes. Loop n times. Each time, calculate shift = p2 modulo 8, read the 2-byte data at the current position of Wd_ba and denote it as ch, and offset by 2 bytes. Perform bit operations on ch. First, shift it left by 16 - (shift % 8) bits, then OR it with the result of shifting it right by shift bits, then perform a bitwise AND operation with 0xffff, and finally XOR it with xk, denoted as dch. Convert the lower 8 bits of dch to char type and add it to the byte array decwords. Shift the higher 8 bits of dch right by 8 bits, convert it to char type, and append it to decwords;
[0149] Convert the byte array decwords to Unicode encoding to obtain a string of commonly used words by the user. Save this string and its word frequency Wd_inf.freq to a queue named decodedWordStr;
[0150] Skip backward from the current position by a data block of size atSize[keyItems[0].at_idx]. If the end of the file has been reached, the loop ends; otherwise, return to the beginning and loop to read the file data until the end of the file;
[0151] Display the result: Output decodedWordStr, and display the commonly used words by the user using the input method and their usage frequencies on the UI to visually present the forensics result.
[0152] Preferably, class template functions such as QT's toUnicode or C++'s wstring_convert that perform conversions between wide strings and byte strings can be used to convert the byte array decwords to Unicode encoding to obtain a string of commonly used words by the user. Save this string and its word frequency Wd_inf.freq to a queue named decodedWordStr.
[0153] Specifically, the analysis of this embodiment:
[0154]
[0155]
[0156] Output decodedWordStr, and display the commonly used words by the user using Sogou Input Method or QQ Input Method and their usage frequencies on the UI.
[0157] Such asFigure 3 As shown, after installing Sogou Input Method 14.11 in a Windows 11 virtual machine and making some inputs, this algorithm is used for evidence collection, and the results are as shown in the figure.
[0158] As Figure 4 shown, after installing QQ Pinyin Input Method 6.6.6304.400 in a Windows 11 virtual machine and making some inputs, this algorithm is used for evidence collection, and the results are as shown in the figure.
[0159] The evidence collection method for multiple different input methods of the present invention includes analyzing and judging the file header data, reading and calculating multiple parameter sets, and a precise parsing method for the user's commonly used words based on these parameters. These steps cooperate with each other to adapt to the changes in the thesaurus structure of different versions of input methods and accurately extract key information.
[0160] In practical applications, if according to different application scenarios or optimization requirements, the execution order of these steps can be reasonably adjusted. As long as the purpose of accurately obtaining parameters and completing the thesaurus parsing can be finally achieved, it should also be regarded as within the scope covered by the present invention. For example, in some scenarios with extremely high requirements for file integrity judgment, the parameter set 2 can be read and its validity can be judged first, or the file validity can be judged first and then other parameter sets can be read; or in some scenarios with a large dependence on a specific parameter set, this parameter set can be preferentially read and then other parameter sets can be read as needed.
[0161] An evidence collection system for multiple different input methods according to an embodiment of the present invention includes:
[0162] Location module: Locate the user's commonly used thesaurus file of the input method;
[0163] Reading and analyzing module: Analyze the file header data, read and calculate the parameter set, determine the validity of the file, and obtain key parameters;
[0164] Parsing module: Perform loop reading and parsing, convert binary data into the user's commonly used word strings and their word frequencies, and display them.
[0165] Furthermore, in this embodiment, the input method is preferably Sogou Input Method or QQ Input Method. Of course, it can also be applied to similar input methods and other input methods.
[0166] In the reading and analyzing module of this embodiment, the first 20-byte data starting from the file header is analyzed, and the file format is read according to the parameter set.
[0167] Further, in the reading and analysis module of this embodiment, the 20-byte data starting from the file header is read and assigned to the 4-byte data header data A at the beginning of the file header, the next 4-byte data header data B, the next 4-byte data header data C, the next 4-byte data header data D, and the next 4-byte data header data E. Conditional judgment is performed. If it is determined that the file data is abnormal and does not meet the parsing conditions, the subsequent operations are exited.
[0168] In the conditional judgment of the reading and analysis module of this embodiment, if the header data A, header data B, header data C, header data D, and header data E <= 0, or the header data B >= the file size, it is determined that the file data is abnormal and does not meet the parsing conditions, and the subsequent operations are exited.
[0169] Further, the parameter set of this embodiment includes: parameter set 1, parameter set 2, parameter set 3, and parameter set 4.
[0170] Further, the reading and analysis module of this embodiment includes: a unit for reading parameter set 1, a unit for reading parameter set 2 and judging the file validity, a unit for reading parameter set 3, and a unit for reading parameter set 4.
[0171] Further, the unit for reading parameter set 1 of this embodiment includes:
[0172] Copy the data from 20 bytes after the current position of the file to the end of the file and name it file block A;
[0173] Define and assign the keyItems array: Loop the number of times of the header data C. For each loop, obtain 2-byte data from after the header of file block A and record it as keyItem.tpdef, then obtain the next 2-byte data and record it as the key type quantity. If the key type quantity > 0, then loop the number of times of the key type quantity again. Each time, obtain the next 2-byte data from file block A and record it as datype, and add it to the datype array of keyItem. Continue to obtain the next 4-byte data from file block A and record it as keyItem.at_idx, the next 4-byte data as keyItem.key_da_idx, the next 4-byte data as keyItem.da_idx, and the next 4-byte data as keyItem.v6. Add the constructed keyItem to the keyItems array set;
[0174] Define and assign values to the atItems array: Loop through the data in the loop header D times. Each time, obtain the next 4-byte data from file block A and denote it as atItem.count, the next 4-byte data as atItem.a2, the next 4-byte data as atItem.da_id, and the next 4-byte data as atItem.b2. Then add the constructed atItem to the atItems array collection;
[0175] Define and assign values to the intItems array: Loop through the data in the loop header E times. Each time, obtain the next 4-byte data from file block A and denote it as intItem.count, and then add the constructed intItem to the intItems array collection.
[0176] Furthermore, the unit for reading parameter set 2 and judging the file validity in this embodiment includes:
[0177] Define the constant block header size = 12 * (atItems.size() + intItem.size() + keyItem.size()) + 24, and the block length = file length - 80;
[0178] Move the file pointer to the 80th byte after the file header, read the next 4-byte data and denote it as tmp1, the next 4-byte data as tmp2, and the next 4-byte data as the total block size;
[0179] Perform a conditional judgment. If it is determined that the file does not conform to the valid format and cannot be parsed, then exit the subsequent operations.
[0180] Furthermore, in the unit for reading parameter set 2 and judging the file validity in this embodiment, perform a conditional judgment: If (total block size <= 0 || block length < 0 || (total block size + block header size + header data B + 8) != block length), then judge that the file does not conform to the valid format and cannot be parsed, and then exit the subsequent operations.
[0181] Furthermore, the unit for reading parameter set 3 in this embodiment includes:
[0182] Define a constant vector array:
[0183] Define the constant vector array datypeHSize[] = {0, 27, 414, 512, -1, -1, 512, 0},
[0184] Define the constant vector array typeSize[] = {4, 1, 1, 2, 1, 2, 2, 4, 4, 8, 4, 4, 4, 0, 0, 0},
[0185] Define an integer vector array keyHSize(500,0,0,0,0,0,0,0,0,0);
[0186] Calculate and define the relevant arrays:
[0187] baseHSize array: For each element in the keyItems array, calculate size = keyItems[i].tpdef >> 2, and then perform a bitwise AND operation with 4; mask_td = keyItems[i].tpdef & 0xFFFFFF8F. If keyHSize.size() == 0, add datypeHSize[mask_td] to baseHSize; if keyHSize[idx_key] > 0, add keyHSize[idx_key] to baseHSize; otherwise, add datypeHSize[mask_td] to baseHSize;
[0188] datypeSize array: If keyItems[i].at_idx < 0, loop keyItems[i].datype.size() times, and record the loop index as isize. When isize > 0 || mask_td != 4, size += typeSize[key[i].datype[isize]]. If keyItems[i].at_idx == -1, then size += 4, and finally add size to the datypeSize array; If keyItems[i].at_idx >= 0, first calculate num_non_at = keyItems[i].datype.size() - at[keyItems[i].at_idx].count, and then loop num_non_at times, and record the loop index as isize. When isize > 0 || mask_td != 4, size += typeSize[keyItems[i].datype[isize]]. If (keyItems[i].tpdef & 0x60) > 0, then size += 4, and add it to the datypeSize array after adding 4;
[0189] The atSize array: Initialize the atSize array to zero. Loop from num_non_at to keyItems[i].datype.size() - 1. Each time, size += typeSize[keyItems[i].datype[isize]]. If the bitwise AND of keyItems[i].tpdef and 0x40 == 0, then size += 4. Finally, assign size to atSize[keyItems[i].at_idx];
[0190] Define other parameter vector arrays:
[0191] The heIitemIndex vector: According to the next 4-byte number sizeB2 in the file, loop sizeB2 times. Each time, obtain the next 4-byte data from the file and record it as head.offset, then the next 4-byte data as head.dasize, and then the next 4-byte data as head.usedDasize. Add the constructed head to the heIitemIndex vector;
[0192] The heItemAt vector: According to the next 4-byte number size4_b2 in the file, loop size4_b2 times. Each time, obtain the next 4-byte data from the file and record it as heAt.offset, then the next 4-byte data as heAt.dasize, and then the next 4-byte data as heAt.usedDasize. Add the constructed head to the heItemAt vector;
[0193] The dastoItems vector: According to the next 4-byte number size5_b2 in the file, loop size5_b2 times. Each time, obtain the next 4-byte data from the file and record it as dastore.offset, then the next 4-byte data as dastore.dasize, and then the next 4-byte data as dastore.usedDasize. Add the constructed dastore to the dastoItems vector;
[0194] The usrHead parameter: Position the file pointer to the 76th byte position of the file, named fr. Create the vector ins array. Loop 19 times. Each time, read the next 4-byte number from fr and add it to the ins array, and at the same time offset fr by 4 bytes. Finally, assign ins
[14] to usrHead.p2 and ins
[15] to usrHead.p3.
[0195] Furthermore, the parameter set 4 reading unit of this embodiment includes:
[0196] haStoreBase: Copy the data from the position offset by heItemAt[0].offset - 8 * baseHSize[0] from the current position of the file to the end of the file to haStoreBase;
[0197] num_at: If heAt[keyItems[0].attr_idx].used_datasize == 0, then num_at = heAt[keyItems[0].attr_idx].dasize; otherwise, num_at = heAt[keyItems[0].attr_idx].usedDasize.
[0198] Further, in the parsing module of this embodiment, parse the user's common words:
[0199] Start looping to read the file data until the end of the file, and the number of loops idx_HasStore ranges from 0 to baseHSize[0] - 1;
[0200] Obtain the first 4-byte number HasStore.offset of haStoreBase and the next 4-byte number HasStore.count after haStoreBase;
[0201] Loop: The number of loops at_id ranges from 0 to HasStore.count - 1, and read at_id_offset: Read the first 4-byte data starting from heItemAt[index_id].offset + HasStore.offset + datypeSize[index_id] * at_id + datypeSize[key_id] - 4 after the current position of the file;
[0202] For each at2_id ranging from 0 to num_at - 1, obtain the data from the position starting from heAt[keyItems[0].at_idx].offset + at_id_offset at the current position of the file to the end of the file, denoted as at2_base;
[0203] Wd_inf reading: Denote the first 4-byte number after the data header of at2_base as Wd_inf.offset, the next 2-byte number as Wd_inf.freq, skip the next 6-byte value, then obtain the next 2-byte number and denote it as Wd_inf.p1, continue to offset 8 bytes, and obtain the next 2-byte number and denote it as Wd_inf.pos;
[0204] Wd_ba Reading: Obtain da_id = atItems[keyItems[0].at_idx].da_id, and record the data from dastoItems[da_id].offset + Wd_inf.offset bytes after the current position of the file to the end of the file as Wd_ba;
[0205] Wd_ba Parsing: Calculate k1 = (Wd_inf.p1 + usrHead.p2) << 2, k2 = (Wd_inf.p1 + usrHead.p3) << 2, xk = k1 + k2 & 0xffff. Record the 2 bytes of data after the data header of Wd_ba as n, and offset 2 bytes. Loop n times. Each time, calculate shift = p2 % 8, read the 2 bytes of data after the current position of Wd_ba as ch, and offset 2 bytes. Perform bit operations on ch. First, left-shift by 16 - (shift % 8) bits, then OR with the right-shifted shift bits, then AND with 0xffff, and finally XOR with xk, and record it as dch. Convert the lower 8 bits of dch to char type and add it to the byte array decwords. Convert the higher 8 bits of dch to char type by right-shifting 8 bits and append it to decwords;
[0206] Convert the byte array decwords to Unicode encoding to obtain a string of commonly used user words, and save the string and its word frequency Wd_inf.freq to a queue named decodedWordStr;
[0207] Skip the data block of size atSize[keyItems[0].at_idx] backward from the current position. If the end of the file has been reached, the loop ends. Otherwise, return to the beginning and loop to read the file data until the end of the file;
[0208] Display the result: Output decodedWordStr, and display the commonly used user words and their usage frequencies of the user's input method on the UI to visually present the forensics result.
[0209] Enlightened by the ideal embodiments of the present application described above, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this application. The technical scope of this application is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
[0210] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0211] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of flows and / or blocks.
[0212] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of flows and / or blocks.
[0213] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of flows and / or blocks.
Claims
1. A forensics method for multiple different input methods, characterized in that Including: Positioning: Locate the user's common word library file of the input method; Reading and analyzing: Analyze the file header data, read and calculate the parameter set, determine the validity of the file and obtain the key parameters; Parsing: Perform loop reading and parsing, convert the binary data into the user's common word string and its word frequency, and display it.
2. The evidence collection method for multiple different input methods according to claim 1, wherein In the step of reading and analyzing, analyze the first 20 bytes of data starting from the file header, and read the file format according to the parameter set. The input method is Sogou Input Method or QQ Input Method.
3. The forensics method for multiple different input methods according to claim 1, characterized in that, In the step of reading and analyzing, read the first 20 bytes of data starting from the file header, and assign them to the first 4 bytes of data at the start of the file header, denoted as header data A, the next 4 bytes of data as header data B, the next 4 bytes of data as header data C, the next 4 bytes of data as header data D, and the next 4 bytes of data as header data E. Conduct a conditional judgment. If it is determined that the file data is abnormal and does not meet the parsing conditions, then exit the subsequent operations.
4. The evidence collection method for multiple different input methods according to claim 3, wherein For the conditional judgment in the step of reading and analyzing, if header data A, header data B, header data C, header data D, and header data E <= 0, or header data B >= the file size, then it is determined that the file data is abnormal and does not meet the parsing conditions, and the subsequent operations are exited.
5. The forensics method for multiple different input methods according to any one of claims 1 to 4, characterized in that, The parameter set includes: Parameter set 1. The reading and analyzing includes reading Parameter set 1. The reading of Parameter set 1 includes: Copy the data from 20 bytes after the current position of the file to the end of the file, and name it file block A. Define and assign the keyItems array: Loop header data C times. For each loop, obtain 2 bytes of data from after the start of file block A, denoted as keyItem.tpdef, then obtain the next 2 bytes of data as the key type quantity. If the key type quantity > 0, then loop key type quantity times. Each time, obtain the next 2 bytes of data from file block A, denoted as datype, and add it to the datype array of keyItem. Then continue to obtain the next 4 bytes of data from file block A, denoted as keyItem.at_idx, the next 4 bytes of data as keyItem.key_da_idx, the next 4 bytes of data as keyItem.da_idx, and the next 4 bytes of data as keyItem.v6. Add the constructed keyItem to the keyItems array set. Define and assign the atItems array: Loop header data D times. Each time, obtain the next 4 bytes of data from file block A, denoted as atItem.count, the next 4 bytes of data as atItem.a2, the next 4 bytes of data as atItem.da_id, and the next 4 bytes of data as atItem.b2. Add the constructed atItem to the atItems array set. Define and assign values to the intItems array: Loop through the data of the loop header E times. Each time, obtain the next 4-byte data from file block A and denote it as intItem.count. Add the constructed intItem to the intItems array set.
6. The evidence collection method for multiple different input methods according to claim 5, characterized in that, The parameter set includes: Parameter set 2. The reading and analysis include reading parameter set 2 and judging the file validity. The reading of parameter set 2 and judging the file validity includes: Define the constant block header size = 12 * (the size of the atItems array + the size of the intItems array + the size of the keyItems array) + 24, and the block length = the file length - 80; Move the file pointer to the 80th byte after the file header. Read the next 4-byte data and denote it as tmp1, then the next 4-byte data as tmp2, and then the next 4-byte data as the total block size; Perform a conditional judgment. If it is determined that the file does not conform to the valid format and cannot be parsed, then exit the subsequent operations.
7. The forensics method for multiple different input methods according to claim 6, characterized in that, In the reading of parameter set 2 and judging the file validity, perform a conditional judgment: If (the total block size <= 0 || the block length < 0 || (the total block size + the block header size + header data B + 8) != the block length), then judge that the file does not conform to the valid format and cannot be parsed, and then exit the subsequent operations.
8. The forensics method for multiple different input methods according to claim 6, characterized in that, The parameter set includes: Parameter set 3, Parameter set 4. The reading and analysis include reading parameter set 3, reading parameter set 4. The reading of parameter set 3 includes: Define a constant vector array: Define a constant vector array datypeHSize[] = {0, 27, 414, 512, -1, -1, 512, 0}, Define a constant vector array typeSize [] = {4, 1, 1, 2, 1, 2, 2, 4, 4, 8, 4, 4, 4, 0, 0, 0}, Define an integer vector array keyHSize(500, 0, 0, 0, 0, 0, 0, 0, 0, 0); Calculate and define relevant arrays: baseHSize array: For each element in the keyItems array, calculate size = keyItems [i].tpdef shifted right by 2 bits, and then perform a bitwise AND operation with 4; mask_td = keyItems [i].tpdef and the hexadecimal number 0xFFFFFF8F for a bitwise AND operation. If keyHSize.size() == 0, then add datypeHSize[mask_td] to baseHSize; if keyHSize[idx_key] > 0, then add keyHSize[idx_key] to baseHSize; otherwise, add datypeHSize[mask_td] to baseHSize; datypeSize array: If keyItems[i].at_idx < 0, loop keyItems[i].datype.size() times, with the loop index denoted as isize. When isize > 0 || mask_td != 4, size += typeSize[key[i].datype[isize]]. If keyItems[i].at_idx == -1, then size += 4. Finally, add size to the datypeSize array; If keyItems[i].at_idx >= 0, first calculate num_non_at = keyItems[i].datype.size() - at[keyItems[i].at_idx].count, then loop num_non_at times, with the loop index denoted as isize. When isize > 0 || mask_td != 4, size += typeSize[keyItems[i].datype[isize]]. If keyItems[i].tpdef bitwise AND with 0x60 > 0, then size += 4, and add it to the datypeSize array after adding 4; atSize array: Initialize the atSize array to zero. Loop from num_non_at to keyItems[i].datype.size() - 1. Each time, size += typeSize[keyItems[i].datype[isize]]. If keyItems[i].tpdef bitwise AND with 0x40 == 0, then size += 4. Finally, assign size to atSize[keyItems[i].at_idx]; Define other parameter vector arrays: heIitemIndex vector: According to the next 4-byte number sizeB2 in the file, loop sizeB2 times. Each time, obtain the next 4-byte data from the file and denote it as head.offset, then the next 4-byte data as head.dasize, and the next 4-byte data as head.usedDasize. Add the constructed head to the heIitemIndex vector; heItemAt vector: According to the next 4-byte number size4_b2 in the file, loop size4_b2 times. Each time, obtain the next 4-byte data from the file and record it as heAt.offset, the next 4-byte data as heAt.dasize, and the next 4-byte data as heAt.usedDasize. Add the constructed head to the heItemAt vector; dastoItems vector: According to the next 4-byte number size5_b2 in the file, loop size5_b2 times. Each time, obtain the next 4-byte data from the file and record it as dastore.offset, the next 4-byte data as dastore.dasize, and the next 4-byte data as dastore.usedDasize. Add the constructed dastore to the dastoItems vector; usrHead parameter: Position the file pointer to the 76th byte position of the file, named fr. Create a vector ins array. Loop 19 times. Each time, read the next 4-byte number from fr and add it to the ins array, and at the same time offset fr by 4 bytes. Finally, assign ins[14] to usrHead.p2 and ins[15] to usrHead.p3; The said read parameter set 4 includes: haStoreBase parameter: Copy the data from the start after offsetting the current position of the file by heItemAt[0].offset - 8 * baseHSize[0] to the end of the file to haStoreBase; num_at parameter: If heAt[keyItems[0].attr_idx].used_datasize == 0, then num_at = heAt[keyItems[0].attr_idx].dasize; otherwise, num_at = heAt[keyItems[0].attr_idx].usedDasize.
9. The evidence collection method for multiple different input methods according to claim 8, characterized in that, In the said parsing, parse the user's common words: Start looping to read the file data until the end of the file. The loop count idx_HasStore ranges from 0 to baseHSize[0] – 1; Obtain the first 4-byte number HasStore.offset of the haStoreBase header and the next 4-byte number HasStore.count of the haStoreBase; Loop: The loop count at_id ranges from 0 to HasStore.count - 1. Read at_id_offset: The first 4-byte data starting from heItemAt[index_id].offset + HasStore.offset + datypeSize[index_id] * at_id + datypeSize[key_id] - 4 after the current position of the file; For each at2_id from 0 to num_at - 1, obtain the data from the starting position heAt[keyItems [0].at_idx].offset + at_id_offset of the current file position to the end of the file, denoted as at2_base; Wd_inf reading: The 4 - byte number after the data header of at2_base is denoted as Wd_inf.offset, the next 2 - byte number is denoted as Wd_inf.freq, skip the next 6 - byte value, then obtain the next 2 - byte number and denote it as Wd_inf.p1, continue to offset 8 bytes, and obtain the next 2 - byte number and denote it as Wd_inf.pos; Wd_ba reading: Obtain da_id = atItems[keyItems[0].at_idx].da_id, and the data from the byte position dastoItems[da_id].offset + Wd_inf.offset after the current file position to the end of the file is denoted as Wd_ba; Wd_ba parsing: Calculate k1 = (Wd_inf.p1 + usrHead.p2) << 2, k2 = (Wd_inf.p1 + usrHead.p3) << 2, xk = (k1 + k2) & 0xffff. Denote the 2 - byte data after the Wd_ba data header as n, and offset 2 bytes. Loop n times. Each time, calculate shift = p2 % 8, read the 2 - byte data after the current Wd_ba position and denote it as ch, and offset 2 bytes. Perform bit operations on ch. First, shift it left by 16 - (shift % 8) bits, then OR it with the result of shifting it right by shift bits, then AND it with 0xffff, and finally XOR it with xk, denoted as dch. Convert the lower 8 bits of dch to char type and add it to the byte array decwords, and convert the higher 8 bits of dch after shifting it right by 8 bits to char type and append it to decwords; Convert the byte array decwords to Unicode encoding to obtain a string of commonly used user words, and save the string and its word frequency Wd_inf.freq to a queue named decodedWordStr; Skip the data block of size atSize[keyItems[0].at_idx] backward from the current position. If the end of the file has been reached, the loop ends; otherwise, return to the beginning and loop to read the file data until the end of the file; Show the result: Output decodedWordStr, and display the commonly used user words and their usage frequencies of the user's input method on the UI to visually present the forensics result.
10. A forensics system for multiple different input methods, characterized in that, Including: Location module: Locate the file of the commonly used word library of the input method; Reading and analysis module: Analyze the file header data, read and calculate the parameter set, determine the validity of the file, and obtain the key parameters; Parsing module: performs loop reading and parsing, converts binary data into user-common word strings and their word frequencies, and performs display.