Method for quickly retrieving evidence obtaining result of instant messaging chatting record
By real-time collection and multimedia information analysis of instant messaging chat records, combined with database indexing and clustering technology, the problems of poor information readability and low search efficiency are solved, and the effect of rapid retrieval and analysis is achieved.
Patent Information
- Application Number
- CN202411686806.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-06
AI Technical Summary
When the prior art collects evidence for instant messaging applications, the information is poor readability, hidden information is difficult to detect, and the search efficiency is low when facing massive data.
By collecting text, pictures and audio and video information in chat records in real time, perform picture information analysis, audio and video information analysis, use OCR, speech recognition and natural language processing technology to extract and clean information, establish a database indexing system, and use the K-means clustering algorithm to cluster keywords, and finally restore the evidence collection results in the form of a web page, providing a quick retrieval function.
It realizes rapid retrieval and analysis of instant messaging chat records, improves the readability and search efficiency of information, can quickly discover hidden information, and reduces the analysis cost after evidence collection.
Smart Images

Figure CN119938953A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication evidence collection, and in particular, mainly relates to a method for quickly retrieving instant messaging chat record evidence collection results. Background Art
[0002] In the prior art, the main way to collect evidence from instant messaging applications is through text, pictures, voice and video. However, the traditional evidence collection method has the following problems:
[0003] 1. Poor information readability: The information obtained is only simply classified and summarized, which is difficult to read and understand directly, and requires a lot of time and energy to analyze.
[0004] 2. Hidden information is difficult to discover: Important information in audio and video is difficult to discover and extract through traditional means. Some key information hidden in the audio and video may be missed during manual analysis, which increases the difficulty of handling cases.
[0005] 3. Large amount of data, difficult to search: Faced with massive amounts of forensic data, manual search efficiency is extremely low, and it is difficult to quickly obtain key information. Even if the keywords are known, it is also difficult to quickly filter out matching information.
[0006] The disadvantage of the existing technology is that it is unable to efficiently process and parse multimedia information in instant messaging, resulting in poor readability of the information, difficulty in discovering hidden information, and low search efficiency when facing large amounts of data. Summary of the invention
[0007] In order to overcome the problems existing in the prior art, a method for quickly retrieving instant messaging chat record evidence collection results is proposed according to one aspect of the present invention, comprising:
[0008] S1. Real-time collection of target chat records, including text, pictures, and audio and video analysis;
[0009] S2, performing image information analysis on the image, wherein the image information analysis includes image preprocessing, and performing text recognition on the preprocessed image using OCR technology to extract text information hidden in the image;
[0010] S3, the audio and video analysis includes audio information analysis and video information analysis, wherein the audio information analysis includes audio preprocessing, and using speech recognition technology to convert the preprocessed audio content into text information;
[0011] S4, the video information analysis includes video preprocessing, image recognition and text extraction for each frame of the video, text conversion of the voice content in the video by combining speech recognition technology, and extraction of keywords in the video by natural language processing technology;
[0012] S5, cleaning and standardizing the text information extracted from the image, the text information converted from the audio, and the keywords extracted from the video, calculating the frequency of each word using the TF-IDF algorithm, selecting words with higher weights to associate and store with the corresponding original data, and establishing a database index system;
[0013] S6. Clustering the keys in the database index system using a K-means clustering algorithm, and re-associating and storing the classified database index system with the original data;
[0014] S7. Using the forensic result viewer, the forensic results in the database index system are read and filled into the preset web chat standard template, and restored in the form of a web page, maintaining the layout and display of the original information.
[0015] Furthermore, the image preprocessing specifically includes: gray-scaling, binarization, and denoising the image to improve the recognition accuracy of the OCR technology.
[0016] The picture information analysis includes:
[0017] 1) Image preprocessing: grayscale, binarize, and denoise the collected images to improve the OCR recognition accuracy.
[0018] 2) Text extraction: Use OCR technology to perform text recognition on the pre-processed images and extract the hidden text information in the images.
[0019] The audio information analysis includes:
[0020] 1) Audio preprocessing: Perform noise reduction, cutting and other processing on the collected audio to improve the accuracy of speech recognition.
[0021] 2) Text conversion: Use speech recognition technology to convert the pre-processed audio content into text information.
[0022] The video information analysis includes:
[0023] 1) Video preprocessing: Frame extraction and denoising are performed on the captured video to improve the accuracy of subsequent processing.
[0024] 2) Keyword extraction: Perform image recognition and text extraction on each frame of the video, convert the voice content in the video into text using speech recognition technology, and finally extract keywords in the video through natural language processing technology.
[0025] Furthermore, the establishment of the database index system specifically includes:
[0026] Image information storage and indexing: The recognized text information is stored in association with the original image data, and the coordinate position of the text in the image is recorded, including the line number and pixel position. After the index is established, the user can quickly locate the original image by searching for the text content, and find the exact location where the text appears in the image.
[0027] Audio information storage and indexing: The text information after speech recognition is associated with the original audio file and stored, recording the specific time point of the text in the audio, including the start time and end time; through indexing, users can jump directly to the time point of related content in the audio file through keyword search, realizing accurate audio content positioning.
[0028] Video information storage and indexing: The extracted keywords are stored in association with the original video, and the frames or time periods corresponding to the keywords in the video are recorded. Through the indexing system, when searching for keywords, users can directly locate the relevant frames or time points of the video, thereby quickly obtaining important content in the video.
[0029] Furthermore, the TF-IDF algorithm calculates the frequency of occurrence of each word, including calculating the word frequency and calculating the inverse document frequency. The formula for the word frequency is expressed as: Among them, t d represents the number of times word t appears in document d, n d represents the total number of words in document d, and TF(t,d) represents the frequency of word t in document d;
[0030] The formula for the inverse document frequency is expressed as: Where N represents the total number of documents, and df(t) represents the number of documents containing word t;
[0031] The comprehensive calculation formula of the TF-IDF algorithm is: TF-IDF=TF(t,d)×IDF(t,d).
[0032] All parsed text information is cleaned and standardized, including the removal of stop words, punctuation marks, and special characters. The TF-IDF (term frequency-inverse document frequency) algorithm is used to calculate the frequency of each word in the document and the inverse document frequency in all documents, and select words with higher weights as keywords. In the case of a large amount of text, this method can extract valuable information and identify keywords that appear frequently in chat records, which is also significantly helpful for analysis after forensics.
[0033] Based on the semantics and contextual information of keywords, the K-means clustering algorithm is used to cluster keywords and group similar keywords into one category. The classified keywords are indexed and stored in association with the original data so that users can quickly find relevant information by keyword. In the forensic result viewer, relevant chat record content can be quickly viewed based on the frequency and classification of keywords.
[0034] After the evidence collection is completed, the parsed evidence collection results are read based on Web technology and filled into the preset web chat standard template, restored in the form of a web page, and the layout and display of the original information are maintained. Different conversations and chat information are sorted by time according to the original website for easy browsing.
[0035] The restored web pages provide a quick search channel, where users can search by entering keywords and directly locate relevant information. Quick search can link keywords to parsed images, audio and video, greatly reducing the analysis cost after forensics.
[0036] According to a second aspect of the present invention, a computer-readable storage medium is provided, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, the above method is implemented.
[0037] The above one or more technical solutions in the embodiments of the present application have at least one of the following technical effects:
[0038] 1. Real-time collection and analysis: Real-time analysis is performed while data is being collected to improve information processing efficiency. After evidence collection is completed, the analyzed key information can be quickly viewed.
[0039] 2. Multimedia information processing: Improve the ability to discover hidden information through image text recognition, speech-to-text and video keyword extraction technology.
[0040] 3. High-frequency keyword classification: Extract and classify high-frequency keywords from the parsed data for quick search.
[0041] 4. Result restoration and display: Restore the forensic results in the form of a web page, maintain the original layout, and provide a quick retrieval channel to improve information readability and search efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and are used together with the description to explain the principles of the present invention. It will be easy to recognize other embodiments and many expected advantages of the embodiments because they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. The same reference numerals refer to corresponding similar parts.
[0043] Figure 1 A schematic diagram of a process for quickly retrieving instant messaging chat record evidence collection results according to an embodiment of the present invention is shown.
[0044] Figure 2 A diagram showing web page forensics results according to an embodiment of the present invention is shown.
[0045] Figure 3 It is a structural diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application. DETAILED DESCRIPTION
[0046] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0047] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0048] Figure 1 The flowchart of the method for quickly retrieving instant messaging chat record evidence collection results according to an embodiment of the present invention is shown as follows: Figure 1 As shown:
[0049] S1. Real-time collection of target chat records, including text, pictures, and audio and video analysis;
[0050] On the basis of traditional web forensics, chat records can be crawled when the account is logged in. According to the characteristics and development methods of different websites, the required chat text, voice, pictures, videos and other chat attachments can be extracted through monitoring of network requests or analysis of pages.
[0051] S2, performing image information analysis on the image, wherein the image information analysis includes image preprocessing, and performing text recognition on the preprocessed image using OCR technology to extract text information hidden in the image;
[0052] The image preprocessing specifically includes: graying, binarizing, and denoising the image to improve the recognition accuracy of the OCR technology.
[0053] The grayscale conversion is to convert a color image into a grayscale image, reduce the amount of data and simplify subsequent processing. The specific method includes: gray =0.299×R+0.587×G+0.114×B, where R, G, and B are the pixel values of the red, green, and blue channels, respectively.
[0054] The binarization is to convert the grayscale image into a black and white image to make the text and background clearer. The specific method uses the global threshold method: a threshold T is preset, and the pixel values greater than the threshold T are set to white, and the pixels less than the threshold T are set to black. The specific formula is:
[0055]
[0056] The denoising is to remove the noise in the image and improve the clarity of the text. The specific method uses median filtering or Gaussian filtering, where the median filtering uses the median of the neighboring pixels to replace the central pixel value; the Gaussian filtering uses the Gaussian kernel for convolution smoothing. The specific formula is: filtered =I×G, where G represents the Gaussian kernel and I represents the input image.
[0057] In some specific embodiments, text recognition also requires pre-positioning of the text area in the image in order to perform OCR more accurately. The text area positioning methods may include: Canny edge detection algorithm, connected region analysis, and using dilation and erosion operations to increase the text area.
[0058] The processed image is used for text recognition using OCR tools. Common OCR tools include Tesseract, Google Cloud Vision API, Microsoft Azure Computer Vision API, etc.
[0059] S3, the audio and video analysis includes audio information analysis and video information analysis, wherein the audio information analysis includes audio preprocessing, and using speech recognition technology to convert the preprocessed audio content into text information;
[0060] In some embodiments, tools such as FFmpeg are used to extract audio streams from videos, DeepSpeech tools are used for speech recognition, text content obtained by image recognition and speech recognition are merged, and then NLP natural language processing technology is used to extract keywords.
[0061] S4, the video information analysis includes video preprocessing, image recognition and text extraction for each frame of the video, text conversion of the voice content in the video by combining speech recognition technology, and extraction of keywords in the video by natural language processing technology;
[0062] In some embodiments, the OpenCV library is used to read the video file and save each frame as an image file. Similarly, the image needs to be preprocessed, including grayscale, binarization, denoising, and text area detection.
[0063] S5, cleaning and standardizing the text information extracted from the image, the text information converted from the audio, and the keywords extracted from the video, calculating the frequency of each word using the TF-IDF algorithm, selecting words with higher weights to associate and store with the corresponding original data, and establishing a database index system;
[0064] The said establishing a database index system specifically includes:
[0065] Image information storage and indexing: The recognized text information is stored in association with the original image data, and the coordinate position of the text in the image is recorded, including the line number and pixel position;
[0066] Audio information storage and indexing: associate the text information after speech recognition with the original audio file and store it, recording the specific time point of the text in the audio, including the start time and end time;
[0067] Video information storage and indexing: The extracted keywords are associated with the original video and stored, and the frames or time periods corresponding to the keywords in the video are recorded.
[0068] The TF-IDF algorithm calculates the frequency of occurrence of each word, including calculating the word frequency and calculating the inverse document frequency. The formula for the word frequency is expressed as: Among them, t d represents the number of times word t appears in document d, n d represents the total number of words in document d, and TF(t,d) represents the frequency of word t in document d;
[0069] The formula for the inverse document frequency is expressed as: Where N represents the total number of documents, and df(t) represents the number of documents containing word t;
[0070] The comprehensive calculation formula of the TF-IDF algorithm is: TF-IDF=TF(t,d)×IDF(t,d).
[0071] By selecting words with higher weights through the TF-IDF algorithm, the size of the index can be significantly reduced, making information retrieval more efficient. Users can quickly locate relevant images, audio or video clips through keywords; secondly, by combining the indexes of images, audio and video, cross-modal searches can be performed, such as searching for related images or video clips through text, which improves the flexibility and comprehensiveness of information retrieval. Using database structured storage, text information is associated with the original data and specific location information is recorded (such as the coordinate position of the text in the picture, the time point in the audio, the frame or time period in the video), making the data more structured. The structured data storage method makes the data easier to manage and maintain, and facilitates operations such as update, deletion and query.
[0072] S6. Clustering the keys in the database index system using a K-means clustering algorithm, and re-associating and storing the classified database index system with the original data;
[0073] Clustering using the K-means clustering algorithm also includes initialization and iteration. Initialization selects the number of clusters (K) according to actual conditions, and then randomly selects K samples as the center of the initial cluster, followed by iteration.
[0074] During the iteration process, each sample is assigned to the nearest cluster center. The specific formula is:
[0075] C i =x j |||x j -μ i || 2 ≤||x j -μ k || 2 k≠i
[0076] Among them, C i represents the i-th cluster, which contains all samples assigned to this cluster, x j represents the jth sample, such as IF-IDF result, usually a high-dimensional vector, μ i represents the center of the i-th cluster, μ k represents the center of the kth sample, ||x j -μ i || 2 Represents sample x j With cluster center μ i The square of the Euclidean distance between them;
[0077] The clustering formula can transform the sample x j Assigned to the cluster center that is closest to it.
[0078] Next, perform cluster update and recalculate the center of each cluster to achieve the effect of iteration. The specific update formula is as follows:
[0079]
[0080] Among them, N i represents the number of samples in the i-th cluster;
[0081] Repeat the assignment and update steps until the cluster center no longer changes or the maximum number of iterations is reached.
[0082] S7. Using the forensic result viewer, the forensic results in the database index system are read and filled into the preset web chat standard template, and restored in the form of a web page, maintaining the layout and display of the original information.
[0083] After the evidence collection is completed, the parsed evidence collection results are read based on Web technology and filled into the preset web chat standard template, restored in the form of a web page, and the layout and display of the original information are maintained. Different conversations and chat information are sorted by time according to the original website for easy browsing.
[0084] The restored web pages provide a quick search channel, where users can search by entering keywords and directly locate relevant information, such as Figure 2 As shown in the figure, the information bar on the left is the search target, that is, the keyword. After clicking the keyword, the detailed content on the right will appear. Figure 2 The forensic results of web chat records are restored in the form of web pages, which is convenient for users to quickly retrieve and locate relevant information, greatly improving work efficiency. When searching, the corresponding pictures, videos and other messages can be highlighted according to the keywords entered by the user, realizing the rapid retrieval of hidden information. In summary, the rapid retrieval can be associated with the parsed pictures, audio and video through keywords, greatly reducing the analysis cost after forensics.
[0085] Reference below Figure 3 , which shows a schematic diagram of the structure of a computer system 300 suitable for implementing an electronic device of an embodiment of the present application. Figure 3 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0086] like Figure 3As shown, the computer system 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage part 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the system 300 are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0087] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed, so that a computer program read therefrom is installed into the storage section 308 as needed.
[0088] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wireline, optical cable, RF, etc., or any suitable combination of the foregoing.
[0089] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0090] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0091] The modules involved in the embodiments of the present application may be implemented by software or by hardware.
[0092] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or it may exist independently and not be assembled into the electronic device. The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device: collects the target chat records in real time, and the collected content includes text, pictures, and audio and video analysis; performs picture information analysis on the pictures, wherein the picture information analysis includes picture preprocessing, and uses OCR technology to perform text recognition on the preprocessed pictures to extract the text information hidden in the pictures; the audio and video analysis includes audio information analysis and video information analysis, wherein the audio information analysis includes audio preprocessing, and uses speech recognition technology to convert the preprocessed audio content into text information; the video information analysis includes video preprocessing, and image recognition and text extraction for each frame of the video, combined with The speech recognition technology converts the speech content in the video into text, and the natural language processing technology is used to extract the keywords in the video; the text information extracted from the picture, the text information converted from the audio, and the keywords extracted from the video are cleaned and standardized, and the TF-IDF algorithm is used to calculate the frequency of each word, and the words with higher weights are selected to be associated with the corresponding original data for storage, and a database index system is established; the K-means clustering algorithm is used to cluster the keys in the database index system, and the classified database index system is re-associated with the original data for storage; the forensic results in the database index system are read and filled into the preset web chat standard template using the forensic results viewer, and restored in the form of a web page to maintain the layout and display of the original information.
[0093] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A method for quickly retrieving instant messaging chat record evidence collection results, characterized in that: The following steps are involved: S1. Real-time collection of target chat records, including text, pictures, and audio and video analysis; S2, performing image information analysis on the image, wherein the image information analysis includes image preprocessing, and performing text recognition on the preprocessed image using OCR technology to extract text information hidden in the image; S3, the audio and video analysis includes audio information analysis and video information analysis, wherein the audio information analysis includes audio preprocessing, and using speech recognition technology to convert the preprocessed audio content into text information; S4, the video information analysis includes video preprocessing, image recognition and text extraction for each frame of the video, text conversion of the voice content in the video by combining speech recognition technology, and extraction of keywords in the video by natural language processing technology; S5, cleaning and standardizing the text information extracted from the image, the text information converted from the audio, and the keywords extracted from the video, calculating the frequency of each word using the TF-IDF algorithm, selecting words with higher weights to associate and store with the corresponding original data, and establishing a database index system; S6. Clustering the keys in the database index system using a K-means clustering algorithm, and re-associating and storing the classified database index system with the original data; S7. Using the forensic result viewer, the forensic results in the database index system are read and filled into the preset web chat standard template, and restored in the form of a web page, maintaining the layout and display of the original information.
2. The method according to claim 1, characterized in that The image preprocessing specifically includes: graying, binarizing, and denoising the image to improve the recognition accuracy of the OCR technology.
3. The method according to claim 1, characterized in that The said establishing a database index system specifically includes: Image information storage and indexing: The recognized text information is stored in association with the original image data, and the coordinate position of the text in the image is recorded, including the line number and pixel position; Audio information storage and indexing: associate the text information after speech recognition with the original audio file and store it, recording the specific time point of the text in the audio, including the start time and end time; Video information storage and indexing: The extracted keywords are associated with the original video and stored, and the frames or time periods corresponding to the keywords in the video are recorded.
4. The method according to claim 1, characterized in that: The TF-IDF algorithm calculates the frequency of occurrence of each word, including calculating the word frequency and calculating the inverse document frequency. The formula for the word frequency is expressed as: Among them, t d represents the number of times word t appears in document d, n d represents the total number of words in document d, and TF(t,d) represents the frequency of word t in document d; The formula for the inverse document frequency is expressed as: Where N represents the total number of documents, and df(t) represents the number of documents containing word t; The comprehensive calculation formula of the TF-IDF algorithm is: TF-IDF=TF(t,d)×IDF(t,d).
5. A computer program product, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
6. A computing system, characterized in that: The method comprises a processor and a memory, wherein the processor is configured to execute the method according to any one of claims 1 to 4.