SSD log analysis method and device based on text classification, equipment and storage medium

Through the SSD log analysis method based on text classification, the problem of insufficient flexibility and compatibility of SSD log analysis in the existing technology is solved, and efficient and intelligent processing and analysis of SSD log data is achieved, which significantly improves the accuracy and efficiency of log analysis, and ensures the reliability and performance optimization of SSD.

CN120144759APending Publication Date: 2025-06-13成都芯忆联信息技术有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510218660.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing SSD log analysis technology has insufficient flexibility and compatibility, and it is difficult to handle diversified log formats, resulting in the neglect of key information, limited event recognition and association analysis capabilities, lack of intelligent analysis engines, unable to make full use of log data, and insufficient multi-source log integration and parallel processing capabilities, making it difficult to meet the needs of real-time monitoring and fast response.

Method used

The SSD log analysis method based on text classification is adopted to achieve efficient and intelligent processing and analysis of SSD log data through classification recognition, networked text analysis model, text error correction processing, factor extraction and feature matching and other steps.

Benefits of technology

It significantly improves the accuracy and efficiency of log analysis, can deeply explore key information in the log, generate intuitive log analysis reports, helping users fully grasp the operation status of SSDs, promptly discover and deal with potential problems, and ensure the reliability and performance optimization of SSDs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144759A_ABST
    Figure CN120144759A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an SSD log analysis method, device and equipment based on text classification and a storage medium. The method comprises the steps that corresponding initial chart information is obtained according to a preset networked text analysis model, and the networked text analysis model comprises an error feature extraction formula, a character library and a matching degree threshold value; recognizing texts contained in the to-be-processed category information page according to a preset networked text analysis model to obtain corresponding initial text information; performing text error correction processing on the initial text information and the initial chart information according to a preset text review model to obtain corresponding error correction text information; and generating a log analysis report according to the result text information, wherein the log analysis report comprises a log structure chart and a text factor analysis chart which are generated based on the result text information. According to the method, efficient processing of SSD log data is achieved, and the problems of log data classification, error recognition, key information extraction and report generation are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, device, and storage medium for SSD log analysis based on text classification. Background Art

[0002] With the wide application of solid-state drives (SSDs) in the field of data storage and processing, the importance of SSD log analysis has become increasingly prominent. In data centers, enterprise-level storage systems, and large database management, SSD logs contain rich operation information and potential problem clues, which play a crucial role in preventing system failures, optimizing storage performance, and ensuring data security. However, there are many limitations in existing SSD log analysis technologies: analysis tools generally lack flexibility and compatibility, making it difficult to process diverse log formats, resulting in a large amount of key information being ignored; the ability to identify and correlate events is limited, making it difficult to accurately capture the root cause of problems; there is a lack of intelligent analysis engines, unable to fully utilize all the information in the log data; in addition, the ability to integrate and process multi-source logs in parallel is insufficient, making it difficult to meet the requirements of real-time monitoring and rapid response, and the functions of visualizing analysis results and generating reports are insufficient, limiting its application effect in actual operation and maintenance. Therefore, improving SSD log analysis technology and overcoming existing problems have become urgent problems to be solved in the field of data storage. Summary of the Invention

[0003] Embodiments of this application provide a method, apparatus, device, and storage medium for SSD log analysis based on text classification, aiming to solve the problem of poor accuracy in the process of SSD log data analysis.

[0004] In a first aspect, an embodiment of the present application provides an SSD log analysis method based on text classification, the method comprising: if a text to be analyzed of an input SSD log is received, classifying and identifying the text to be analyzed to obtain a category information page to be processed; identifying the text content contained in each category information page to be processed according to a preset networked text analysis model to obtain corresponding initial chart information, the networked text analysis model including an error feature extraction formula, a character library and a matching threshold; identifying the text content contained in each category information page to be processed according to a preset networked text analysis model to obtain corresponding initial text information; performing text error correction processing on the initial text information and the initial chart information according to a preset text review model to obtain corresponding error correction text information; extracting corresponding text element information from the error correction text information according to preset element extraction rules; obtaining text features corresponding to each text unit from each category information page to be processed, the text features including iconic text characters obtained from each text unit; extracting a text feature extraction formula for each The text character feature vector corresponding to the text feature of the text unit is calculated; the character matching degree between the text character feature vector of each text unit and the template character preset in the character library is calculated; it is determined whether the number of characters whose character matching degree between the text character feature vector is greater than the matching degree threshold is greater than zero; if the number of characters whose matching degree between the text character feature vector is greater than the matching degree threshold is greater than zero, the character with the highest matching degree between the text character feature vector is obtained as the target character matching each text character feature vector; if the number of characters whose matching degree between the text character feature vector is greater than the matching degree threshold is not greater than zero, the text unit corresponding to the text character feature vector is determined as a qualified text unit; according to the character template of the target character corresponding to each text character feature vector in the character library, the text content contained in the text unit corresponding to each text character feature vector is identified to obtain the corresponding result text information; a log analysis report is generated according to the result text information, and the log analysis report includes a log structure chart and a text factor analysis chart generated based on the result text information.

[0005] Second aspect, an SSD log analysis device based on text classification is further provided in an embodiment of the present application. The device includes a classification and recognition unit, configured to perform classification and recognition on the text to be analyzed in the input SSD log to obtain the to-be-processed category information page therein; a first content recognition unit, configured to recognize the text content included in each to-be-processed category information page according to a preset networked text analysis model to obtain the corresponding initial chart information; a second content recognition unit, configured to recognize the text content included in each to-be-processed category information page according to a preset networked text analysis model to obtain the corresponding initial text information; an error correction unit, configured to perform text error correction processing on the initial text information and the initial chart information according to a preset text review model to obtain the corresponding error-corrected text information; a feature extraction unit, configured to extract the corresponding text feature information from the error-corrected text information according to a preset feature extraction rule; a feature acquisition unit, configured to acquire the text feature corresponding to each text unit from each to-be-processed category information page; a first calculation unit, configured to calculate the text character feature vector corresponding to the text feature of each text unit according to an error feature extraction formula; a second calculation unit, configured to calculate the character matching degree between the text character feature vector of each text unit and the preset template character in the character library; a judgment unit, configured to judge whether the number of characters with a character matching degree greater than the matching degree threshold with the text character feature vector is greater than zero; a target character acquisition unit, configured to, if the number of characters with a matching degree greater than the matching degree threshold with the text character feature vector is greater than zero, acquire the character with the highest matching degree with the text character feature vector as the target character matching each text character feature vector; a text determination unit, configured to, if the number of characters with a matching degree greater than the matching degree threshold with the text character feature vector is not greater than zero, determine the text unit corresponding to the text character feature vector as a qualified text unit; a text information acquisition unit, configured to recognize the text content included in the text unit corresponding to each text character feature vector according to the character template of the target character corresponding to each text character feature vector in the character library to obtain the corresponding result text information; a report generation unit, configured to generate a log analysis report according to the result text information, and the log analysis report includes a log structure chart and a text factor analysis chart generated based on the result text information.

[0006] Third aspect, a computer device is further provided in an embodiment of the present application. The computer device includes a memory and a processor. A computer program is stored on the memory. When the processor executes the computer program, the above method is implemented.

[0007] Fourth aspect, a computer-readable storage medium is further provided in an embodiment of the present application. The storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by the processor, the above method can be implemented.

[0008] The embodiment of the present application provides an SSD log analysis method based on text classification, the method comprising: if a text to be analyzed of an input SSD log is received, classifying and identifying the text to be analyzed to obtain a category information page to be processed; identifying the text content contained in each category information page to be processed according to a preset networked text analysis model to obtain corresponding initial chart information, the networked text analysis model including an error feature extraction formula, a character library and a matching threshold; identifying the text content contained in each category information page to be processed according to a preset networked text analysis model to obtain corresponding initial text information; performing text error correction processing on the initial text information and the initial chart information according to a preset text review model to obtain corresponding error correction text information; extracting corresponding text element information from the error correction text information according to preset element extraction rules; obtaining text features corresponding to each text unit from each category information page to be processed, the text features including iconic text characters obtained from each text unit; extracting the corresponding text element information from the error correction text information according to the error feature extraction formula The text character feature vector corresponding to the text feature of the element is calculated; the character matching degree between the text character feature vector of each text unit and the template character preset in the character library is calculated; it is determined whether the number of characters whose character matching degree between the text character feature vector is greater than the matching degree threshold is greater than zero; if the number of characters whose matching degree between the text character feature vector is greater than the matching degree threshold is greater than zero, the character with the highest matching degree between the text character feature vector is obtained as the target character matching each text character feature vector; if the number of characters whose matching degree between the text character feature vector is greater than the matching degree threshold is not greater than zero, the text unit corresponding to the text character feature vector is determined as a qualified text unit; according to the character template of the target character corresponding to each text character feature vector in the character library, the text content contained in the text unit corresponding to each text character feature vector is identified to obtain the corresponding result text information; according to the result text information, a log analysis report is generated, and the log analysis report includes a log structure chart and a text factor analysis chart generated based on the result text information. This method proposes an SSD log analysis tool based on a text classification network, and realizes efficient and intelligent processing and analysis of SSD log data through refined step design. This method effectively solves the technical problems of log data classification, error identification, key information extraction and report generation, and significantly improves the intelligent level of SSD operation and maintenance management. By providing accurate analysis results and intuitive display forms, this method helps users fully understand the operation status of SSDs, discover and deal with potential problems in a timely manner, thereby ensuring the reliability and performance optimization of SSDs, and providing strong technical support for SSD maintenance and management. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0010] Figure 1 It is a schematic flowchart of the SSD log analysis method based on text classification provided by the embodiments of the present application;

[0011] Figure 2 It is a schematic sub - flowchart of the SSD log analysis method based on text classification provided by the embodiments of the present application;

[0012] Figure 3 It is another schematic sub - flowchart of the SSD log analysis method based on text classification provided by the embodiments of the present application;

[0013] Figure 4 It is a schematic block diagram of the SSD log analysis device based on text classification provided by the embodiments of the present application;

[0014] Figure 5 It is a schematic block diagram of the computer device provided by the embodiments of the present application. Specific embodiments

[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0016] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0017] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0018] It should be further understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0019] The embodiments of the present application provide a method, device, computer device, and storage medium for SSD log analysis based on text classification.

[0020] The execution subject of the SSD log analysis method based on text classification may be the SSD log analysis device provided in the embodiments of the present application. Among them, the SSD log analysis device based on text classification may be implemented in a hardware or software manner, and the above method is executed by the unit modules configured in the above device. Among them, the server may be a computer device, and both the server and the serial communication device may be implemented as a computer device as Figure 5 shown.

[0021] The SSD log analysis method based on text classification is applied to Figure 5 the computer device 500 in

[0022] Figure 1 is a schematic flowchart of the SSD log analysis method based on text classification provided by the embodiments of the present application. The method includes the following steps S110 - S1130.

[0023] S110. If the text to be analyzed of the input SSD log is received, classify and identify the text to be analyzed to obtain the page of category information to be processed therein.

[0024] S120. Identify the text content included in each page of category information to be processed according to the preset networked text analysis model to obtain the corresponding initial chart information.

[0025] The networked text analysis model includes an error feature extraction formula, a character library, and a matching degree threshold.

[0026] S130. Identify the text content included in each page of category information to be processed according to the preset networked text analysis model to obtain the corresponding initial text information.

[0027] S140. Perform text error correction processing on the initial text information and the initial chart information according to the preset text review model to obtain the corresponding corrected text information.

[0028] S150. Extract the corresponding text element information from the corrected text information according to the preset element extraction rules.

[0029] S160: Obtain text features corresponding to each text unit from each category information page to be processed, where the text features include iconic text characters obtained from each text unit.

[0030] S170, calculating the text character feature vector corresponding to the text feature of each text unit according to the error feature extraction formula.

[0031] S180, calculating the character matching degree between the text character feature vector of each text unit and the template characters preset in the character library.

[0032] S190, determining whether the number of characters whose character matching degree with the text character feature vector is greater than a matching degree threshold is greater than zero.

[0033] S1100: If the number of characters whose matching degree with the text character feature vector is greater than the matching degree threshold is greater than zero, obtain the character with the highest matching degree with the text character feature vector as the target character matching each text character feature vector.

[0034] S1110: If the number of characters whose matching degree with the text character feature vector is greater than the matching degree threshold is not greater than zero, the text unit corresponding to the text character feature vector is determined as a qualified text unit.

[0035] S1120. According to the character template of the target character corresponding to each text character feature vector in the character library, the text content contained in the text unit corresponding to each text character feature vector is recognized to obtain corresponding result text information.

[0036] S1130. Generate a log analysis report according to the result text information, where the log analysis report includes a log structure chart and a text factor analysis chart generated based on the result text information.

[0037] Specifically, the SSD log text uploaded by the user is received and the text to be analyzed is preliminarily classified using text classification technology to identify different information categories. The networked text analysis model is used to identify the text content in each information category. Initial text information and chart information are generated to provide basic data for subsequent analysis. The initial text and chart information are corrected using the text review model to ensure the accuracy of the analysis data. According to the preset element extraction rules, key information elements are extracted from the corrected text. The iconic text characters are extracted from each text unit as features. The text character feature vector is calculated using the error feature extraction formula. The matching degree between the text character feature vector and the preset template character is calculated. It is determined whether the character matching degree exceeds the threshold and whether the text unit is a qualified text unit. The characters with high matching degree are identified as target characters. According to the character template of the target character, the text content is identified and the result text information is generated. Finally, a log analysis report containing a log structure chart and a text factor analysis chart is generated. The accuracy and reliability of log analysis are improved through error correction processing and feature matching. The method realizes the automation of log analysis, reduces manual intervention, and improves analysis efficiency. Through feature extraction and feature matching, we can deeply mine the key information in the log. The generated log analysis report is presented in a graphical form, which is easy for users to understand and operate.

[0038] The specific embodiment of the method realizes efficient and intelligent analysis of SSD logs through a series of carefully designed steps, and achieves significant technical effects. First, by classifying and identifying the text to be analyzed, it ensures that log information of different categories can be correctly distinguished and processed, laying the foundation for subsequent in-depth analysis. Then, the networked text analysis model is used to identify the text content, extract the initial text and chart information, and perform error correction through the text review model, which greatly improves the accuracy and reliability of the data. On this basis, the key information elements are extracted through the element extraction rules, combined with the extraction and matching of text features, to further ensure the accuracy of the analysis results. This method can not only automatically identify and correct errors in the text, but also determine the eligibility of the text unit through character matching judgment, thereby screening out high-quality data for generating analysis reports. Finally, the log analysis report generated based on the result text information shows the structure of the log and the text factor analysis in an intuitive chart form, so that users can easily understand the operation status of the SSD and promptly discover and solve potential problems. Overall, this method not only improves the intelligence level of SSD operation and maintenance, improves the efficiency and accuracy of log analysis, but also provides strong technical support for SSD reliability assurance and performance optimization, and greatly enhances users' ability to monitor and manage the operating status of SSD hard drives.

[0039] The above method can effectively solve the massiveness and complexity faced in the process of processing log data: through text classification and feature extraction technologies, it can effectively process massive log data and simplify the analysis process. Accurate identification of log information: through error correction processing and precise character matching, it ensures the accurate identification of log information and reduces false positives and false negatives. Performance bottleneck and fault prediction: through in-depth analysis of logs, it can quickly locate performance bottlenecks and predict potential fault trends. User experience and operation convenience: it provides easy-to-understand reports and charts, reduces the technical threshold for users, and improves the user experience. The above solution, through a series of refined steps, not only solves the technical problems in SSD log analysis, but also provides an efficient and accurate analysis tool, which is of great significance for improving the SSD operation and maintenance management level and ensuring system stability.

[0040] In summary, the present method proposes an SSD log analysis tool based on a text classification network. Through refined step design, it realizes the efficient and intelligent processing and analysis of SSD log data. This method effectively solves the technical problems of log data classification, error identification, key information extraction, and report generation, and significantly improves the intelligent level of SSD operation and maintenance management. By providing accurate analysis results and intuitive display forms, this method helps users comprehensively master the operating conditions of SSDs, timely discover and handle potential problems, thereby ensuring the reliability and performance optimization of SSDs, and providing strong technical support for the maintenance and management of SSDs.

[0041] Generally speaking, in a more specific embodiment, as Figure 2 shown, when executing method S110, it further specifically includes executing steps S111 - S114.

[0042] S111. Determine whether the log topic words of each page of the document in the text to be analyzed are the same as the preset standard log topic words.

[0043] S112. If the log topic words of the document are not the same as the standard log topic words, correct the document so that the log topic words are the same as the standard log topic words.

[0044] S113. Determine whether the proportion of the text in the document corresponding to the standard log topic words is greater than the preset proportion threshold.

[0045] S114. Obtain the document with the text proportion greater than the proportion threshold and determine it as the page of the category information to be processed corresponding to the text to be analyzed.

[0046] Specifically, perform a page-by-page analysis on the input SSD log text, use natural language processing techniques to extract the keywords of each page of the document, and form a list of log topic words. Compare the extracted log topic words with a preset set of standard log topic words. The standard log topic words are predefined based on the common topics and keywords of SSD logs. For each page of the document, use an algorithm to compare whether its log topic words are exactly the same as or belong to the same category as the standard log topic words. For those documents whose log topic words do not match the standard log topic words, the system will execute a correction process. The correction process may involve using techniques such as synonym replacement, stemming, and lemmatization to adjust the keywords in the document to match the standard log topic words. If no matching standard topic words can be determined during the correction process, the system may mark the document for manual review. After the log topic words are matched or corrected, the system calculates the proportion of the text content in each document. This proportion is calculated by dividing the number of text characters in the document by the total number of characters (including spaces, punctuation, etc.). The system compares the calculated text proportion with a preset proportion threshold. This threshold is set based on experience or specific analysis requirements to ensure that the analyzed documents are information-rich. Based on the judgment result of the text proportion, the system filters out those documents whose text proportion exceeds the preset threshold. These documents are considered to be information-sufficient and relevant, and are thus confirmed as the information pages of the category to be processed. The system records the indexes or identifiers of these pages so that these pages containing key information can be preferentially processed in subsequent analysis steps. Through these refinement steps, method S110 ensures that only those documents that are content-related and information-sufficient will be sent to the subsequent analysis process, thereby improving the efficiency and accuracy of the overall log analysis tool.

[0047] For each page of the document in the text to be analyzed, first extract the log topic words. Compare the extracted log topic words with the preset standard log topic words. Determine whether the log topic words of each page of the document match the preset standard log topic words. If it is found in step S111 that the log topic words of the document do not match the standard log topic words, perform a correction operation. The correction operation may include text cleaning, keyword replacement, format adjustment, etc., to ensure that the log topic words are consistent with the standard log topic words. For the documents that have been matched or matched the standard log topic words after correction, further determine the proportion of text in these documents. Compare whether the proportion of text in the document is greater than the preset proportion threshold. The proportion threshold is a parameter preset according to the log analysis requirements, and is used to screen out documents with sufficient information. According to the judgment result of step S113, screen out the documents with the proportion of text greater than the proportion threshold. Determine these documents as the pages of the category information to be processed corresponding to the text to be analyzed, that is, it is considered that these pages contain sufficient information and are suitable for subsequent in-depth analysis and processing. Through the above steps, method S110 ensures that each page of the document in the text to be analyzed can be correctly classified and identified, laying a solid foundation for subsequent text classification network analysis and log data analysis.

[0048] In a more specific embodiment, as Figure 3 shown, executing method S140 further specifically includes executing steps S141 - S143.

[0049] S141. Perform encoding conversion on the characters included in each text segment in the initial text information according to the encoding information library in the text review model to obtain the corresponding text encoding information.

[0050] S142. Sequentially input the encoding sequences corresponding to each text segment in the text encoding information into the error correction neural network in the text review model to obtain the corresponding error correction encoding sequences.

[0051] S143. Perform inverse encoding conversion on each error correction encoding sequence according to the encoding information library to obtain the corresponding error correction text information.

[0052] Specifically, first, the system uses the encoding information library in the text review model to perform encoding conversion on the initial text information in the SSD log. The encoding information library contains the mapping relationship between characters and encodings, which can be ASCII code, Unicode code, or other custom encoding formats. The system processes each text segment one by one, and converts each character in each text segment into the corresponding encoding information according to the mapping relationship in the encoding information library. In this way, each text segment is converted into an encoding sequence, which represents the encoding form of the characters in the text segment. Then, the system sequentially inputs each encoding sequence in the obtained text encoding information into the error correction neural network in the text review model. The error correction neural network is a deep learning model that is trained to identify and correct errors in text, such as spelling mistakes and grammar errors. By learning a large amount of correct text data, the neural network can identify abnormal patterns in the input encoding sequence and output a corrected error correction encoding sequence. In this process, the neural network may use structures such as attention mechanism, recurrent neural network (RNN), or transformer to improve the accuracy and efficiency of error correction. Finally, the system performs inverse encoding conversion on each error correction encoding sequence output by the error correction neural network according to the encoding information library. The inverse encoding conversion remaps the error correction encoding sequence back to the original character form, thereby obtaining the error-corrected text information. Through this process, the errors in the original text are corrected, and more accurate and reliable error-corrected text information is generated. These error-corrected text information will be used for subsequent log analysis to ensure the accuracy and reliability of the analysis results. By performing these three steps of S140, the text errors in the SSD log are effectively identified and corrected, thereby improving the quality and usability of log data analysis. This is crucial for ensuring that the log analysis tool can accurately capture key information and identify potential problems.

[0053] In a more specific embodiment, extracting the corresponding text element information from the error-corrected text information according to the preset element extraction rules includes locating the element symbols corresponding to each element in the error-corrected text information according to the element labels corresponding to each element in the element comparison table set in the element extraction rules; verifying the element symbols corresponding to each element according to the element verification formula set in the element extraction rules to obtain a verification result of whether it passes; obtaining the element symbols with the verification result of passing and determining them as the text element information corresponding to the error-corrected text information.

[0054] Specifically, according to the preset element extraction rules, the system first refers to the element comparison table, which lists various elements to be extracted and their corresponding element tags. The system traverses the error correction text information and uses the element tags corresponding to each element in the element comparison table to locate the element symbols in the text that correspond to each element. Element symbols can be specific keywords, data formats, code snippets, etc., which mark the position and scope of elements in the text. Once the element symbols are located, the system will verify these element symbols according to the element verification formula set in the element extraction rules. The element verification formula is a set of predefined rules or patterns used to verify whether the element symbols conform to the expected format or content. The verification process may include regular expression matching, data type verification, value range checking, etc., to ensure the correctness and effectiveness of the element symbols. For each element symbol, the system will obtain a verification result of whether it passes. If the element symbol passes the verification, that is, the verification result is "passed", the system will confirm it as valid text element information. If the element symbol fails the verification, the system may mark it as an error or an anomaly, or perform further processing as needed, such as attempting to correct or ignore it. Finally, the system will extract all element symbols with a verification result of "passed" from the error correction text information and determine them as the text element information corresponding to the error correction text information. These text element information will be used to generate the final log analysis report, providing in-depth insights into the SSD operating conditions. Through this process, the system can accurately extract key information from the corrected log text, which is crucial for understanding the behavior of the SSD, diagnosing problems, and predicting future failures. This method improves the accuracy and efficiency of log analysis and provides users with more reliable and practical log analysis results.

[0055] Furthermore, before calculating the text character feature vector corresponding to the text feature of each text unit according to the error feature extraction formula, calculate the character feature vector corresponding to each character in the unrecognized document according to the error feature extraction formula; calculate the matching degree between each character feature vector and the feature vector corresponding to each character in the preset character set, and the unrecognized document is a document with a text proportion not greater than the proportion threshold.

[0056] Furthermore, before encoding and converting the characters included in each text segment in the initial text information according to the encoding information library in the text review model to obtain the corresponding text encoding information, the method includes: filtering out the disordered characters in the initial text information to obtain the corresponding ordered text information; segmenting the ordered text information according to the text statements included in the initial text information to obtain the preprocessed text information including multiple text segments.

[0057] Before calculating the text features of a text unit according to the error feature extraction formula, the feature vector calculation is first performed on each character in the unrecognized document. This involves defining a set of features for each character, such as the font, size, position, context, etc. of the character, and converting these features into a multi-dimensional character feature vector. Then, the matching degree between each character feature vector and the feature vectors corresponding to each character in the preset character set is calculated. The matching degree calculation can use various methods, such as cosine similarity, Euclidean distance, or other suitable similarity metrics. The purpose of this step is to identify the possible error characters in the document, that is, those characters with a low matching degree with the standard feature vectors in the preset character set. Before performing the encoding conversion on the initial text information, the disordered characters in the text are first filtered out. The disordered characters may be caused by transmission errors, encoding problems, or other reasons, and they may interfere with the subsequent analysis process. Through a certain algorithm or rule, these disordered characters are identified and removed to obtain ordered text information. According to the text statements contained in the initial text information, the ordered text information is segmented. The segmentation can be achieved by identifying the punctuation marks at the end of the statement, specific delimiters, or the logical structure based on the text content. In this way, the entire text is divided into multiple text segments, and each text segment is the basic unit for subsequent analysis and encoding conversion. Through these preprocessing steps, the system provides cleaner and more ordered input data for subsequent text review and error feature extraction, thereby improving the accuracy and efficiency of the entire log analysis tool. These steps ensure that text errors can be more accurately identified and corrected during the analysis process, and key information can be more effectively extracted.

[0058] Figure 4 It is a schematic block diagram of an SSD log analysis device based on text classification provided by an embodiment of the present application. As shown in the figure, corresponding to the above SSD log analysis method based on text classification, the present application also provides an SSD log analysis device 100 based on text classification. The SSD log analysis device based on text classification includes units for executing the above SSD log analysis method based on text classification. Among them, the SSD log analysis device based on text classification can be implemented in a hardware or software manner, and the above method is executed through the unit modules configured in the above device, where the server can be a computer device. Specifically, please refer to Figure 4, the SSD log analysis device 100 based on text classification includes a classification and recognition unit 110, which is used to, if receiving the text to be analyzed of the input SSD log, classify and recognize the text to be analyzed to obtain the to-be-processed category information page therein; a first content recognition unit 120, which is used to recognize the text content included in each to-be-processed category information page according to a preset networked text analysis model to obtain the corresponding initial chart information; a second content recognition unit 130, which is used to recognize the text content included in each to-be-processed category information page according to a preset networked text analysis model to obtain the corresponding initial text information; an error correction unit 140, which is used to perform text error correction processing on the initial text information and the initial chart information according to a preset text review model to obtain the corresponding error-corrected text information; a feature extraction unit 150, which is used to extract the corresponding text feature information from the error-corrected text information according to a preset feature extraction rule; a feature acquisition unit 160, which is used to acquire the text feature corresponding to each text unit from each to-be-processed category information page; a first calculation unit 170, which is used to calculate the text character feature vector corresponding to the text feature of each text unit according to an error feature extraction formula; a second calculation unit 180, which is used to calculate the character matching degree between the text character feature vector of each text unit and the preset template character in the character library; a judgment unit 190, which is used to judge whether the number of characters with a character matching degree greater than the matching degree threshold with the text character feature vector is greater than zero; a target character acquisition unit 1100, which is used to, if the number of characters with a matching degree greater than the matching degree threshold with the text character feature vector is greater than zero, acquire the character with the highest matching degree with the text character feature vector as the target character matching each text character feature vector; a text determination unit 1110, which is used to, if the number of characters with a matching degree greater than the matching degree threshold with the text character feature vector is not greater than zero, determine the text unit corresponding to the text character feature vector as a qualified text unit; a text information acquisition unit 1120, which is used to recognize the text content included in the text unit corresponding to each text character feature vector according to the character template of the target character corresponding to each text character feature vector in the character library to obtain the corresponding result text information; a report generation unit 1130, which is used to generate a log analysis report according to the result text information, and the log analysis report includes a log structure chart and a text factor analysis chart generated based on the result text information.

[0059] Among them, the classification and recognition unit further includes a subject word judgment unit for judging whether the log subject word of each page of the document in the text to be analyzed is the same as the preset standard log subject word; a correction unit for correcting the document to make the log subject word the same as the standard log subject word if the log subject word of the document is not the same as the standard log subject word; a threshold judgment unit for judging whether the proportion of the text in the document corresponding to the standard log subject word is greater than the preset proportion threshold; a threshold determination unit for obtaining the document with the text proportion greater than the proportion threshold and determining it as the page of the to-be-processed category information corresponding to the text to be analyzed.

[0060] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above SSD log analysis device based on text classification and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the convenience and conciseness of description, they will not be elaborated here.

[0061] The above SSD log analysis device based on text classification can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 5 shown.

[0062] Please refer to Figure 5 , which shows a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 executes the above method through the unit modules configured in the device. The server can be an independent server or a server cluster composed of multiple servers.

[0063] The computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.

[0064] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions. When the program instructions are executed, the processor 502 can execute an SSD log analysis method based on text classification.

[0065] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0066] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute an SSD log analysis method based on text classification.

[0067] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand,Figure 5 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. Specifically, the computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0068] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0069] Those of ordinary skill in the art can understand that all or part of the processes in the methods of implementing the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0070] Therefore, this application also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, where the computer program includes program instructions. When the program instructions are executed by the processor, the processor is caused to execute the following steps:

[0071] If the text to be analyzed of the input SSD log is received, the text to be analyzed is classified and identified to obtain the category information page to be processed; the text content contained in each category information page to be processed is identified according to the preset networked text analysis model to obtain the corresponding initial chart information, and the networked text analysis model includes an error feature extraction formula, a character library and a matching threshold; the text content contained in each category information page to be processed is identified according to the preset networked text analysis model to obtain the corresponding initial text information; the initial text information and the initial chart information are corrected according to the preset text review model to obtain the corresponding error-corrected text information; the corresponding text element information is extracted from the error-corrected text information according to the preset element extraction rules; the text features corresponding to each text unit are obtained from each category information page to be processed, and the text features include the iconic text characters obtained from each text unit; the text character features corresponding to the text features of each text unit are extracted according to the error feature extraction formula The method comprises the following steps: calculating a quantity of characters; calculating a character matching degree between a text character feature vector of each text unit and a template character preset in a character library; judging whether the number of characters whose character matching degree between the text character feature vector and the template character preset in a character library is greater than zero; if the number of characters whose matching degree between the text character feature vector and the template character preset in a character library is greater than zero, obtaining the character with the highest matching degree between the text character feature vector and the template character preset in a character library as a target character matching each text character feature vector; determining a text unit corresponding to the text character feature vector as a qualified text unit if the number of characters whose matching degree between the text character feature vector and the template character preset in a character library is not greater than zero; identifying the text content contained in the text unit corresponding to each text character feature vector according to the character template of the target character corresponding to each text character feature vector in the character library, and obtaining the corresponding result text information; generating a log analysis report according to the result text information, the log analysis report comprising a log structure chart and a text factor analysis chart generated based on the result text information.

[0072] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, etc., which are computer-readable storage media that can store program codes.

[0073] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0074] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0075] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of this application can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0076] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application.

[0077] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A SSD log analysis method based on text classification, characterized in that: The method comprises: If the input text to be analyzed of the SSD log is received, the text to be analyzed is classified and identified to obtain a category information page to be processed therein; Identifying the text content contained in each of the to-be-processed category information pages according to a preset networked text analysis model to obtain corresponding initial chart information, wherein the networked text analysis model includes an error feature extraction formula, a character library, and a matching degree threshold; Identifying the text content contained in each category information page to be processed according to the preset networked text analysis model to obtain corresponding initial text information; Performing text error correction processing on the initial text information and the initial chart information according to a preset text review model to obtain corresponding error correction text information; Extracting corresponding text element information from the error correction text information according to a preset element extraction rule; Acquire text features corresponding to each text unit from each of the category information pages to be processed, wherein the text features include iconic text characters acquired from each of the text units; Calculating the text character feature vector corresponding to the text feature of each of the text units according to the error feature extraction formula; Calculating the character matching degree between the text character feature vector of each text unit and the template characters preset in the character library; Determining whether the number of characters whose character matching degree with the text character feature vector is greater than the matching degree threshold is greater than zero; If the number of characters whose matching degree with the text character feature vector is greater than the matching degree threshold is greater than zero, obtaining the character with the highest matching degree with the text character feature vector as the target character matching each of the text character feature vectors; If the number of characters whose matching degree with the text character feature vector is greater than the matching degree threshold is not greater than zero, determining the text unit corresponding to the text character feature vector as a qualified text unit; According to the character template of the target character corresponding to each of the text character feature vectors in the character library, the text content contained in the text unit corresponding to each of the text character feature vectors is recognized to obtain corresponding result text information; A log analysis report is generated according to the result text information, wherein the log analysis report includes a log structure chart and a text factor analysis chart generated based on the result text information.

2. The SSD log analysis method based on text classification according to claim 1 is characterized in that: The step of classifying and identifying the text to be analyzed to obtain the category information page to be processed includes: Determining whether the log subject words of each page of the document in the text to be analyzed are the same as the preset standard log subject words; If the log subject words of the document are not the same as the standard log subject words, correcting the document so that the log subject words are the same as the standard log subject words; Determine whether the text ratio in the document corresponding to the standard log keyword is greater than a preset ratio threshold; The document whose text ratio is greater than the ratio threshold is obtained and determined as the category information page to be processed corresponding to the text to be analyzed.

3. The SSD log analysis method based on text classification according to claim 1 is characterized in that: The performing text error correction processing on the initial text information and the initial chart information according to the preset text review model to obtain corresponding error correction text information includes: According to the encoding information library in the text review model, the characters contained in each text segment in the initial text information are encoded and converted to obtain corresponding text encoding information; Inputting the coding sequence corresponding to each of the text segments in the text coding information into the error correction neural network in the text review model in sequence to obtain the corresponding error correction coding sequence; According to the coding information library, each error correction coding sequence is subjected to inverse coding conversion to obtain corresponding error correction text information.

4. The SSD log analysis method based on text classification according to claim 1 is characterized in that: The extracting corresponding text element information from the error correction text information according to the preset element extraction rule includes: Locate the error correction text information and the element symbol corresponding to each element according to the element label compared to each element in the element comparison table set in the element extraction rule; Testing the element symbol corresponding to each element according to the element test formula set in the element extraction rule to obtain a test result of whether it passes or not; The element symbol with a passed inspection result is obtained and determined as text element information corresponding to the error correction text information.

5. The SSD log analysis method based on text classification according to claim 1 is characterized in that: Before calculating the text character feature vector corresponding to the text feature of each text unit according to the error feature extraction formula, the method includes: The character feature vector corresponding to each character in the unrecognized document is calculated according to the error feature extraction formula, and the unrecognized document is a document in which the text ratio is not greater than the ratio threshold. The matching degree between each of the character feature vectors and the feature vector corresponding to each character in the preset character set is calculated.

6. The SSD log analysis method based on text classification according to claim 3 is characterized in that: Before converting the characters contained in each text segment in the initial text information according to the encoding information library in the text review model to obtain corresponding text encoding information, the method includes: Filter out-of-order characters in the initial text information to obtain corresponding ordered text information; The ordered text information is segmented according to the text sentences contained in the initial text information to obtain preprocessed text information containing a plurality of text segments.

7. A device for extracting element information based on text classification, characterized in that: The device comprises: A classification and identification unit, configured to classify and identify the text to be analyzed to obtain a category information page to be processed therein if a text to be analyzed of the input SSD log is received; A first content recognition unit is used to recognize the text content contained in each of the to-be-processed category information pages according to a preset networked text analysis model to obtain corresponding initial chart information; A second content recognition unit is used to recognize the text content contained in each of the to-be-processed category information pages according to a preset networked text analysis model to obtain corresponding initial text information; An error correction unit, used for performing text error correction processing on the initial text information and the initial chart information according to a preset text review model to obtain corresponding error correction text information; An element extraction unit, used to extract corresponding text element information from the error correction text information according to a preset element extraction rule; A feature acquisition unit, used for acquiring text features corresponding to each text unit from each of the category information pages to be processed; A first calculation unit, used for calculating a text character feature vector corresponding to a text feature of each of the text units according to the error feature extraction formula; A second calculation unit, used for calculating the character matching degree between the text character feature vector of each text unit and the template characters preset in the character library; A judging unit, used to judge whether the number of characters whose character matching degree with the text character feature vector is greater than the matching degree threshold is greater than zero; A target character acquisition unit, configured to acquire, if the number of characters whose matching degree with the text character feature vector is greater than the matching degree threshold is greater than zero, the character with the highest matching degree with the text character feature vector as a target character matching each of the text character feature vectors; A text determination unit, configured to determine a text unit corresponding to the text character feature vector as a qualified text unit if the number of characters having a matching degree greater than the matching degree threshold with the text character feature vector is not greater than zero; A text information acquisition unit, used to recognize the text content contained in the text unit corresponding to each of the text character feature vectors according to the character template of the target character corresponding to each of the text character feature vectors in the character library, and obtain the corresponding result text information; A report generating unit is used to generate a log analysis report according to the result text information, wherein the log analysis report includes a log structure chart and a text factor analysis chart generated based on the result text information.

8. The device for extracting element information based on text classification according to claim 7, characterized in that: The classification and identification unit also includes: A subject word judgment unit, used to judge whether the log subject words of each page of the document in the text to be analyzed are the same as the preset standard log subject words; a correction unit, configured to correct the document so that the log subject words are the same as the standard log subject words if the log subject words of the document are not the same as the standard log subject words; A threshold judgment unit, used to judge whether the text ratio in the document corresponding to the standard log subject word is greater than a preset ratio threshold; The threshold determination unit is used to obtain a document whose text ratio is greater than the ratio threshold and determine it as a to-be-processed category information page corresponding to the to-be-analyzed text.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the SSD log analysis method based on text classification as described in any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the SSD log analysis method based on text classification as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Element information extraction method and device based on text recognition, equipment and medium

    CN113536771A

  • Error log processing method and device, electronic equipment and readable storage medium

    CN117874236A

  • Structured document analyzing device and method, structured document analyzing program, and storage medium with the program stored

    JP2003296344A

  • Text error correction method and apparatus, computer-readable storage medium and system

    WO2021212614A1