Solid state disk pre-reading method and device based on host I / O software and medium
Through host I/O software, the access prediction model is used to dynamically adjust the pre-reading strategy of the solid-state drive, solving the problems of inaccurate and inefficient pre-reading in the existing technology, and achieving more efficient data reading and system resource utilization.
Patent Information
- Application Number
- CN202510086160.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The existing solid-state drive pre-read technology cannot fully match the variable data access mode, resulting in invalid pre-reading, increasing data read latency and waste of cache resources.
The host I/O software collects application information, user operation behavior information and system resource information, integrates it into context information, and enters the pre-constructed access prediction model for prediction, and dynamically adjusts the pre-read strategy to optimize data read into the cache.
It improves data reading efficiency, reduces invalid pre-reading, optimizes system resource utilization, significantly accelerates data reading speed, and improves user experience.
Smart Images

Figure CN120010775A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of solid state hard disks, and in particular to a solid state hard disk pre-reading method, device and medium based on host I / O software. Background Art
[0002] At present, some solid-state drives use a sequential pre-reading strategy to optimize data access performance. Based on the assumption that data is usually accessed sequentially, the reading speed is accelerated by pre-reading continuous data blocks. However, in complex application scenarios, this strategy exposes obvious limitations. For example, database queries and multimedia editing operations often require random access to data rather than sequential reading, which results in a large amount of unnecessary data being pre-read into the cache, occupying valuable cache space and possibly preventing the really needed data from entering the cache in time, thereby increasing the read latency and affecting the overall performance. In addition, some solid-state drives try to adjust the pre-reading strategy based on historical access frequency, giving priority to high-frequency access data blocks. However, this method is difficult to adapt to changes in application scenarios in real time. For example, the operation modes of office software users on documents vary significantly in different time periods. It is difficult to accurately predict actual needs by relying solely on historical access frequencies, which can easily cause inaccurate pre-reading problems and further limit data reading efficiency. It can be seen that the existing pre-reading technology (including sequential pre-reading and pre-reading based on historical access frequency) often leads to invalid pre-reading and increases data reading latency because it cannot fully match the changing data access mode. Especially when system resources are tight, the problem of cache resource waste or insufficient resource allocation for key applications caused by excessive pre-reading is particularly prominent. These limitations often cause users to experience slow file loading and data editing when using devices equipped with solid-state drives. Therefore, it is necessary to propose a new pre-reading mechanism to more efficiently utilize cache resources, improve data reading efficiency, and meet the needs of diverse application scenarios. Summary of the invention
[0003] The embodiments of the present invention provide a solid state hard disk pre-reading method, device and medium based on host I / O software, aiming to solve the problems of poor accuracy and low efficiency of existing solid state hard disk pre-reading.
[0004] In a first aspect, an embodiment of the present invention provides a solid state hard disk pre-reading method based on host I / O software, which is applied to a solid state hard disk, and the method includes:
[0005] receiving context information collected by the I / O software from the host;
[0006] Inputting the context information into a pre-built access prediction model for prediction to obtain pre-access related parameters;
[0007] Dynamically adjusting the current pre-reading strategy according to the pre-access related parameters;
[0008] A data pre-reading operation is performed according to the adjusted pre-reading strategy to read the data into a cache.
[0009] In a second aspect, an embodiment of the present invention provides a solid state hard disk pre-reading method based on host I / O software, which is applied to a host, and the method includes:
[0010] Obtain application information, user operation behavior information, and system resource information;
[0011] The application information, the user operation behavior information and the system resource information are integrated into context information in a standard data format, and the context information is sent to the solid state drive so that the solid state drive performs pre-reading according to the context information.
[0012] In a third aspect, an embodiment of the present invention further provides a solid state hard disk pre-reading method device based on host I / O software, comprising a unit for executing the method of the first aspect or the second aspect.
[0013] In a fourth aspect, an embodiment of the present invention further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method of the first aspect or the second aspect when executing the computer program.
[0014] In a fifth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program, when executed by a processor, can implement the method of the first aspect or the second aspect mentioned above.
[0015] The embodiment of the present invention provides a method, device and medium for pre-reading a solid-state hard disk based on host I / O software. The method is applied to a host, and the method includes: obtaining application information, user operation behavior information and system resource information; integrating the application information, the user operation behavior information and the system resource information into context information in a standard data format, and sending the context information to a solid-state hard disk, so that the solid-state hard disk performs pre-reading according to the context information. The method is applied to a solid-state hard disk, and the method includes: receiving context information collected by the I / O software of the host; inputting the context information into a pre-built access prediction model for prediction to obtain pre-access related parameters; dynamically adjusting the current pre-reading strategy according to the pre-access related parameters; performing a pre-reading operation of data according to the adjusted pre-reading strategy to read the data into a cache. The present application collects various types of context information through the host I / O software and transmits it to the solid-state hard disk. The solid-state hard disk analyzes and processes the context information with the help of the access prediction model to predict the pre-access related parameters, thereby dynamically adjusting the pre-reading strategy accordingly, so that the pre-read data is more in line with actual needs, improves the pre-reading hit rate, improves the overall data reading efficiency, and significantly speeds up the data reading speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0017] Figure 1 A schematic diagram of a process flow of a solid state drive pre-reading method based on host I / O software provided by an embodiment of the present invention;
[0018] Figure 2 A simplified logic processing diagram of a solid state drive pre-reading method based on host I / O software provided in an embodiment of the present invention;
[0019] Figure 3 A schematic flow chart of a solid state drive pre-reading method based on host I / O software provided by another embodiment of the present invention;
[0020] Figure 4 A simplified logic processing diagram of a solid state disk pre-reading method based on host I / O software provided by another embodiment of the present invention;
[0021] Figure 5 A schematic block diagram of a solid state disk pre-reading device based on host I / O software provided by an embodiment of the present invention;
[0022] Figure 6 A schematic block diagram of a solid state disk pre-reading device based on host I / O software provided by another embodiment of the present invention;
[0023] Figure 7 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0026] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0027] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0028] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0029] With the rapid development of information technology, solid state drives (SSDs) have become increasingly important in the storage field and are widely used in many electronic devices and various computing scenarios. However, despite the many advantages of SSDs, their pre-reading technology still faces many challenges in practical applications, and the rich information held by the host I / O software has not been fully utilized to coordinately optimize the pre-reading process, which limits the further improvement of SSD performance.
[0030] The sequential pre-reading strategy currently adopted by some solid-state drives is designed based on the assumption that data access is usually sequential. Its operation method is generally to pre-read a certain number of continuous data blocks in a fixed order. However, in actual complex application scenarios, this strategy has obvious drawbacks. For example, in database applications, query operations often involve randomly jumping to different storage locations based on indexes to obtain data, rather than sequential reading; in multimedia editing software, when users process video, audio or image materials, they may frequently switch operations between different time points and different clips, and need to randomly access the corresponding data blocks. In these cases, the large amount of data obtained by sequential pre-reading is not actually required, which will take up valuable cache space, resulting in the inability to store the really needed data in the cache in time, thereby increasing data reading delays and affecting overall performance.
[0031] At present, some SSDs determine the pre-reading strategy based on the historical access frequency, that is, counting the number of times the data blocks have been accessed in the past, and giving priority to pre-reading those data blocks that are accessed frequently. However, this method is difficult to adapt to the current usage scenario changes of the application in real time. Taking office software as an example, the user's purpose of operating the same document in different time periods may be different. In the morning, it may be the first editing and creation, and frequently access the beginning of the document to add content, while in the afternoon it may be an overall proofreading, and it is necessary to randomly check the content of various locations in the document. Relying solely on the historical access frequency, it is difficult to accurately judge the data that should actually be pre-read at the moment, and it is easy to have inaccurate pre-reading, which ultimately restricts the data reading efficiency of the SSD.
[0032] Host I / O software plays a key role in the data interaction process of the entire computer system. As a bridge between the operating system and storage devices (including SSDs), it can obtain a rich variety of information. On the one hand, it can know the details of each data request issued by the application, such as which application the request comes from, the type of data requested, the specific data range, and the time of the request. On the other hand, it can also capture the user's operation behavior at the operating system level in real time, such as the use of keyboard shortcuts, the click of the mouse in different interface areas, and other operations. The intention information contained in the operation. At the same time, the host I / O software can also fully grasp the overall resource usage of the system, including CPU usage, memory usage, disk I / O and other indicators. However, most of the current SSD pre-reading technologies do not fully cooperate with the host I / O software, fail to fully explore and use this comprehensive information to optimize the pre-reading strategy, and miss a good opportunity to further improve the reading performance of SSDs, resulting in the failure to maximize the advantages of SSDs in practical applications.
[0033] Therefore, the embodiments of the present invention propose a solid-state hard disk pre-reading method, device and medium based on host I / O software to solve the problem of insufficient pre-reading accuracy, improve data reading efficiency, optimize system resource utilization, ensure stable system operation, and enhance user experience.
[0034] The specific solution is to establish a deep coordination mechanism between the host I / O software and the SSD, make full use of the comprehensive context information such as the application running status, user operation intention and system resources collected by the host I / O software, and use machine learning algorithms for precise analysis, so as to dynamically and accurately adjust the pre-reading strategy to ensure that the pre-read data is highly matched with the actual subsequent access requirements, thereby significantly reducing the amount of invalid pre-read data, effectively improving the data reading efficiency of the SSD, and allowing the SSD to respond to data access requests more quickly in various application scenarios. In this way, on the one hand, based on the real-time system resource information transmitted by the host I / O software, the SSD can intelligently adjust the pre-read data amount, trigger timing and other strategies according to the overall resource tension of the system and the pre-reading operation space supported by the remaining resources, and cooperate with the optimized cache management mechanism to avoid invalid occupation of cache resources, prevent excessive pre-reading from having a negative impact on system performance, ensure that the system can run stably under different load conditions, and improve the comprehensive utilization efficiency of system resources. On the other hand, by improving data reading efficiency and optimizing system resource utilization, the application response speed is significantly accelerated. Whether it is daily office operations, multimedia processing or running large applications, it can effectively shorten user waiting time and allow users to have a smoother and more efficient operating experience, promoting solid-state drives to better serve various users and application scenarios.
[0035] See also Figure 1 , Figure 1 It is a flow chart of a solid state drive pre-reading method based on host I / O software provided by an embodiment of the present invention. The solid state drive pre-reading method based on host I / O software is described in detail below, and the method is applied to the host. Among them, the host I / O software, as a key intermediate layer for the interaction between the operating system and the solid state drive, is responsible for collecting multi-dimensional context information and passing it to the solid state drive through a communication channel, while receiving feedback information from the solid state drive to collaboratively optimize the pre-reading operation of the solid state drive. The host I / O software includes an application information parsing module, a user operation behavior capture module, a system resource monitoring module and a communication interface module. The application information parsing module is used to obtain application identification information with the help of operating system related interfaces to determine the type, use function interception technology to analyze the application's operation on the file system, determine its running stage, and provide application-level data access references for subsequent pre-reading strategies. The user operation behavior capture module is used to integrate with the operating system user interface and input device driver, capture keyboard and mouse operations, interpret user intentions, such as identifying shortcut keys, tracking mouse behavior, etc., to assist in determining possible data access requirements. The system resource monitoring module is used to integrate the operating system's built-in performance monitoring tools and hardware monitoring data, collect the usage of the system as a whole and hardware-related resources, and let the SSD know the current resource status in order to reasonably plan the pre-reading strategy. The communication interface module is used to establish a dedicated communication channel with the SSD, and realize efficient and stable data interaction based on a custom communication protocol to ensure the two-way flow of information. Figure 1 As shown, the method includes the following steps S110-S120.
[0036] S110, obtaining application information, user operation behavior information and system resource information;
[0037] In this embodiment, the host I / O software integrates multiple functional modules, including an application information parsing module, a user operation behavior capture module, and a system resource monitoring module. The application information is obtained through the application information parsing module, the user operation behavior information is obtained through the user operation behavior capture module, and the system resource information is obtained through the system resource monitoring module. Rich context information is collected from different angles, and then transmitted to the solid-state drive in real time through a custom communication protocol and a specially established communication channel to ensure the comprehensiveness and timeliness of the information. For example, the application information parsing module can accurately determine the operating stage of the application, the user operation behavior capture module can accurately capture the user's operation intention, and the system resource monitoring module can fully grasp the system resource status.
[0038] In one embodiment, the step S110 includes: S111 - S113.
[0039] S111. Obtaining the process identifier and executable file path of the running application in real time through the process management interface provided by the system to determine the type of the application; and intercepting the function call of the application to access the file system, recording the data block location, operation time, and operation type of each operation to determine the running stage of the application;
[0040] In this embodiment, the process management interface includes functions such as Enum Processes and OpenProcess in Windows environment and file access interfaces related to the / proc directory in Linux environment. These process management interfaces are provided by the operating system, and various information of the application can be obtained through the interface. The application information parsing module first obtains detailed identification information of the running application in real time through the process management interface provided by the operating system, including the process identifier (PID), executable file path, etc., so as to determine the type of the application, such as office software, graphic processing software, database management system, etc. Then, the application information parsing module uses function interception technology, such as dynamic link library injection or system hook technology, to intercept and analyze the function calls of the application accessing the file system (such as Read File, Write File, Seek, etc. in Windows, read, write, lseek, etc. in Linux), and records the data block location, operation time, operation type (read or write) and other information of each operation, so as to determine which operation stage the application is in, such as file loading stage, data editing stage, saving stage, etc. Of course, it is understandable that in addition to obtaining application file system operation information through the above-mentioned function interception technology, the application audit function provided by the operating system can also be used (some operating systems come with audit logs to record key application operations) to extract the required application running stage and data access related information from the audit log.
[0041] S112, monitoring the user's keyboard operation, obtaining the pressing and releasing operations of a single key and a shortcut key combination; monitoring the user's mouse operation, obtaining the mouse movement track, click position, click number, and long press operation;
[0042] In this embodiment, the user operation behavior capture module is deeply integrated with the user interface layer and input device driver (keyboard driver, mouse driver, etc.) of the operating system. Among them, for keyboard operation, not only the pressing and releasing of a single key is monitored, but also complex shortcut key combinations can be identified to obtain the corresponding keyboard operation information. Based on this, the user's next possible data access needs can be inferred based on the pre-established shortcut key and application function mapping table, such as "Ctrl+F" corresponding to the search function in the word processing software, "Ctrl+S" corresponding to the save function, etc.), combined with the current focus and running state of the application. For mouse operation, the mouse movement trajectory, click position, number of clicks, and long press operation behavior information are accurately tracked. Based on this, in the graphics processing software, if it is detected that the mouse has been right-clicked multiple times in a certain image area and a specific editing option is selected, it can be judged that the user is likely to frequently access the data in the area and its surroundings in the future. Of course, it is understandable that in addition to being deeply integrated with the user interface layer and input device driver (keyboard driver, mouse driver, etc.) of the operating system, the user operation behavior capture module can also use the event monitoring mechanism of the operating system, such as the message queue mechanism in Windows. By registering a specific message processing function to capture message events related to user operations to obtain user operation behavior information, instead of directly integrating with the input device driver, the purpose of capturing user operations and analyzing their intentions can also be achieved. Therefore, by capturing user operation behaviors, rich context information can be obtained.
[0043] S113. Obtain CPU usage, memory usage, disk I / O, memory bandwidth, and hard disk interface bandwidth through a monitoring tool.
[0044] In this embodiment, the system resource monitoring module integrates multiple resource monitoring methods. On the one hand, the CPU usage, memory usage, disk I / O and other indicators can be obtained through the performance monitoring tools provided by the operating system, such as Windows' Performance Monitor; under Linux, command tools such as top, htop, and iostat can be used to obtain corresponding resource information and regularly collect the overall resource usage of the system. On the other hand, the system resource monitoring module interacts with the hardware monitoring chip on the motherboard to obtain hardware-related resource information such as memory bandwidth and hard disk interface bandwidth, so as to fully grasp the system resource status, and organize and standardize this information.
[0045] S120, integrating the application information, the user operation behavior information and the system resource information into context information in a standard data format, and sending the context information to a solid state drive, so that the solid state drive performs pre-reading according to the context information.
[0046] In this embodiment, the communication interface module is responsible for establishing a dedicated communication channel with the solid-state drive, and adopts NVMe / SAS / SATA and other protocols to achieve efficient and stable data interaction. The communication interface module encapsulates the collected application information, user operation behavior information and system resource information in the standard data format specified by the protocol, and sends it to the solid-state drive in real time, and receives feedback information from the solid-state drive, such as pre-reading execution status, cache status, etc., so that each module can further adjust the working strategy according to the feedback. Specifically, first, each type of information extracts key data according to its characteristics, such as the ID and type of the application, the timestamp and behavior type of the user operation, the utilization rate and status of system resources, etc. Then, these data are encoded in a predefined format to ensure that the information is structured and compact. Finally, the encoded data is encapsulated into a package (i.e., context information) and transmitted to the solid-state drive in real time, so as to facilitate the subsequent accurate analysis of each application scenario, optimize the pre-reading strategy, and improve data processing efficiency.
[0047] In summary, if Figure 2 , through this embodiment, during the operation of the system, the application information parsing module, the user operation behavior capture module and the system resource monitoring module work in parallel and continuously collect the information they are responsible for. Once new information is collected, it is immediately passed to the communication interface module, which integrates and encodes this information to obtain context information and then sends it to the solid-state drive to maintain the real-time and consistency of information transmission. For example, when the user performs a series of operations in the graphics processing software, the user operation behavior capture module quickly captures the relevant behaviors, the application information parsing module also updates the current running status information of the software, and the system resource monitoring module obtains the system resource usage at this time. This information is sent to the solid-state drive via the communication interface module in a very short time, triggering the subsequent analysis and strategy adjustment process.
[0048] See also Figure 3 , Figure 3It is a flow chart of a solid-state hard disk pre-reading method based on host I / O software provided by an embodiment of the present invention. The solid-state hard disk pre-reading method based on host I / O software is described in detail below, and the method is applied to a solid-state hard disk. Among them, the function of the solid-state hard disk is mainly to receive context information transmitted by the host I / O software, adjust the pre-reading strategy accordingly after analysis and processing, and coordinate the cache management module to optimize the cache to improve data reading efficiency and overall system performance. The solid-state hard disk includes a communication receiving module, a context analysis module, a pre-reading strategy adjustment module and a cache management collaboration module. The communication receiving module is responsible for receiving context information from the host I / O software, and after verification, unpacking and other operations, different types of information are stored in the corresponding buffer, providing an accurate data basis for subsequent modules. The context analysis module is used to integrate and extract the received context information features, use a machine learning prediction model to predict pre-access related parameters based on feature vectors, and output results to provide a basis for pre-reading strategy adjustment, and can update the model in real time to adapt to changes. The pre-reading strategy adjustment module is used to dynamically adjust the pre-reading strategy in terms of the pre-reading address range, data volume, and triggering timing according to the prediction results, to ensure that the pre-reading operation meets the actual data access requirements and avoids resource waste and performance issues. The cache management collaboration module is used to establish multi-data records for cache data blocks based on context, manage the cache through a comprehensive cache replacement algorithm and preheating and cleaning mechanisms, ensure that the cache stores data that meets the current application access requirements, and improve cache utilization efficiency. Figure 3 As shown, the method includes the following steps S210-S240.
[0049] S210, receiving context information collected by I / O software from a host;
[0050] In this embodiment, the communication receiving module is responsible for receiving various context information sent by the host I / O software through the communication channel, performing verification and unpacking operations on the received data, and storing the parsed application information, user operation behavior information and system resource information in corresponding buffers respectively, to ensure the integrity and accuracy of the data and provide a reliable data basis for the processing of subsequent modules.
[0051] S220, inputting the context information into a pre-built access prediction model for prediction to obtain pre-access related parameters;
[0052] In this embodiment, the pre-built access prediction model adopts at least one of the following algorithms: vector machine, decision tree, recurrent neural network, naive Bayes algorithm, convolutional neural network. Of course, it can be understood that other algorithms can also be used, as long as the pre-access related parameters can be accurately predicted based on the extracted context feature vector. In addition, in terms of model training, different training data set division methods can be used, such as dividing the training set, verification set and test set in chronological order, or using stratified sampling division, etc., as well as different model evaluation indicators. In addition to the common accuracy rate, recall rate, etc., F1 value, mean square error, etc. can also be considered to optimize the model performance so that it can better adapt to the application scenario of this embodiment. The model parameters of the access prediction model of this embodiment are obtained by offline training based on a large amount of actual application scenario monitoring data, and are stored in the non-volatile storage area of the solid-state hard disk, such as a specific reserved partition in the flash memory. The pre-trained model parameters are used to predict the next pre-access related parameters. In addition, as the system continues to run and generate new context information, incremental learning techniques, such as random gradient descent variants in online learning algorithms, can be used to update the model in real time so that it can dynamically adapt to changes in application programs and user behaviors.
[0053] In one embodiment, the step S220 includes: S221-S222.
[0054] S221, performing feature extraction on the context information according to a preset feature extraction rule to obtain a plurality of feature information, and combining the plurality of feature information to obtain a feature vector;
[0055] S222. Use the feature vector as the input of a pre-constructed access prediction model, perform prediction through the access prediction model, and output pre-access related parameters, wherein the pre-access related parameters include the pre-accessed data block position, access order, and data volume.
[0056] In this embodiment, the preset feature extraction rules refer to a series of standards set according to historical data analysis and application scenario requirements, which are used to determine which information can be extracted as features. The feature vector is a numerical sequence composed of these extracted features in a certain order, which is convenient for subsequent model processing. In the data access process, in order to more accurately predict the access behavior of users or applications, it is necessary to extract valuable feature information from context information. First, the collected context information is screened according to the preset feature extraction rules to extract multiple feature information reflecting user behavior habits, application characteristics and system load conditions. Then, these feature information are standardized according to their importance and relevance and combined into a feature vector. The feature vector is input into the pre-built access prediction model, which calculates the most likely future access path according to the internal algorithm and outputs the corresponding pre-access related parameters. The pre-access related parameters refer to the results output by the model, specifically including the location of the data block expected to be accessed, the access order, and the expected amount of data to be accessed. Since the access prediction model can accurately predict the user's access intention based on the feature vector in the current environment, it can help the SSD prepare the required data in advance, avoid unnecessary cache occupation, effectively reduce latency, enhance the overall performance of the system, and ultimately achieve a smoother user experience.
[0057] In one embodiment, the step S221 specifically includes: extracting the type of application, the current operating stage, and the discrete degree of recent request data from the application information as application feature information according to the application information extraction rule; extracting the trigger frequency and trigger time interval of key shortcut key usage and specific area mouse operation from the user operation behavior information as user operation behavior feature information according to the user operation behavior information extraction rule; extracting the CPU idle rate, memory remaining space ratio, solid state drive cache availability, and hard disk interface bandwidth remaining amount from the system resource information as system resource feature information according to the system resource information extraction rule; and combining the application feature information, the user operation behavior feature information, and the system resource feature information into a feature vector.
[0058] Specifically, in order to optimize the data pre-reading strategy of the solid-state drive, it is necessary to extract key feature information from three aspects: application, user operation behavior and system resources to accurately predict pre-access related parameters. Through detailed feature extraction rules, it is ensured that the selected features can fully reflect the actual needs of the current application scenario, thereby improving data access efficiency. Among them, application feature information includes application type (such as office software, multimedia editing tools, etc.), current operation stage (such as editing, saving or closing, etc.), and the discrete degree of recent request data (measurement of whether data requests are concentrated). This information helps to understand the behavior pattern of the application. User operation behavior feature information includes the frequency of use of key shortcut keys and the trigger frequency and time interval of mouse operations in specific areas. This type of information reveals the user's operating habits and is of great value in predicting subsequent operations. System resource feature information involves CPU idle rate, memory remaining space ratio, solid-state drive cache availability and hard disk interface bandwidth remaining, which is used to evaluate the overall load status of the system and its impact on data access performance. According to their respective information extraction rules, the type, current running stage and data request discreteness are extracted from the application information; the shortcut key usage, mouse operation frequency and time interval are obtained from the user operation behavior information; and the state parameters of the CPU, memory, cache and hard disk interface are collected from the system resource information. Then, the above three types of feature information are standardized and combined into a comprehensive feature vector in a predetermined order. By accurately extracting and combining feature information from different dimensions to form a feature vector, the characteristics of the current application scenario can be fully described. Based on such a feature vector, the subsequent access prediction model can more accurately identify potential pre-access related parameters. For example, if it is detected that an application is in a high-load editing stage and the user frequently uses a specific shortcut key, and the system resources are tight, it can be inferred that the upcoming data access will be intensive, and the pre-reading strategy can be adjusted accordingly to prioritize the loading of data blocks that may be needed. This precise matching not only reduces the waste of resources caused by invalid pre-reading, but also improves the data access speed, improves the user experience, and achieves an overall improvement in system performance.
[0059] S230, dynamically adjusting the current pre-reading strategy according to the pre-access related parameters;
[0060] In this embodiment, the pre-reading strategy refers to the rules and algorithms for pre-loading data into the cache of the solid-state drive in order to increase the data reading speed. For example, it includes specific measures such as which data blocks to select for pre-reading and when to start the pre-reading operation. First, the pre-access related parameters output by the access prediction model are obtained. Then, the specific content of these parameters is analyzed, such as whether the data blocks to be accessed are concentrated, whether the access sequence has a specific pattern, and the estimated amount of data required. Next, based on these analysis results, the existing pre-reading strategy is dynamically adjusted. For example, if the prediction shows that the data blocks to be accessed are relatively scattered, the strategy can be adjusted to increase support for random reading; if the estimated amount of data to be accessed is large, more cache space is reserved in advance. By adjusting the pre-reading strategy according to the pre-access related parameters in real time, the needs of actual applications can be met more accurately.
[0061] In one embodiment, the step S230 includes: S231-S236.
[0062] S231, if the access sequence is sequential access and the current system resources are sufficient, expanding the pre-read continuous address range and increasing the pre-read data amount according to the storage unit multiple of the flash memory;
[0063] S232: If the access order is random access, locate the corresponding discrete flash memory address according to the data block position, and search and update the pre-read flash memory physical address through a preset address mapping table;
[0064] In this embodiment, the access sequence refers to the data access mode determined according to the prediction result output by the context analysis module, including sequential access and random access. The sufficiency of system resources can be judged by setting corresponding resource thresholds, for example, the memory usage rate is lower than 70%, the cache idle rate is higher than 30%, and the hard disk interface bandwidth allows and other conditions, which are used to judge whether the current system resources are sufficient. The flash memory physical storage structure, such as the page size of NAND flash memory, is 4KB, 8KB, etc., and the block size is 128KB, 256KB, etc. These parameters affect how to reasonably increase the amount of pre-read data. The address mapping table is used to find and update the pre-read flash memory physical address to ensure that the data block position with high probability of access can be accurately located. Specifically, if the access sequence is predicted to be sequential access and the system resources are sufficient, the corresponding resource thresholds (such as memory usage, cache idle rate, etc.) can be set to judge whether the system resources are sufficient, then the pre-read continuous address range is expanded, and the pre-read data amount is increased according to the storage unit multiple of the flash memory (such as page or block size). For example, if the initial setting is to pre-read 10 consecutive data blocks starting from the current position, after the range is expanded, 20 or more consecutive data blocks may be pre-read, so that more data that may be needed can be loaded into the cache in advance, reducing the response time of future continuous read requests, thereby improving data processing efficiency and speed. If the access sequence is predicted to be random access, the corresponding discrete flash memory address is accurately located according to the predicted high probability access data block position. The pre-read flash memory physical address is searched and updated using the address mapping table to avoid unnecessary pre-reading of data at irrelevant addresses, thereby improving the pre-read hit rate. For example, suppose a multimedia editing software is processing a video file, and the prediction result shows that the next operation will be mainly sequential access. At the same time, the system checks and finds that the memory usage rate is 60% and the cache idle rate is 40%, which meets the condition of sufficient resources. At this time, the system will expand the pre-read continuous address range and increase the pre-read data amount according to the flash memory block size (assuming 128KB), load more required data in advance, and speed up the subsequent operations. On the contrary, if the user performs a query operation in the database management system, the prediction is displayed as a random access mode. The system will accurately locate and update the pre-read flash memory physical address based on the predicted high probability access data block location through the address mapping table, avoiding pre-reading of data at irrelevant addresses and improving the accuracy and efficiency of data access. This strategy not only reduces waiting time, but also optimizes the overall performance of the system. Through this embodiment, when the system determines that it is a sequential access and resources are sufficient, by increasing the amount of pre-read data, the response time of continuous data requests that may occur in the future can be effectively reduced, thereby improving the overall data processing efficiency. For random access, accurate address positioning and updating reduce invalid pre-reading, improve cache utilization and hit rate, and thus reduce latency.
[0065] S233, if the current system resources meet the resource sufficiency condition, increasing the data volume according to a preset ratio;
[0066] S234, if the current system resources meet the resource shortage condition, reducing the data amount, and adjusting the pre-reading order according to a preset priority rule;
[0067] In this embodiment, the resource sufficient condition may be, for example, a memory usage rate lower than a certain threshold (e.g., 80%), a cache idle rate higher than a certain ratio (e.g., 20%), etc., indicating that the system has sufficient resources for extended pre-reading. The resource shortage condition may be, for example, when the memory usage rate exceeds a certain limit (e.g., 80%), or the cache idle rate is lower than a certain level (e.g., 20%), which means that the system resources are close to saturation and the pre-reading operation needs to be strictly controlled. The preset priority rule is used to determine which applications should obtain priority pre-reading services, for example, the operating system core process and the application currently being interacted by the user are given priority pre-reading. The preset ratio is the ratio between the cache free space and the amount of pre-read data. When the system detects that the resources are sufficient, the amount of pre-read data is increased according to the preset ratio. For example, for every 10% increase in cache free space, the amount of pre-read data is increased by 15% accordingly, so as to more efficiently utilize the available resources and improve the data processing speed. If the system is in a resource shortage state, the amount of pre-read data is reduced, and the pre-reading order is adjusted according to the preset priority rule to ensure that key applications can obtain the necessary data support and avoid system freezes or other performance problems caused by excessive pre-reading.
[0068] For example, suppose a graphic design software is running, and the current system memory usage is 60%, and the cache free rate is 40%, which meets the resource sufficient condition. At this time, the system can increase the amount of pre-read data according to the set ratio. For example, for every additional 10% of cache free space, the amount of pre-read data increases by 15%, thereby loading large design files faster and improving work efficiency. On the contrary, if the system memory usage reaches 85% and the cache free rate is only 15%, it is a resource-constrained situation. At this time, the system will reduce the amount of pre-read data, and give priority to the data reading needs of the design software running in the foreground and the core processes of the operating system, to prevent excessive pre-reading of background applications from causing jamming in the foreground application, and to ensure the smoothness of the user experience. This flexible resource management method effectively balances the performance requirements in different application scenarios.
[0069] Through this embodiment, the pre-reading strategy is dynamically adjusted to fully utilize system resources to accelerate data access when resources are sufficient, and to protect the performance of key applications from being affected when resources are limited. Increasing the amount of pre-read data helps to speed up continuous data access and reduce latency; while limiting pre-reading prevents unnecessary cache occupation and ensures that the data needs of key applications are responded to in a timely manner.
[0070] S235, if the currently running application meets the real-time and predictable conditions, the triggering time of pre-accessing the pre-reading of the data is set in advance;
[0071] S236: If the currently running application meets the idle condition or the system meets the high load and resource shortage condition, the triggering timing of pre-reading the pre-access data is delayed or suspended.
[0072] In this embodiment, the real-time and predictable conditions refer to the fact that the application has a high timeliness requirement for data access, and its access mode can be accurately predicted (such as regularly executed tasks, periodic database queries, etc.). The idle condition is that the application has not generated new operation behaviors for a long time. For example, if there is no operation for more than 5 minutes, it is considered that the application is in an idle state, indicating that frequent data pre-reading is not required at present. High load and resource shortage conditions are situations such as high system memory usage (such as more than 80%) and low cache idle rate (such as less than 20%), indicating that system resources are close to saturation and resource allocation needs to be strictly controlled. If the currently running application meets the real-time and predictable conditions, the pre-reading triggering time of its pre-access data is advanced. The task scheduling mechanism of the operating system, such as the Windows task scheduler or the Linux cron scheduled task mechanism, is combined with the event notification mechanism of the application itself, such as the preparation event of the database application or the material loading notification of the graphics processing software, to accurately grasp the best pre-reading time and ensure that the data is immediately available when needed. If the currently running application meets the idle condition, or the system is in a high load and resource shortage state, the triggering time of pre-reading is delayed or suspended. In this way, limited system resources can be reasonably allocated without affecting the user experience, and the execution of key tasks can be prioritized. For example, suppose a video editing software is processing a large project, and the software has high real-time and predictability requirements for data access. After the operating system detects this situation, it triggers the pre-reading operation in advance according to the preparation event issued by the video editing software, so that the required materials can be quickly loaded during the editing process, greatly improving work efficiency. On the other hand, during non-working hours at night, if it is detected that the office document processing software has no new operation behavior for more than 5 minutes, the system will judge it as idle and delay or suspend the relevant pre-reading operation. At the same time, if the system is performing background updates or other maintenance tasks that cause resource shortages at this time, the same measures will be taken to ensure that system resources are reasonably allocated to more needed tasks, thereby maintaining the stability and efficiency of the overall system. This flexible management method effectively improves the utilization efficiency of system resources and also ensures the smoothness of user experience. Through this embodiment, the precise control of the pre-reading triggering timing can maximize the utilization efficiency of system resources without affecting the user experience. For applications with high real-time requirements, triggering pre-reading in advance can significantly reduce waiting time and improve user satisfaction; while for idle or resource-constrained situations, delaying or suspending pre-reading can help save resources and ensure the smooth execution of other important tasks.
[0073] S240: Execute a data pre-reading operation according to the adjusted pre-reading strategy to read the data into a cache.
[0074] In this embodiment, the pre-reading operation of data is performed according to the adjusted pre-reading strategy, and the location and number of data blocks to be loaded are first identified. Then, the solid-state drive controller reads the specified data block from the flash memory according to the optimized strategy and transmits it to the cache through the internal bus. In this process, if the system resources are tight, the data requirements of key applications are given priority to ensure that high-priority data can be loaded into the cache in time. For sequential access mode, the amount of data in the continuous address range is increased; for random access, discrete data blocks are accurately located and loaded. This not only improves the data access speed, but also effectively utilizes the cache space, reduces unnecessary I / O operations, and improves overall performance.
[0075] In one embodiment, the method further includes steps: S251-S252.
[0076] S251, establishing a context-associated metadata record for each data block read in the cache, the metadata record including a most recent access time, an access frequency, and a subsequent access probability evaluation value, and assigning a weight factor to the most recent access time, the access frequency, and the subsequent access probability evaluation value;
[0077] S252. If the current cache is full and new pre-read data needs to be stored, the importance of each data block is determined based on the weight factor of the most recent access time, the access frequency and the subsequent access probability evaluation value of each data block, and the data in the cache is replaced in order from low to high importance.
[0078] In this embodiment, the traditional cache replacement algorithm (such as LRU) only considers the most recent access time and frequency of the data block, and fails to fully evaluate the possibility of the data being accessed again in the future, resulting in a low cache hit rate in some cases. This embodiment introduces a subsequent access probability evaluation value and comprehensively evaluates the importance of each data block in combination with a weight factor to optimize the cache replacement strategy. The metadata record contains the most recent access time, access frequency and subsequent access probability evaluation value of each data block. Among them, the subsequent access probability evaluation value is the possibility of the data block being accessed again in the future calculated based on the prediction model. Weight factors are assigned to the indicators of the above three metadata, such as a subsequent access probability weight of 0.4, a recent access time weight of 0.3, and an access frequency weight of 0.3, which is used to quantify the impact of each factor on the importance of the data block. The cache replacement strategy of this embodiment not only considers the traditional access time and frequency, but also adds a prediction of the possibility of future access, so as to more accurately determine which data should be retained in the cache. Specifically, a metadata record is established for each data block in the cache, including the most recent access time, access frequency and subsequent access probability evaluation value, and corresponding weight factors are assigned to these attributes. When the cache is full and new pre-read data needs to be stored, the importance of each data block is calculated based on its metadata record and its weight factor. Data is replaced in order of importance from low to high, ensuring that the cache always stores the data blocks that best meet the current application requirements.
[0079] Exemplarily, assume that the metadata record of each data block is assigned weight factors as follows: the last access time is 0.3, the access frequency is 0.3, and the subsequent access probability evaluation value is 0.4. The last access time can be converted into a score from 0 to 1, where 0 means it was accessed a long time ago and 1 means it has just been accessed; the access frequency can also be converted into a similar ratio, indicating how frequently the data block is accessed relative to all possible maximum access frequencies; the subsequent access probability evaluation value adopts a value between 0 and 1 given by the context analysis module. The context analysis module can predict the output result according to the context information through the above-mentioned prediction model to obtain the subsequent access probability evaluation value. Assume that there are three data blocks A, B and C in the cache, and it is necessary to decide which one will be replaced to make room for new pre-read data. The metadata records of each data block are as follows: Data block A: the last access time is 5 minutes ago (converted to a ratio of 0.2), the access frequency is 10 times per hour (converted to a ratio of 0.8), and the subsequent access probability evaluation value is 0.6. Data block B: The last access time is 1 hour ago (converted to a ratio of 0.1), the access frequency is 5 times per hour (converted to a ratio of 0.4), and the subsequent access probability evaluation value is 0.9. Data block C: The last access time is just now (converted to a ratio of 1.0), the access frequency is 2 times per hour (converted to a ratio of 0.2), and the subsequent access probability evaluation value is 0.3. Based on the above information, we can calculate the importance score of each data block according to the following formula: Importance score = (last access time × 0.3) + (access frequency × 0.3) + (subsequent access probability evaluation value × 0.4). Calculate for each data block:
[0080] Data block A: (0.2x0.3)+(0.8x0.3)+(0.6x0.4)=0.06+0.24+0.24=0.54
[0081] Data block B: (0.1x0.3)+(0.4x0.3)+(0.9x0.4)=0.03+0.12+0.36=0.51.
[0082] Data block C: (1.0x0.3)+(0.2x0.3)+(0.3x0.4)=0.3+0.06+0.12=0.48
[0083] Therefore, according to the calculation results, data block C has the lowest importance score and should be removed from the cache as a priority to make room for new pre-read data. This method ensures that the cache stores the data blocks that best meet the current application requirements, improving the cache hit rate and system performance. By introducing subsequent access probability evaluation values and adjusting weight factors, this strategy can more accurately identify truly important data blocks and avoid frequent replacement of data that may be accessed again in the future even though it has not been used recently. This not only improves the cache hit rate and reduces unnecessary I / O operations, but also effectively improves the overall system performance and user experience.
[0084] In one embodiment, the method further includes steps: S261-S262.
[0085] S261: If the running stage of the application is initial running, determining pre-read initial data according to the type of the application, and pre-reading the initial data into a cache;
[0086] S262: If the running stage of the application is about to end, clear the space occupied by the application in the cache.
[0087] In this embodiment, when the application is started, the user often needs to wait for the data to load, which affects the user experience. At the same time, after the application ends, if the cache space occupied by it is not cleaned up in time, it may cause a waste of cache resources and affect the performance of other applications. This embodiment aims to optimize the cache pre-reading and cleaning strategy by analyzing the different running stages of the application (initial running and about to end running) to improve the overall efficiency of the system. When it is detected that the application is in the initial running stage, the initial data that is likely to be accessed is predicted and pre-read into the cache according to its type. For example, for office software, information such as configuration files and recently used document lists can be pre-loaded. If it is monitored that the application is about to end running, such as receiving a shutdown command or no operation for more than 10 minutes, the cache cleaning mechanism is triggered to release the cache space occupied by the application so that it can be reallocated to other active applications. For example, suppose a user is starting an office software, which usually loads the user's preferences and the list of recently opened documents when it starts. According to the type of this application, the system automatically pre-reads these data into the cache during its initial running stage, so that the software can quickly respond to the user's first operation, greatly shortening the startup time and improving the speed of the first access. On the other hand, when the user finishes the work and does not perform any operation on the office software for a long time (for example, more than 10 minutes), or directly issues a shutdown command, the system recognizes that the application is about to end its operation and immediately starts to clean up its related data blocks in the cache. This step ensures that the cache space is effectively released and can be used to support other applications that are running or about to be opened, thereby maintaining the efficient operation of the system. By dynamically adjusting the cache content, the best use of limited cache resources is achieved, and the user experience and system performance are improved. By pre-reading the initial data, the application startup time and the initial data access delay can be significantly reduced, the application can be "preheated", and a smoother user experience can be provided. Timely cleaning of cache data that is no longer needed ensures the effective use of cache resources and avoids the problem of new data not being able to be loaded in time due to a full cache. This method not only improves the response speed of a single application, but also enhances the resource management efficiency of the entire system.
[0088] In summary, if Figure 4, the communication receiving module of this embodiment continuously monitors the communication channel, and once the context information sent by the host I / O software is received, it notifies the context analysis module to process. The context analysis module extracts the feature vector according to the above process and inputs it into the prediction model for analysis, and passes the obtained prediction result to the pre-reading strategy adjustment module. After the pre-reading strategy adjustment module adjusts the various parameters of pre-reading according to the prediction results, it sends instructions to the main control chip of the solid-state hard disk. The main control chip initiates a pre-reading operation to the flash memory according to the new pre-reading strategy and reads the data into the cache. At the same time, the cache management collaboration module monitors the cache status in real time, and manages the data in the cache accordingly according to the cache replacement algorithm and the preheating and cleaning mechanism. The whole process is repeated, and the pre-reading and cache management work is continuously optimized according to the real-time context information. Through this embodiment, accurate pre-reading can be achieved, the data reading efficiency is greatly improved, it can adapt to complex and changeable application scenarios, optimize the utilization of system resources, ensure the stable and efficient operation of the system, significantly improve the user experience, and enhance the competitiveness of products.
[0089] Figure 5 4 is a schematic block diagram of a solid state disk pre-reading device 400 based on host I / O software provided by an embodiment of the present invention. Figure 5 As shown, corresponding to the above-mentioned solid state hard disk pre-reading method based on host I / O software, the present invention also provides a solid state hard disk pre-reading device 400 based on host I / O software. The solid state hard disk pre-reading device 400 based on host I / O software includes a unit for executing the above-mentioned solid state hard disk pre-reading method based on host I / O software, and the device can be configured in a solid state hard disk. Specifically, please refer to Figure 5 The solid state disk pre-reading device 400 based on host I / O software includes: a receiving unit 401, an access prediction unit 402, an adjustment unit 403 and a pre-reading unit 404.
[0090] Among them, the receiving unit 401 is used to receive context information collected by the I / O software from the host; the access prediction unit 402 is used to input the context information into a pre-built access prediction model for prediction to obtain pre-access related parameters; the adjustment unit 403 is used to dynamically adjust the current pre-reading strategy according to the pre-access related parameters; the pre-reading unit 404 is used to perform a pre-reading operation of data according to the adjusted pre-reading strategy to read the data into the cache.
[0091] In one embodiment, the access prediction unit 402 includes: a feature unit and a model unit.
[0092] Among them, the feature unit is used to extract features from the context information according to preset feature extraction rules to obtain multiple feature information, and combine the multiple feature information to obtain a feature vector; the model unit is used to use the feature vector as the input of a pre-constructed access prediction model, perform prediction through the access prediction model, and output pre-access related parameters, wherein the pre-access related parameters include the pre-accessed data block position, access order and data amount.
[0093] In one embodiment, the feature unit includes: a first extraction unit, a second extraction unit, a third extraction unit and a combination unit.
[0094] Among them, the first extraction unit is used to extract the type of application, the current operating stage, and the discrete degree of recent request data from the application information as application feature information according to the application information extraction rules; the second extraction unit is used to extract the trigger frequency and trigger time interval of key shortcut key usage and specific area mouse operation from the user operation behavior information as user operation behavior feature information according to the user operation behavior information extraction rules; the third extraction unit is used to extract the CPU idle rate, memory remaining space ratio, solid state drive cache availability and hard disk interface bandwidth remaining amount from the system resource information as system resource feature information according to the system resource information extraction rules; the combination unit is used to combine the application feature information, the user operation behavior feature information and the system resource feature information into a feature vector.
[0095] In one embodiment, the model unit includes: the pre-built access prediction model adopts at least one of the following algorithms: vector machine, decision tree, recurrent neural network, naive Bayes algorithm, convolutional neural network.
[0096] In one embodiment, the adjustment unit 403 includes: a first adjustment subunit, a second adjustment subunit, a third adjustment subunit, a fourth adjustment subunit, a fifth adjustment subunit and a sixth adjustment subunit.
[0097] Among them, the first adjustment subunit is used to expand the pre-read continuous address range and increase the pre-read data volume according to the storage unit multiple of the flash memory if the access order is sequential access and the current system resources are sufficient; the second adjustment subunit is used to locate the corresponding discrete flash memory address according to the data block position if the access order is random access, and search and update the pre-read flash memory physical address through a preset address mapping table; the third adjustment subunit is used to increase the data volume according to a preset ratio if the current system resources meet the resource sufficient condition; the fourth adjustment subunit is used to reduce the data volume and adjust the pre-read order according to the preset priority rule if the current system resources meet the resource shortage condition; the fifth adjustment subunit is used to advance the triggering time of pre-reading of pre-access data if the currently running application meets the real-time and predictable conditions; the sixth adjustment subunit is used to delay or suspend the triggering time of pre-reading of its pre-access data if the currently running application meets the idle condition or the system meets the high load and resource shortage conditions.
[0098] In one embodiment, the host I / O software-based solid state drive pre-reading device 400 further includes: a recording unit and a replacement unit.
[0099] Among them, the recording unit is used to establish a context-associated metadata record for each data block read in the cache, and the metadata record includes the most recent access time, access frequency and subsequent access probability evaluation value; the replacement unit is used to replace the data in the cache in the order of the subsequent access probability evaluation values from low to high if the current cache is full and new pre-read data needs to be stored.
[0100] In one embodiment, the host I / O software-based solid state drive pre-reading device 400 further includes: a preheating unit and a cleaning unit.
[0101] Among them, the preheating unit is used to determine the pre-read initial data according to the type of the application if the running stage of the application is the initial running, and pre-read the initial data into the cache; the cleaning unit is used to clean up the space occupied by the application in the cache if the running stage of the application is about to end.
[0102] Figure 6 is a schematic block diagram of a solid state disk pre-reading device 500 based on host I / O software provided by an embodiment of the present invention. Figure 6As shown, corresponding to the above-mentioned solid state disk pre-reading method based on host I / O software, the present invention also provides a solid state disk pre-reading device 500 based on host I / O software. The solid state disk pre-reading device 500 based on host I / O software includes a unit for executing the above-mentioned solid state disk pre-reading method based on host I / O software, and the device can be configured in the host. Specifically, please refer to Figure 6 The solid state drive pre-reading device 500 based on host I / O software includes: a collection unit 501 and an integration unit 502 .
[0103] Among them, the collection unit 501 is used to obtain application information, user operation behavior information and system resource information; the integration unit 502 is used to integrate the application information, the user operation behavior information and the system resource information into context information in a standard data format, and send the context information to the solid state drive so that the solid state drive can pre-read according to the context information.
[0104] In one embodiment, the collection unit 501 includes: an application information collection unit, a user operation behavior information collection unit, and a system resource information collection unit.
[0105] Among them, the application information collection unit is used to obtain the process identifier and executable file path of the running application in real time through the process management interface provided by the system to determine the type of application; and to intercept the function call of the application to access the file system, and record the data block location, operation time, and operation type of each operation to determine the running stage of the application; the user operation behavior information collection unit is used to monitor the user's keyboard operations, obtain the pressing and releasing operations of single keys and shortcut key combinations; monitor the user's mouse operations, obtain the mouse movement trajectory, click position, number of clicks, and long press operations; the system resource information collection unit is used to obtain CPU utilization, memory utilization, disk I / O, memory bandwidth, and hard disk interface bandwidth through monitoring tools.
[0106] The above-mentioned host I / O software-based solid state disk pre-reading device 400, 500 can be implemented in the form of a computer program, and the computer program can be used in a computer program such as Figure 7 Runs on the computer device shown.
[0107] See also Figure 7 , Figure 7 300 is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device 300 is a solid state hard disk.
[0108] See also Figure 7The computer device 300 includes a processor 302 , a memory and a network interface 305 connected via a system bus 301 , wherein the memory may include a non-volatile storage medium 303 and an internal memory 304 .
[0109] The non-volatile storage medium 303 can store an operating system 3031 and a computer program 3032. When the computer program 3032 is executed, the processor 302 can execute a solid state drive pre-reading method based on host I / O software.
[0110] The processor 302 is used to provide computing and control capabilities to support the operation of the entire computer device 300 .
[0111] The internal memory 304 provides an environment for the operation of the computer program 3032 in the non-volatile storage medium 303. When the computer program 3032 is executed by the processor 302, the processor 302 can execute a solid state hard disk pre-reading method based on the host I / O software.
[0112] The network interface 305 is used to communicate with other devices over the network. Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device 300 to which the solution of the present invention is applied. The specific computer device 300 may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0113] The processor 302 is configured to run a computer program 3032 stored in a memory to implement any embodiment of the above method.
[0114] It should be understood that in the embodiment of the present invention, the processor 302 may be a central processing unit (CPU), and the processor 302 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0115] It is understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiment of the above method.
[0116] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program. When the computer program is executed by a processor, the processor executes any embodiment of the above method.
[0117] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, etc., which are computer-readable storage media that can store program codes.
[0118] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0119] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0120] The steps in the method of the embodiment of the present invention can be adjusted in order, combined and deleted according to actual needs. The units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0121] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device to execute all or part of the steps of the method described in each embodiment of the present invention.
[0122] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0123] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
[0124] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A solid state hard disk pre-reading method based on host I / O software, applied to a solid state hard disk, characterized in that: The method comprises: receiving context information collected by the I / O software from the host; Inputting the context information into a pre-built access prediction model for prediction to obtain pre-access related parameters; Dynamically adjusting the current pre-reading strategy according to the pre-access related parameters; A data pre-reading operation is performed according to the adjusted pre-reading strategy to read the data into a cache.
2. The method according to claim 1, characterized in that The step of inputting the context information into a pre-built access prediction model for prediction to obtain pre-access related parameters includes: Performing feature extraction on the context information according to a preset feature extraction rule to obtain a plurality of feature information, and combining the plurality of feature information to obtain a feature vector; The feature vector is used as the input of a pre-constructed access prediction model, and prediction is performed through the access prediction model to output pre-access related parameters, wherein the pre-access related parameters include the pre-accessed data block position, access sequence and data volume.
3. The method according to claim 2, characterized in that The context information includes application information, user operation behavior information and system resource information. The step of extracting features from the context information according to a preset feature extraction rule to obtain a plurality of feature information, and combining the plurality of feature information to obtain a feature vector includes: Extracting the type of application, the current running stage, and the discrete degree of recent request data from the application information as application feature information according to the application information extraction rule; Extracting the use of key shortcut keys and the triggering frequency and triggering time interval of mouse operation in a specific area from the user operation behavior information as user operation behavior feature information according to the user operation behavior information extraction rule; According to the system resource information extraction rule, the CPU idle rate, the remaining memory space ratio, the solid state drive cache availability rate and the remaining hard disk interface bandwidth are extracted from the system resource information as the system resource feature information; The application feature information, the user operation behavior feature information and the system resource feature information are combined into a feature vector.
4. The method according to claim 2, characterized in that: The step of dynamically adjusting the current pre-reading strategy according to the pre-access related parameters includes: If the access sequence is sequential access and the current system resources are sufficient, then the pre-read continuous address range is expanded and the pre-read data amount is increased according to the storage unit multiple of the flash memory; If the access sequence is random access, the corresponding discrete flash memory address is located according to the data block position, and the pre-read flash memory physical address is searched and updated through a preset address mapping table; If the current system resources meet the resource sufficiency condition, the data volume is increased according to a preset ratio; If the current system resources meet the resource shortage condition, the data amount is reduced, and the pre-reading order is adjusted according to the preset priority rule; If the currently running application meets the real-time and predictable conditions, the pre-access data pre-reading trigger timing is advanced; If the currently running application meets the idle condition or the system meets the high load and resource shortage condition, the triggering timing of pre-reading the pre-access data is delayed or suspended.
5. The method according to claim 1, characterized in that The method further comprises: Establishing a context-associated metadata record for each data block read from the cache, the metadata record including a most recent access time, an access frequency, and a subsequent access probability evaluation value, and assigning a weight factor to the most recent access time, the access frequency, and the subsequent access probability evaluation value; If the current cache is full and new pre-read data needs to be stored, the importance of each data block is determined based on the weight factor of the most recent access time, the access frequency and the subsequent access probability evaluation value of each data block, and the data in the cache is replaced in order from low to high importance.
6. A solid state hard disk pre-reading method based on host I / O software, applied to a host, characterized in that: The method comprises: Obtain application information, user operation behavior information, and system resource information; The application information, the user operation behavior information and the system resource information are integrated into context information in a standard data format, and the context information is sent to the solid state drive so that the solid state drive performs pre-reading according to the context information.
7. The method according to claim 6, characterized in that The step of obtaining application information, user operation behavior information and system resource information includes: Through the process management interface provided by the system, the process identifier and executable file path of the running application are obtained in real time to determine the type of application; and the function call of the application accessing the file system is intercepted, and the data block location, operation time, and operation type of each operation are recorded to determine the running stage of the application; Monitor the user's keyboard operations to obtain the pressing and releasing operations of single keys and shortcut key combinations; monitor the user's mouse operations to obtain the mouse movement trajectory, click position, number of clicks, and long press operations; Use monitoring tools to obtain CPU usage, memory usage, disk I / O, memory bandwidth, and hard disk interface bandwidth.
8. A solid state hard disk pre-reading method and device based on host I / O software, characterized in that: The method comprises a unit for executing the method according to any one of claims 1 to 5, or a unit for executing the method according to any one of claims 6 to 7.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 5, or implements the method according to any one of claims 6 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the computer program can implement the method according to any one of claims 1 to 5, or implement the method according to any one of claims 6 to 7.
Citation Information
Cited By
Method, device and equipment for reading data of solid state disk and storage medium
CN120469653A
Simulator starting acceleration method and device, equipment and storage medium
CN120492053A
Solid state disk data pre-reading method based on access frequency
CN120508263A
Pre-reading method of solid state disk
CN120743182A
File processing method and electronic equipment
CN120805845A