Action parameter value determination method and device, equipment and medium

By building a signal transmission parameter adjustment model and dynamically adjusting the data transmission parameters of SSD, the signal integrity and transmission reliability problems of SSD during high frequency and large data transmission are solved, and more efficient and reliable data transmission is achieved.

CN119987665APending Publication Date: 2025-05-13SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510065113.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When SSD is transmitted at high frequency and large data volume, it faces the problems of signal integrity and transmission reliability, which affects its performance and life.

Method used

By obtaining the SSD's status parameter set and action parameter value set, a signal transmission parameter adjustment model is built, and the data transmission channel and chip enable signals are dynamically adjusted to optimize data transmission.

Benefits of technology

It improves the transmission efficiency and performance of SSD, speeds up data reading speed, improves the accuracy of data reading and system flexibility and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987665A_ABST
    Figure CN119987665A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses an action parameter value determination method and device, equipment and a medium, and the method comprises the steps: obtaining state parameter sets of an SSD at a previous moment and a current moment, an instant reward value of the previous moment, and a first action parameter value set; constructing a signal transmission parameter adjustment model according to the parameters; inputting the first state data combination, the instant reward value, the second state data combination and the corresponding different action parameter value set into a model to obtain an expected reward value; when it is determined that the preset condition is not met, the first action parameter value set is adjusted; determining read data through the data transmission channel influenced by the second action parameter value set and the chip enable signal; performing consistency analysis on the read data and the original data; and when the accuracy cannot meet an iteration stopping condition, taking the first analysis result as a reward for feedback, stopping the iteration operation until the accuracy reaches a preset threshold value, and obtaining an optimal action parameter value set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, device, equipment and medium for determining an action parameter value. Background Art

[0002] With the rapid development of big data and cloud computing technologies, solid state drives (SSDs) are increasingly being used in the field of data storage. SSDs are gradually replacing traditional mechanical hard drives and becoming the mainstream storage device due to their advantages such as high speed, low power consumption, and no noise.

[0003] However, when SSDs transmit large amounts of data at high frequencies, they often face problems with signal integrity and transmission reliability, which affects their performance and lifespan. Summary of the invention

[0004] In view of this, the present invention provides a method, device, equipment and medium for determining an action parameter value to solve the problem in the related art that the related parameters cannot be dynamically adjusted to improve the SSD performance.

[0005] In a first aspect, the present invention provides a method for determining an action parameter value, the method comprising:

[0006] In the current iteration loop, the state parameter set of the SSD at the previous moment, the state parameter set at the current moment, the instant reward value corresponding to the previous moment, and the first action parameter value set corresponding to the state parameter set at the previous moment are obtained, wherein the state parameter set includes an operating frequency, a data transmission channel, a chip enable signal, and a target storage chip corresponding to the chip enable signal, and the action parameter value set includes at least one of a physical parameter of the data transmission channel and a conversion number of each data transmission, and the parameters in the state parameter set and the parameters in the action parameter value set correspond to multiple parameter values ​​respectively. When the parameters in the state parameter set and the parameters in the action parameter value set correspond to different parameter values ​​respectively, different state data combinations are constituted;

[0007] According to the state parameter set of SSD at the previous moment, the immediate reward value corresponding to the previous moment, the state parameter set at the current moment, and the first action parameter value set, a signal transmission parameter adjustment model is constructed;

[0008] Inputting the first state data combination of the SSD at the previous moment, the instant reward value corresponding to the previous moment, the second state data combination corresponding to the SSD at the current moment, and the different action parameter value sets corresponding to the second state data combination into the signal transmission parameter adjustment model to obtain the expected reward value of the SSD at the previous moment, wherein the first state data combination is any state data combination among the different state data combinations corresponding to the SSD at the previous moment, and the operation frequency in the second state data combination is the same as the operation frequency in the first state data combination;

[0009] When it is determined that the preset condition is not satisfied according to the expected reward value and the target reward value at the previous moment, the first action parameter value set is adjusted to generate a second action parameter value set;

[0010] Determining to read data in a target storage chip through a data transmission channel and a chip enable signal affected by a second action parameter value set;

[0011] Performing data consistency analysis on the read data and the original data to obtain a first analysis result, where the first analysis result is used to indicate the accuracy of the data read from the target storage chip through the data transmission channel;

[0012] When the accuracy of the read data cannot meet the condition for stopping iteration, the first analysis result is fed back as a reward to select a new set of action parameter values. After the data transmission channel is adjusted based on the final action parameter value set obtained, the accuracy of the read data reaches a preset threshold, and the iterative operation is stopped to obtain the optimal action parameter value set corresponding to the first state data combination at the previous moment.

[0013] The method for determining an action parameter value provided by the present invention has the following advantages:

[0014] In the above method steps, the state parameter set of the SSD at the last moment, the state parameter set at the current moment, the instant reward value corresponding to the last moment, and the first action parameter value set corresponding to the state parameter set at the last moment are obtained, and a signal transmission parameter adjustment model is constructed based on these parameters. By obtaining the state parameter set, instant reward value and action parameter value set of the SSD at the last moment and the current moment, the working state and performance of the SSD can be understood. Then, the expected reward value at the last moment is obtained using the model, and based on the expected reward value and the instant reward value, it is determined whether the first action parameter value model needs to be adjusted. Thereby optimizing the data transmission channel. Finally, by continuously iteratively adjusting the action parameter value set and performing data consistency analysis, the optimal action parameter value set under a specific state can be found, which can improve the transmission efficiency and performance of the SSD and speed up the data reading speed. Performing consistency analysis on the read data and adjusting the action parameters according to the feedback of the analysis results can help improve the accuracy of data reading and reduce data errors. Adaptive adjustment is performed according to the real-time state of the SSD to adapt to different working conditions and requirements, thereby improving the flexibility and adaptability of the system.

[0015] In an optional implementation, the first state data combination of the SSD at the previous moment, the immediate reward value corresponding to the previous moment, the second state data combination corresponding to the SSD at the current moment, and the different action parameter value sets corresponding to the second state data combination are input into the signal transmission parameter adjustment model to obtain the expected reward value of the SSD at the previous moment, specifically including:

[0016] Determine the maximum expected reward value corresponding to the second state data combination at the current moment according to the second state data combination corresponding to the SSD at the current moment and each action parameter value corresponding to the second state data combination;

[0017] The expected reward value of the SSD at the previous moment is determined according to the first state data combination, the immediate reward value corresponding to the previous moment, and the maximum expected reward value at the current moment corresponding to the second state data combination.

[0018] Specifically, by inputting the second state data combination and the corresponding action parameter value set at the current moment into the transmission parameter adjustment model, the maximum expected reward value corresponding to the state combination at the current moment can be obtained. The best action under different states is predicted more accurately, thereby improving the performance and efficiency of the SSD. The expected reward value is determined according to the first state data combination, the instant reward value at the previous moment, and the maximum expected reward value at the current moment corresponding to the second state data combination. The decision-making process is made more scientific and reasonable, and better decisions can be made based on historical data and current status, further optimizing the performance of the SSD. Adaptive adjustment is performed according to the real-time status of the SSD. By continuously inputting new state data combinations and action parameter value sets, the model can dynamically learn and adapt to changes in the SSD, improving the flexibility and adaptability of the system. Accurate prediction and optimization decisions help improve the reliability of the SSD. By selecting the best action parameter value set, errors and delays in data transmission can be reduced, and the accuracy and integrity of data can be improved. By determining the expected reward value, system resources can be used more effectively. The model can select the most valuable action based on the prediction results, avoid unnecessary waste of resources, and improve the overall efficiency of the system. That is, the above method steps combine the state data and action parameters of the SSD, and make accurate predictions and optimize decisions through the transmission parameter adjustment model, thereby improving the performance, reliability and resource utilization of the SSD.

[0019] In an optional implementation, a signal transmission parameter adjustment model is constructed according to the state parameter set of the SSD at the previous moment, the state parameter set at the current moment, and the first action parameter value set, and is expressed by the following formula:

[0020] Q1(s,a)←Q(s,a)+α[γmax a' Q(s',a')-Q(s,a)]

[0021] Among them, Q1(s,a) is the expected reward value corresponding to the previous moment, Q(s,a) is the immediate reward value corresponding to the previous moment, s is the first state data combination, a is the action parameter value set corresponding to the previous moment, α is the learning rate, γ is the reward decay coefficient, s' is the second state parameter combination corresponding to the current moment, Q(s',a') is the maximum expected reward value at the current moment generated at the current moment according to the different action parameter value sets corresponding to the second state parameter combination, among which the learning rate and the reward decay coefficient are both constants.

[0022] In an optional implementation, when it is determined that the preset condition is not satisfied according to the expected reward value and the target reward value at the previous moment, adjusting the first action parameter value set to generate the second action parameter value set specifically includes:

[0023] Determine the absolute difference between the expected reward value at the previous moment and the target reward value;

[0024] When the absolute difference is greater than the target difference, it is determined that the preset condition is not met;

[0025] Determine the adjustment trend based on the difference between the expected reward value at the previous moment and the target reward value;

[0026] According to the adjustment trend and the preset action parameter value adjustment amount, the first action parameter value set is adjusted to generate a second action parameter value set.

[0027] Specifically, by determining the absolute difference between the expected reward value and the target reward value, the gap between the current state and the target state can be accurately judged, so that adjustments can be made more targeted. When the absolute difference is greater than or equal to the target difference, it is determined that the preset condition is not met to ensure the necessity and rationality of the adjustment. According to the size relationship between the expected reward value and the target reward value, the adjustment trend is determined, so that the adjustment direction is clearer, which helps to reach the target state faster. According to the adjustment trend and the preset action parameter value adjustment amount, the first action parameter value set is adjusted to generate a second action parameter value set. This moderate adjustment can avoid over-adjustment or under-adjustment and improve the accuracy and effectiveness of the adjustment. By continuously adjusting the action parameter value set, the performance of the SSD is gradually optimized, and the data transmission efficiency and accuracy are improved. The method can be dynamically adjusted according to the actual operation of the SSD and the target requirements, and adapt to different working conditions and demand changes. The optimized action parameter value set helps to improve the reliability of the SSD and reduce the occurrence of data transmission errors and failures. By reasonably adjusting the action parameter value set, the resources of the SSD are fully utilized to improve the overall performance and efficiency of the system.

[0028] In an optional implementation, when the first action parameter value set includes both the physical parameters of the data transmission channel and the number of conversions of each data transmission, the method further includes:

[0029] Based on the default value of the first action parameter in the first action parameter value set, the second action parameter is adjusted according to the preset action parameter value adjustment amount corresponding to the second action parameter other than the first action parameter, wherein when the first action parameter is a physical parameter of a data transmission channel, the second action parameter is the number of conversions for each data transmission; or, when the first action parameter is the number of conversions for each data transmission, the second action parameter is the physical parameter of the data transmission channel;

[0030] Using the data transmission channel and the chip enable signal after the second action parameter is adjusted, determining to read data in the target storage chip;

[0031] Performing data consistency analysis on the read data and the original data to obtain a second analysis result;

[0032] When the data are determined to be consistent according to the second analysis result, determining a window lower limit of a parameter value corresponding to the second action parameter;

[0033] Continue to adjust the second action parameter according to the preset action parameter value adjustment amount corresponding to the second action parameter until it is determined that the read data is inconsistent with the original data, and then determine the window upper limit of the parameter value corresponding to the second action parameter;

[0034] Determine an initial parameter value corresponding to the second action parameter according to the window upper limit and the window lower limit;

[0035] The operation is stopped after the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions for each data transmission are determined respectively.

[0036] Specifically, based on the default value of the first action parameter, the second action parameter is adjusted according to the preset adjustment amount, so that the action parameter can be accurately controlled. By adjusting the physical parameters of the data transmission channel and the number of conversions of each data transmission, it is possible to adapt to different storage chips and working conditions, and improve the efficiency and accuracy of data transmission. The read data is analyzed for consistency to ensure the accuracy and integrity of the data. Only when the data is consistent, the lower limit of the window of the parameter value is determined, which further ensures the reliability of the adjustment. By continuously adjusting the second action parameter, the upper and lower limits of the window of the parameter value are determined, thereby determining the initial parameter value. In this way, the optimal parameter setting can be found to improve the performance of the SSD. By determining the initial values ​​of the physical parameters of the data transmission channel and the number of conversions of each data transmission, the performance of the SSD can be optimized, and the data transmission speed and reliability can be improved. The method can flexibly adjust the action parameters according to different needs and situations to meet different application scenarios. By optimizing the parameter setting, the resources of the SSD can be fully utilized to improve the overall performance and efficiency of the system.

[0037] In an optional implementation, the state parameter set further includes an operating temperature and a power supply voltage of the SSD; and the method further includes:

[0038] Obtain the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage;

[0039] According to the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage, the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission are adjusted respectively.

[0040] Specifically, considering the state parameters such as the operating temperature and power supply voltage of the SSD, the physical parameters of the data transmission channel and the initial values ​​of the number of conversions can be adjusted according to the actual working conditions, thereby achieving more accurate optimization. Among them, the changes in the operating temperature and power supply voltage may affect the performance and stability of the SSD. By adjusting according to these parameters, the stability and reliability of the SSD under different working conditions can be improved. By adjusting according to the actual operating temperature and power supply voltage, the resources of the SSD can be better utilized to avoid resource waste or performance degradation caused by unreasonable parameter settings. The method can adapt to different working environments and requirements and has high flexibility. The reference temperature and reference voltage can be set according to specific circumstances to meet different optimization goals. By adjusting the initial value, the performance of the SSD can be further optimized and the efficiency and accuracy of data transmission can be improved. The changes in the operating temperature and power supply voltage may lead to an increase in the error rate of data transmission. By adjusting the parameters, the error rate can be reduced and the reliability of the data can be improved. Reasonable parameter adjustment can reduce the loss of the SSD during operation and extend its service life. Optimizing the performance of the SSD can improve the performance of the entire system and enhance user experience.

[0041] In an optional implementation, the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission are adjusted according to the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage, respectively, as shown in the following formula:

[0042] PHY new =PHY init +k T ·(TT ref )+k v ·(VV ref )

[0043] NTODT new =NTODT init +k T (TT ref )+k v (VV ref )

[0044] Among them, PHY new is the initial value of the physical parameters of the data transmission channel after adjustment, PHY init To adjust the initial value of the physical parameters of the data transmission channel before, k T and k v are the adjustment coefficients of PHY and NTODT for temperature and voltage, respectively. ref and V refare the reference temperature and reference voltage values, respectively. T is the operating temperature of the SSD at the last moment, and V is the power supply voltage of the SSD at the last moment. NTODT new The initial value of the number of conversions per data transfer after adjustment, NTODT init The initial value for the number of conversions per data transfer before adjustment.

[0045] In a second aspect, the present invention provides a device for determining an action parameter value, the device comprising:

[0046] An acquisition module is used to acquire, within a current iteration loop, a state parameter set of the SSD at a previous moment, a state parameter set at a current moment, an immediate reward value corresponding to a previous moment, and a first action parameter value set corresponding to the state parameter set at a previous moment, wherein the state parameter set includes an operating frequency, a data transmission channel, a chip enable signal, and a target storage chip corresponding to the chip enable signal; the action parameter value set includes at least one of a physical parameter of a data transmission channel and a number of conversions of each data transmission; the parameters in the state parameter set and the parameters in the action parameter value set correspond to multiple parameter values ​​respectively; when the parameters in the state parameter set and the parameters in the action parameter value set correspond to different parameter values ​​respectively, different state data combinations are constituted;

[0047] A model building module, used to build a signal transmission parameter adjustment model according to the state parameter set of the SSD at the previous moment, the immediate reward value corresponding to the previous moment, the state parameter set at the current moment, and the first action parameter value set;

[0048] a processing module, configured to input a first state data combination of the SSD at a previous moment, an immediate reward value corresponding to the previous moment, a second state data combination corresponding to the SSD at a current moment, and a set of different action parameter values ​​corresponding to the second state data combination into a signal transmission parameter adjustment model, to obtain an expected reward value of the SSD at a previous moment, wherein the first state data combination is any state data combination among different state data combinations corresponding to the SSD at a previous moment, and an operation frequency in the second state data combination is the same as an operation frequency in the first state data combination;

[0049] An adjustment module, configured to adjust the first action parameter value set and generate a second action parameter value set when it is determined that the preset condition is not satisfied according to the expected reward value and the target reward value at the previous moment;

[0050] A determination module, used to determine to read data in a target storage chip through a data transmission channel and a chip enable signal affected by a second action parameter value set;

[0051] An analysis module, used for performing data consistency analysis on the read data and the original data to obtain a first analysis result, where the first analysis result is used to indicate the accuracy of the data read from the target storage chip through the data transmission channel;

[0052] The processing module is also used to provide feedback of the first analysis result as a reward when the accuracy of the read data cannot meet the condition for stopping iteration, so as to select a new set of action parameter values, until the data transmission channel is adjusted based on the final action parameter value set obtained, and the accuracy of the read data reaches a preset threshold, then the iterative operation is stopped, and the optimal set of action parameter values ​​corresponding to the first state data combination at the previous moment is obtained.

[0053] The device for determining an action parameter value provided by the present invention has the following advantages:

[0054] The state parameter set of the SSD at the last moment, the state parameter set at the current moment, the immediate reward value corresponding to the last moment, and the first action parameter value set corresponding to the state parameter set at the last moment are obtained, and a signal transmission parameter adjustment model is constructed based on these parameters. By obtaining the state parameter set, immediate reward value, and action parameter value set of the SSD at the last moment and the current moment, the working state and performance of the SSD can be understood. Then, the expected reward value at the last moment is obtained using the model, and based on the expected reward value and the immediate reward value, it is determined whether the first action parameter value model needs to be adjusted. This optimizes the data transmission channel. Finally, by continuously iteratively adjusting the action parameter value set and performing data consistency analysis, the optimal action parameter value set under a specific state can be found, which can improve the transmission efficiency and performance of the SSD and speed up data reading. Performing consistency analysis on the read data and adjusting the action parameters based on the feedback of the analysis results can help improve the accuracy of data reading and reduce data errors. Adaptive adjustment is performed according to the real-time state of the SSD to adapt to different working conditions and requirements, thereby improving the flexibility and adaptability of the system.

[0055] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the method for determining the action parameter value of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0056] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method for determining action parameter values ​​of the first aspect or any corresponding embodiment thereof.

[0057] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the method for determining an action parameter value of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0059] Figure 1 It is a flowchart of a method for determining an action parameter value provided by an embodiment of the present invention;

[0060] Figure 2 is a flowchart of another method for determining an action parameter value provided by an embodiment of the present invention;

[0061] Figure 3 is a flowchart of another method for determining an action parameter value provided by an embodiment of the present invention;

[0062] Figure 4 is a structural block diagram of a device for determining an action parameter value provided by an embodiment of the present invention;

[0063] Figure 5 It is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0065] With the rapid development of big data and cloud computing technology, SSD is increasingly used in the field of data storage. SSD has gradually replaced traditional mechanical hard disks and become a mainstream storage device due to its advantages such as high speed, low power consumption and no noise.

[0066] However, SSDs often face problems with signal integrity and transmission reliability when transmitting data at high frequencies and large amounts of data. Fixed parameter configuration methods in related technologies are difficult to cope with complex working environments and constantly changing working conditions.

[0067] To solve the above problems, an embodiment of the present invention provides an embodiment of determining an action parameter value. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system (computer device) including, for example, a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0068] In this embodiment, a method for determining an action parameter value is provided, which can be used in the above-mentioned terminal equipment, such as a mobile phone, a tablet computer, etc. Figure 1 is a flow chart of a method for determining an action parameter value provided by an embodiment of the present invention, such as Figure 1 As shown, the process includes the following steps. Before introducing the above method steps of the present application, the preparatory work before executing the method steps of the present application is first introduced. See the following for details:

[0069] Step 1: Data setup;

[0070] A data sample of a page size (e.g., 16KB) is pre-written on DDR to ensure data consistency.

[0071] Step 2: Power on and start;

[0072] The SSD is powered on and an initial operating frequency is set to, for example, 5 MHz to ensure stability in the initial state.

[0073] Among them, this frequency selection is to ensure that the system can run stably during the initial startup phase and avoid unstable factors caused by high frequency.

[0074] Step 3: Initialize the NAND flash memory chip;

[0075] Through the interface provided by the manufacturer, configure the initialization parameters of the NAND flash memory, including the number of data transfer channels (ch), chip enable (ce), logical unit number (lun), etc. Ensure that all parameters are configured correctly so that subsequent read and write operations can proceed smoothly.

[0076] Step 4: Data writing;

[0077] The preset 16KB data is written to each channel, each chip, and each logic unit through a write operation to ensure that the data is evenly distributed.

[0078] This process needs to ensure that data is evenly distributed across different channels, chips, and logic units so that subsequent reading and analysis can provide comprehensive coverage.

[0079] After performing the above preparations, the method steps of the method embodiment of the present application are performed, as shown below:

[0080] Step S101, in the current iteration loop, obtain the state parameter set of the SSD at the previous moment, the state parameter set at the current moment, the immediate reward value corresponding to the previous moment, and the first action parameter value set corresponding to the state parameter set at the previous moment.

[0081] Specifically, the state parameter set includes an operating frequency, a data transmission channel, a chip enable signal, and a target storage chip corresponding to the chip enable signal; the action parameter value set includes at least one of a physical parameter of the data transmission channel and a conversion number of each data transmission; the parameters in the state parameter set and the parameters in the action parameter value set each correspond to multiple parameter values; when the parameters in the state parameter set and the parameters in the action parameter value set correspond to different parameter values, they constitute different state data combinations.

[0082] In a specific example, different operating frequencies can be set according to the SSD type, such as 5 MHz, 800 MHz, 1600 MHz, 2400 MHz, etc. Select different operating frequencies f∈[f min ,f max ], where f min =5MHZ, and f max It can be set according to the type of SSD and is tentatively set to 2400MHZ. The data transmission channel can include multiple channels for transmitting data. The chip enable signal is used to select the target memory chip, which includes multiple logic units. The number of transitions per data transfer (NTODT) and phy parameters (Physical Layer parameters) are important parameters in SSD, which have a direct impact on the performance of data transmission and storage. By adjusting these parameters, the working mode and performance characteristics of the SSD can be changed.

[0083] Step S102, constructing a signal transmission parameter adjustment model according to the state parameter set of the SSD at the previous moment, the instant reward value corresponding to the previous moment, the state parameter set at the current moment, and the first action parameter value set.

[0084] Specifically, a signal transmission parameter adjustment model capable of obtaining the optimal decision is constructed based on the state parameter set of the SSD at the previous moment, the immediate reward value corresponding to the previous moment, the state parameter set at the current moment, and the first action parameter value set.

[0085] In this model, the state parameter set is used to represent the environment or situation in which the system is located. The state provides the SSD with an understanding of the environment and is the basis for decision-making. Actions are actions or decisions that the system can take. The system selects appropriate actions based on the current state. Rewards are feedback signals that the system receives from the environment. Rewards can be positive (indicating beneficial results), negative (indicating unfavorable results), or zero (indicating neutral results). The goal is to maximize the accumulated rewards by selecting actions.

[0086] Step S103, input the first state data combination of the SSD at the previous moment, the immediate reward value corresponding to the previous moment, the second state data combination corresponding to the SSD at the current moment, and the different action parameter value sets corresponding to the second state data combination into the signal transmission parameter adjustment model to obtain the expected reward value of the SSD at the previous moment.

[0087] The first state data combination is any state data combination among different state data combinations corresponding to the SSD at the previous moment, and the operating frequency in the second state data combination is the same as the operating frequency in the first state data combination.

[0088] Specifically, by inputting the first state data combination of the SSD at the previous moment, the previous instant reward value, the second state data combination corresponding to the SSD at the current moment, and the different action parameter value sets corresponding to the second state data combination into the signal transmission parameter adjustment model, the expected reward value of the SSD at the previous moment can be obtained.

[0089] The first state data combination is any state data combination among different state data combinations corresponding to the SSD at the previous moment, and the operating frequency in the second state data combination is the same as the operating frequency in the first state data combination.

[0090] In a specific example, for example, the first state data combination is a state data combination consisting of different numbers of transmission channels, different chip enable signals, and different logical unit numbers (equivalent to different memory chips) at a frequency of 1200M, then the operating frequency in the second state data combination will also be 1200M.

[0091] Step S104: when it is determined that the preset condition is not satisfied according to the expected reward value and the target reward value at the previous moment, the first action parameter value set is adjusted to generate a second action parameter value set.

[0092] Specifically, if the expected reward value does not reach the preset standard, for example, the gap between the expected reward value and the target reward value is relatively large, it means that the preset condition is not met, so it is necessary to adjust the first action parameter value set to generate the second action parameter value set. Specifically, for example, an increase or decrease operation is performed on each action parameter value in the first action parameter value set, or other algorithm operations are performed to obtain a new action parameter value.

[0093] Step S105 , determining to read data in the target memory chip through the data transmission channel and the chip enable signal affected by the second action parameter value set.

[0094] Specifically, the data transmission channel and the chip enable signal actually constitute a data transmission channel, through which the target memory chip can be selected, and the logical unit number in the target memory chip can be determined according to the above configuration, so that data can be read from the specific logical unit in the target memory chip.

[0095] Step S106: Perform data consistency analysis on the read data and the original data to obtain a first analysis result.

[0096] Specifically, the first analysis result is used to indicate the accuracy of data read from the target storage chip through the data transmission channel.

[0097] Specifically, when performing consistency data analysis, it can be achieved by data consistency comparison. For example, by comparing the number of different characters at corresponding positions in two character strings. In an embodiment of the present application, for example, the character string of the original data and the character string of the read data are compared to obtain the number of different characters. Thus, the data consistency is calculated. The specific analysis result is used to indicate the accuracy of the data read through the data transmission channel to the target storage chip.

[0098] Step S107, when the accuracy of the read data cannot meet the condition for stopping iteration, the first analysis result is fed back as a reward to select a new set of action parameter values, until the data transmission channel is adjusted based on the final action parameter value set, and the accuracy of the read data reaches a preset threshold, then the iterative operation is stopped to obtain the optimal action parameter value set corresponding to the first state data combination at the previous moment.

[0099] Specifically, for example, when data consistency is a ratio, the ratio of the number of different characters to the total number of characters can be calculated to obtain a ratio value. The ratio value is then compared with the preset ratio value standard to determine whether the accuracy of the read data meets the condition for stopping iteration. If not, the first analysis result is fed back as a reward to select a new set of action parameter values ​​until the data transmission channel is adjusted based on the final set of action parameter values. After the accuracy of the read data reaches a preset threshold, the iterative operation is stopped to obtain the optimal set of action parameter values ​​corresponding to the first state data combination at the previous moment.

[0100] The method for determining the action parameter value provided in this embodiment, in the above method steps, the state parameter set of the SSD at the last moment, the state parameter set at the current moment, the instant reward value corresponding to the last moment, and the first action parameter value set corresponding to the state parameter set at the last moment are obtained, and a signal transmission parameter adjustment model is constructed based on these parameters. By obtaining the state parameter set, instant reward value and action parameter value set of the SSD at the last moment and the current moment, the working state and performance of the SSD can be understood. Then, the expected reward value at the last moment is obtained using the model, and based on the expected reward value and the instant reward value, it is determined whether the first action parameter value model needs to be adjusted. Thereby optimizing the data transmission channel. Finally, by continuously iteratively adjusting the action parameter value set and performing data consistency analysis, the optimal action parameter value set under a specific state can be found, which can improve the transmission efficiency and performance of the SSD and speed up the data reading speed. Performing consistency analysis on the read data and adjusting the action parameters according to the feedback of the analysis results helps to improve the accuracy of data reading and reduce data errors. Adaptive adjustment is performed according to the real-time state of the SSD to adapt to different working conditions and requirements, and improve the flexibility and adaptability of the system.

[0101] In this embodiment, a method for determining an action parameter value is provided, which can be used in the above-mentioned mobile terminal, such as a mobile phone, a tablet computer, etc. Figure 2 is a flow chart of another method for determining action parameter values ​​provided by an embodiment of the present invention. Figure 2 As shown, based on the above-mentioned embodiment, the first state data combination of the SSD at the previous moment, the instant reward value corresponding to the previous moment, the second state data combination corresponding to the SSD at the current moment, and the different action parameter value sets corresponding to the second state data combination are input into the signal transmission parameter adjustment model to obtain the expected reward value of the SSD at the previous moment. In a specific example, the following method steps may be included:

[0102] Step S201, according to the second state data combination corresponding to the SSD at the current moment and each action parameter value corresponding to the second state data combination, determine the maximum expected reward value corresponding to the second state data combination at the current moment.

[0103] Specifically, the second state data combination corresponding to the SSD at the current moment and each action parameter value corresponding to the second state data combination are input into the pre-built action function estimator to obtain an expected reward value. Then the maximum expected reward value is selected from multiple expected rewards. Among them, the action function estimator can be various forms of models, such as neural networks, decision trees, linear regression, etc. These models try to approximate the real action function by learning historical data.

[0104] Step S202, determining the expected reward value of the SSD at the previous moment according to the first state data combination, the instant reward value corresponding to the previous moment, and the maximum expected reward value at the current moment corresponding to the second state data combination.

[0105] Specifically, in an optional example, the expression of the signal transmission parameter adjustment model is as follows:

[0106] Q1(s,a)←Q(s,a)+α[γmax a' Q(s',a')-Q(s,a)]

[0107] Among them, Q1(s,a) is the expected reward value corresponding to the previous moment, Q(s,a) is the immediate reward value corresponding to the previous moment, s is the first state data combination, a is the action parameter value set corresponding to the previous moment, α is the learning rate, γ is the reward decay coefficient, s' is the second state parameter combination corresponding to the current moment, a' is the action parameter value in the action parameter value set at the current moment, Q(s',a') is the maximum expected reward value at the current moment generated at the current moment according to the different action parameter value sets corresponding to the second state parameter combination, wherein the learning rate and the reward decay coefficient are both constants.

[0108] Among them, when s' does not change, if a' changes, then the Q(s',a') obtained may be different. Therefore, the maximum value can be obtained from different Q(s',a') values, and then substituted into the above formula corresponding to the signal transmission parameter adjustment model to obtain the expected reward value of SSD at the previous moment.

[0109] A method for determining an action parameter value provided by an embodiment of the present invention can obtain a maximum expected reward value corresponding to the state combination at the current moment by inputting the second state data combination and the corresponding action parameter value set at the current moment into a transmission parameter adjustment model. The optimal action under different states is predicted more accurately, thereby improving the performance and efficiency of the SSD. The expected reward value is determined according to the first state data combination, the instant reward value at the previous moment, and the maximum expected reward value at the current moment corresponding to the second state data combination. The decision-making process is made more scientific and reasonable, and better decisions can be made based on historical data and the current state, further optimizing the performance of the SSD. Adaptive adjustment is performed according to the real-time state of the SSD. By continuously inputting new state data combinations and action parameter value sets, the model can dynamically learn and adapt to changes in the SSD, thereby improving the flexibility and adaptability of the system. Accurate prediction and optimization decisions help improve the reliability of the SSD. By selecting the optimal action parameter value set, errors and delays in data transmission can be reduced, and the accuracy and integrity of data can be improved. By determining the expected reward value, system resources can be used more effectively. The model can select the most valuable action based on the prediction results, avoid unnecessary waste of resources, and improve the overall efficiency of the system. That is, the above method steps combine the state data and action parameters of the SSD, and make accurate predictions and optimize decisions through the transmission parameter adjustment model, thereby improving the performance, reliability and resource utilization of the SSD.

[0110] In this embodiment, a method for determining an action parameter value is provided, which can be used in the above-mentioned mobile terminal, such as a mobile phone, a tablet computer, etc. Figure 3 is a flow chart of another method for determining an action parameter value provided by an embodiment of the present invention, such as Figure 3 As shown, based on the method corresponding to any of the foregoing embodiments, in an optional example, when it is determined that the preset condition is not satisfied according to the expected reward value and the target reward value at the previous moment, the first action parameter value set is adjusted to generate the second action parameter value set, and the following method steps can be referred to:

[0111] Step S301, determining the absolute difference between the expected reward value at the previous moment and the target reward value.

[0112] Step S302: When the absolute difference is greater than the target difference, it is determined that the preset condition is not met.

[0113] Specifically, when the absolute difference is less than or equal to the target difference, it means that the expected reward value at the previous moment is already within the upper and lower floating range corresponding to the target reward value. The first action parameter value set may not be adjusted temporarily. Go directly to the next step, that is, determine to read data in the target storage chip through the data transmission channel and chip enable signal affected by the first action parameter value set. And perform other subsequent operations.

[0114] When the absolute difference is greater than the target difference, it means that there is still a certain gap between the expected reward at the previous moment and the target reward value, that is, the action parameter value set in the current signal transmission parameter adjustment model is not optimal and further adjustment is needed.

[0115] Step S303, determining the adjustment trend according to the difference between the expected reward value and the target reward value at the last moment.

[0116] Specifically, it is equivalent to determining whether to adjust the action parameter value upward or downward based on the size between the expected reward value and the target reward value at the previous moment. For example, if the expected reward value is greater than the target reward value, you can consider adjusting the action parameter value downward. On the contrary, adjust the action parameter value upward. However, in actual applications, it is not such a mechanical operation, it may also be the opposite. It is necessary to define whether to operate upward or downward according to the actual situation. Of course, in one possibility, it is also possible to multiply by a certain coefficient, etc.

[0117] Step S304: adjusting the first action parameter value set according to the adjustment trend and the preset action parameter value adjustment amount to generate a second action parameter value set.

[0118] Specifically, the action parameter value adjustment amount can be obtained based on a large amount of experimental data statistics and is set when performing the initialization work in the previous text.

[0119] By determining the absolute difference between the expected reward value and the target reward value, the gap between the current state and the target state can be accurately judged, so that adjustments can be made more targeted. When the absolute difference is greater than or equal to the target difference, it is determined that the preset condition is not met to ensure the necessity and rationality of the adjustment. According to the size relationship between the expected reward value and the target reward value, the adjustment trend is determined, so that the adjustment direction is clearer, which helps to reach the target state faster. According to the adjustment trend and the preset action parameter value adjustment amount, the first action parameter value set is adjusted to generate a second action parameter value set. This moderate adjustment can avoid over-adjustment or under-adjustment and improve the accuracy and effectiveness of the adjustment. By continuously adjusting the action parameter value set, the performance of the SSD is gradually optimized, and the data transmission efficiency and accuracy are improved. The method can be dynamically adjusted according to the actual operation of the SSD and the target requirements, and adapt to different working conditions and demand changes. The optimized action parameter value set helps to improve the reliability of the SSD and reduce the occurrence of data transmission errors and failures. By reasonably adjusting the action parameter value set, the resources of the SSD are fully utilized to improve the overall performance and efficiency of the system.

[0120] In an optional example, based on any of the foregoing embodiments, when the first action parameter value set includes both the physical parameters of the data transmission channel and the number of conversions of each data transmission, the method may further include the following method steps:

[0121] Step a1: based on the default value of the first action parameter in the first action parameter value set, the second action parameter is adjusted according to the preset action parameter value adjustment amount corresponding to the second action parameter other than the first action parameter.

[0122] When the first action parameter is a physical parameter of a data transmission channel, the second action parameter is the number of conversions for each data transmission; or when the first action parameter is the number of conversions for each data transmission, the second action parameter is the physical parameter of the data transmission channel.

[0123] Step a2, using the data transmission channel and the chip enable signal after the second action parameter is adjusted, to determine to read data in the target storage chip.

[0124] Step a3: Perform data consistency analysis on the read data and the original data to obtain a second analysis result.

[0125] Step a4: when it is determined according to the second analysis result that the data are consistent, a lower limit of a window of a parameter value corresponding to the second action parameter is determined.

[0126] Step a5, continue to adjust the second action parameter according to the preset action parameter value adjustment amount corresponding to the second action parameter, until it is determined that the read data is inconsistent with the original data, and determine the window upper limit of the parameter value corresponding to the second action parameter.

[0127] Step a6: determining an initial parameter value corresponding to the second action parameter according to the window upper limit and the window lower limit.

[0128] Step a7, the operation is stopped until the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission are determined respectively.

[0129] Specifically, the first action parameter takes the NTODT parameter as an example, and the second action parameter takes the phy parameter (such as read_dqs_delay in the phy parameter) as an example, and the default value is its minimum value. Then, the NTODT value and the number of opens are also set to the minimum value. When read_dqs_delay gradually increases, if the NTODT value and the number of opens can make the read operation read correctly at this time, the configuration is recorded as the window offline until an error occurs in the read operation reading data, and the configuration is also recorded, and the initial parameter value corresponding to the second action parameter is determined according to the upper and lower limits of the window.

[0130] Similarly, the initial parameter value corresponding to the second action parameter is obtained in a similar manner as above.

[0131] In an optional example, an initial parameter value corresponding to the second action parameter is determined according to the window upper limit and the window lower limit, as shown in the following formula:

[0132]

[0133] Among them, PHY init The initial value of the physical parameters of the data transmission channel, PHY max The lower limit of the window of the physical parameters of the data transmission channel, PHY min NTODT is the upper limit of the window of the physical parameters of the data transmission channel. init The initial value of the number of conversions for each data transfer, NTODT max The lower limit of the window for the number of conversions per data transmission, NTODT min The upper limit of the window for the number of conversions per data transfer.

[0134] In an optional example, the state parameter set also includes the operating temperature and power supply voltage of the SSD. Considering that the performance and data transmission rate of the SSD may be affected in different working environments, it is also necessary to adaptively adjust the action parameter value. Therefore, the method may also include the following method steps:

[0135] Step b1, obtaining the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage.

[0136] Step b2, adjusting the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission respectively according to the working temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage.

[0137] In a specific example, according to the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage, the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission are adjusted respectively, as shown in the following formula:

[0138] PHY new =PHY init +k T ·(TT ref )+k v ·(VV ref )

[0139] NTODT new =NTODT init +kT (TT ref )+k v (VV ref )

[0140] Among them, PHY new is the initial value of the physical parameters of the data transmission channel after adjustment, PHY init To adjust the initial value of the physical parameters of the data transmission channel before, k T and k v are the adjustment coefficients of PHY and NTODT for temperature and voltage, respectively. ref and V ref are the reference temperature and reference voltage values, respectively. T is the operating temperature of the SSD at the last moment, and V is the power supply voltage of the SSD at the last moment. NTODT new The initial value of the number of conversions per data transfer after adjustment, NTODT init The initial value for the number of conversions per data transfer before adjustment.

[0141] In the specific application process, the following steps are included:

[0142] (a) Environmental monitoring: real-time monitoring of the SSD operating temperature T and power supply voltage V.

[0143] (b) The intermediate values ​​of the parameter windows determined by phy and NTODT are used as the initial configuration.

[0144] (c) Dynamic adjustment: Dynamically adjust PHY parameters and NTODT settings according to environmental changes.

[0145] After adjustment, the signal transmission parameter adjustment model is added to determine the expected reward value at the previous moment. If the conditions are not met, it can be adjusted according to the adjustment trend and adjustment amount, and then determined again until the expected reward value and the target reward value meet the preset conditions. This means that the current action parameter value set meets the conditions of the first stage and no further adjustment is required. According to the current action parameter value, the data transmission channel is adjusted, and then the data reading operation is performed to determine the accuracy of the read data. If the conditions are also met, it means that the current action parameter value is the optimal action parameter value set corresponding to the first state data combination.

[0146] Specifically, considering the state parameters such as the operating temperature and power supply voltage of the SSD, the physical parameters of the data transmission channel and the initial values ​​of the number of conversions can be adjusted according to the actual working conditions, thereby achieving more accurate optimization. Among them, the changes in the operating temperature and power supply voltage may affect the performance and stability of the SSD. By adjusting according to these parameters, the stability and reliability of the SSD under different working conditions can be improved. By adjusting according to the actual operating temperature and power supply voltage, the resources of the SSD can be better utilized to avoid resource waste or performance degradation caused by unreasonable parameter settings. The method can adapt to different working environments and requirements and has high flexibility. The reference temperature and reference voltage can be set according to specific circumstances to meet different optimization goals. By adjusting the initial value, the performance of the SSD can be further optimized and the efficiency and accuracy of data transmission can be improved. The changes in the operating temperature and power supply voltage may lead to an increase in the error rate of data transmission. By adjusting the parameters, the error rate can be reduced and the reliability of the data can be improved. Reasonable parameter adjustment can reduce the loss of the SSD during operation and extend its service life. Optimizing the performance of the SSD can improve the performance of the entire system and enhance user experience.

[0147] In the present embodiment, a device for determining the action parameter value is also provided, and the device is used to implement the above-mentioned embodiment and the preferred implementation mode, and the description has been made no further. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware is also possible and conceived.

[0148] This embodiment provides a device for determining an action parameter value, such as Figure 4 As shown, it includes: an acquisition module 401, a model building module 402, a processing module 403, an adjustment module 404, a determination module 405 and an analysis module 406.

[0149] The acquisition module 401 is used to acquire, within the current iteration loop, the state parameter set of the SSD at the previous moment, the state parameter set at the current moment, the instant reward value corresponding to the previous moment, and the first action parameter value set corresponding to the state parameter set at the previous moment, wherein the state parameter set includes an operating frequency, a data transmission channel, a chip enable signal, and a target storage chip corresponding to the chip enable signal; the action parameter value set includes at least one of a physical parameter of the data transmission channel and a conversion number of each data transmission; the parameters in the state parameter set and the parameters in the action parameter value set correspond to multiple parameter values ​​respectively; when the parameters in the state parameter set and the parameters in the action parameter value set correspond to different parameter values ​​respectively, different state data combinations are constituted;

[0150] A model building module 402, for building a signal transmission parameter adjustment model according to a state parameter set of the SSD at a previous moment, an immediate reward value corresponding to the previous moment, a state parameter set at a current moment, and a first action parameter value set;

[0151] The processing module 403 is used to input the first state data combination of the SSD at the previous moment, the instant reward value corresponding to the previous moment, the second state data combination corresponding to the SSD at the current moment, and the different action parameter value sets corresponding to the second state data combination into the signal transmission parameter adjustment model to obtain the expected reward value of the SSD at the previous moment, wherein the first state data combination is any state data combination among the different state data combinations corresponding to the SSD at the previous moment, and the operation frequency in the second state data combination is the same as the operation frequency in the first state data combination;

[0152] An adjustment module 404, configured to adjust the first action parameter value set to generate a second action parameter value set when it is determined that the preset condition is not satisfied according to the expected reward value and the target reward value at the previous moment;

[0153] A determination module 405, configured to determine to read data in a target storage chip through a data transmission channel and a chip enable signal affected by a second action parameter value set;

[0154] An analysis module 406 is used to perform data consistency analysis on the read data and the original data to obtain a first analysis result, where the first analysis result is used to indicate the accuracy of the data read from the target storage chip through the data transmission channel;

[0155] The processing module 403 is also used to provide feedback of the first analysis result as a reward when the accuracy of the read data cannot meet the condition for stopping the iteration, so as to select a new set of action parameter values, until the data transmission channel is adjusted based on the final action parameter value set obtained, and the accuracy of the read data reaches a preset threshold, then the iterative operation is stopped, and the optimal set of action parameter values ​​corresponding to the first state data combination at the previous moment is obtained.

[0156] In an optional example, the processing module 403 is specifically used to: determine the maximum expected reward value corresponding to the second state data combination at the current moment according to the second state data combination corresponding to the SSD at the current moment and each action parameter value corresponding to the second state data combination;

[0157] The expected reward value of the SSD at the previous moment is determined according to the first state data combination, the immediate reward value corresponding to the previous moment, and the maximum expected reward value at the current moment corresponding to the second state data combination.

[0158] In an optional example, the model building module 402 builds a signal transmission parameter adjustment model according to the state parameter set of the SSD at the previous moment, the state parameter set at the current moment, and the first action parameter value set, which is expressed by the following formula:

[0159] Q1(s,a)←Q(s,a)+α[γmax a' Q(s',a')-Q(s,a)]

[0160] Among them, Q1(s,a) is the expected reward value corresponding to the previous moment, Q(s,a) is the immediate reward value corresponding to the previous moment, s is the first state data combination, a is the action parameter value set corresponding to the previous moment, α is the learning rate, γ is the reward decay coefficient, s' is the second state parameter combination corresponding to the current moment, Q(s',a') is the maximum expected reward value at the current moment generated at the current moment according to the different action parameter value sets corresponding to the second state parameter combination, among which the learning rate and the reward decay coefficient are both constants.

[0161] In an optional example, the adjustment module 404 is specifically configured to: when it is determined that the preset condition is not satisfied according to the expected reward value and the target reward value at the previous moment, determine the absolute difference between the expected reward value and the target reward value at the previous moment;

[0162] When the absolute difference is greater than the target difference, it is determined that the preset condition is not met;

[0163] Determine the adjustment trend based on the difference between the expected reward value at the previous moment and the target reward value;

[0164] According to the adjustment trend and the preset action parameter value adjustment amount, the first action parameter value set is adjusted to generate a second action parameter value set.

[0165] In an optional example, when the first action parameter value set includes both the physical parameter of the data transmission channel and the number of conversions of each data transmission, the adjustment module 404 is further used to adjust the second action parameter based on the default value of the first action parameter in the first action parameter value set and according to the preset action parameter value adjustment amount corresponding to the second action parameter other than the first action parameter, wherein when the first action parameter is the physical parameter of the data transmission channel, the second action parameter is the number of conversions of each data transmission; or, when the first action parameter is the number of conversions of each data transmission, the second action parameter is the physical parameter of the data transmission channel;

[0166] The determination module 405 is further used to determine to read data in the target storage chip using the data transmission channel and the chip enable signal after the second action parameter is adjusted;

[0167] The analysis module 406 is further used to perform data consistency analysis on the read data and the original data to obtain a second analysis result; when the data are determined to be consistent according to the second analysis result, determine a window lower limit of the parameter value corresponding to the second action parameter;

[0168] The adjustment module 404 is further used to continue adjusting the second action parameter according to a preset action parameter value adjustment amount corresponding to the second action parameter;

[0169] The analysis module 406 is further used to determine the upper limit of the window of the parameter value corresponding to the second action parameter when it is determined that the read data is inconsistent with the original data;

[0170] The processing module 403 is further used to determine the initial parameter value corresponding to the second action parameter according to the window upper limit and the window lower limit; and stop the operation after the initial value of the physical parameter of the data transmission channel and the initial value of the number of conversions of each data transmission are determined respectively.

[0171] In an optional example, the state parameter set further includes the operating temperature and power supply voltage of the SSD; the acquisition module 401 is further used to acquire the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage;

[0172] The adjustment module 404 is further used to adjust the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission according to the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage.

[0173] In an optional example, the adjustment module 404 adjusts the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission according to the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage, respectively, as shown in the following formula:

[0174] PHY new =PHY init +k T ·(TT ref )+k v ·(VV ref )

[0175] NTODT new =NTODT init +k T (TT ref )+k v (VV ref )

[0176] Among them, PHY newis the initial value of the physical parameters of the data transmission channel after adjustment, PHY init To adjust the initial value of the physical parameters of the data transmission channel before, k T and k v are the adjustment coefficients of PHY and NTODT for temperature and voltage, respectively. ref and V ref are the reference temperature and reference voltage values, respectively. T is the operating temperature of the SSD at the last moment, and V is the power supply voltage of the SSD at the last moment. NTODT new The initial value of the number of conversions per data transfer after adjustment, NTODT init The initial value for the number of conversions per data transfer before adjustment.

[0177] The device for determining the action parameter value in this embodiment is presented in the form of a functional module, where the module refers to an application specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0178] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0179] An embodiment of the present invention provides a device for determining an action parameter value, which obtains a state parameter set of an SSD at the last moment, a state parameter set at the current moment, an immediate reward value corresponding to the last moment, and a first action parameter value set corresponding to the state parameter set at the last moment, and constructs a signal transmission parameter adjustment model based on these parameters. By obtaining the state parameter set, immediate reward value, and action parameter value set of the SSD at the last moment and the current moment, the working state and performance of the SSD can be understood. Then, the expected reward value at the last moment is obtained using the model, and based on the expected reward value and the immediate reward value, it is determined whether the first action parameter value model needs to be adjusted. Thereby optimizing the data transmission channel. Finally, by continuously iteratively adjusting the action parameter value set and performing data consistency analysis, the optimal action parameter value set under a specific state can be found, which can improve the transmission efficiency and performance of the SSD and speed up the data reading speed. Performing consistency analysis on the read data and adjusting the action parameters according to the feedback of the analysis results helps to improve the accuracy of data reading and reduce data errors. Adaptive adjustment is performed according to the real-time state of the SSD to adapt to different working conditions and requirements, thereby improving the flexibility and adaptability of the system.

[0180] The embodiment of the present invention also provides a computer device having the above Figure 4 The action parameter values ​​shown are determined by the device.

[0181] See also Figure 5 , Figure 5 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.

[0182] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include an integrated circuit. The integrated circuit may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0183] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0184] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0185] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0186] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 5 The example of connecting through bus is taken in the following.

[0187] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.

[0188] The embodiment of the present invention also provides a computer-readable storage medium. The method provided in the above embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or is implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0189] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.

[0190] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A method for determining an action parameter value, characterized in that: The method comprises: In the current iteration loop, a state parameter set of the SSD at the previous moment, a state parameter set at the current moment, an immediate reward value corresponding to the previous moment, and a first action parameter value set corresponding to the state parameter set at the previous moment are obtained, wherein the state parameter set includes an operating frequency, a data transmission channel, a chip enable signal, and a target storage chip corresponding to the chip enable signal, and the action parameter value set includes at least one of a physical parameter of the data transmission channel and a conversion number of each data transmission, and the parameters in the state parameter set and the parameters in the action parameter value set correspond to multiple parameter values ​​respectively, and when the parameters in the state parameter set and the parameters in the action parameter value set correspond to different parameter values ​​respectively, different state data combinations are constituted; Constructing a signal transmission parameter adjustment model according to the state parameter set of the SSD at the previous moment, the instant reward value corresponding to the previous moment, the state parameter set at the current moment, and the first action parameter value set; Inputting a first state data combination of the SSD at the previous moment, an immediate reward value corresponding to the previous moment, a second state data combination corresponding to the SSD at the current moment, and a set of different action parameter values ​​corresponding to the second state data combination into the signal transmission parameter adjustment model to obtain an expected reward value of the SSD at the previous moment, wherein the first state data combination is any state data combination among different state data combinations corresponding to the SSD at the previous moment, and an operating frequency in the second state data combination is the same as an operating frequency in the first state data combination; When it is determined that the preset condition is not satisfied according to the expected reward value and the target reward value at the previous moment, adjusting the first action parameter value set to generate a second action parameter value set; Determining to read data in a target storage chip through a data transmission channel and a chip enable signal affected by the second action parameter value set; Performing data consistency analysis on the read data and the original data to obtain a first analysis result, where the first analysis result is used to indicate the accuracy of the data read from the target storage chip through the data transmission channel; When the accuracy of the read data cannot meet the condition for stopping iteration, the first analysis result is fed back as a reward to select a new set of action parameter values. After the data transmission channel is adjusted based on the final action parameter value set obtained, the accuracy of the read data reaches a preset threshold, and the iterative operation is stopped to obtain the optimal set of action parameter values ​​corresponding to the first state data combination at the previous moment.

2. The method according to claim 1, characterized in that The step of inputting the first state data combination of the SSD at the previous moment, the immediate reward value corresponding to the previous moment, the second state data combination corresponding to the SSD at the current moment, and different action parameter value sets corresponding to the second state data combination into the signal transmission parameter adjustment model to obtain the expected reward value of the SSD at the previous moment specifically includes: Determine the maximum expected reward value corresponding to the second state data combination at the current moment according to the second state data combination corresponding to the SSD at the current moment and each action parameter value corresponding to the second state data combination; The expected reward value of the SSD at the previous moment is determined according to the first state data combination, the immediate reward value corresponding to the previous moment, and the maximum expected reward value at the current moment corresponding to the second state data combination.

3. The method according to claim 1 or 2, characterized in that: The signal transmission parameter adjustment model is constructed according to the state parameter set of the SSD at the previous moment, the state parameter set at the current moment, and the first action parameter value set, and is expressed by the following formula: Q1(s,a)←Q(s,a)+α[γmax a' Q(s',a')-Q(s,a)] Among them, Q1(s,a) is the expected reward value corresponding to the previous moment, Q(s,a) is the immediate reward value corresponding to the previous moment, s is the first state data combination, a is the action parameter value set corresponding to the previous moment, α is the learning rate, γ is the reward decay coefficient, s' is the second state parameter combination corresponding to the current moment, Q(s',a') is the maximum expected reward value at the current moment generated at the current moment according to the different action parameter value sets corresponding to the second state parameter combination, wherein the learning rate and the reward decay coefficient are both constants.

4. The method according to claim 1 or 2, characterized in that: When it is determined according to the expected reward value and the target reward value at the previous moment that the preset condition is not satisfied, adjusting the first action parameter value set to generate a second action parameter value set specifically includes: Determine the absolute difference between the expected reward value at the previous moment and the target reward value; When the absolute difference is greater than the target difference, determining that the preset condition is not met; Determining an adjustment trend according to the difference between the expected reward value at the previous moment and the target reward value; According to the adjustment trend and the preset action parameter value adjustment amount, the first action parameter value set is adjusted to generate a second action parameter value set.

5. The method according to claim 1 or 2, characterized in that: When the first action parameter value set includes both the physical parameters of the data transmission channel and the number of conversions of each data transmission, the method further includes: Based on the default value of the first action parameter in the first action parameter value set, the second action parameter is adjusted according to the preset action parameter value adjustment amount corresponding to the second action parameter other than the first action parameter, wherein when the first action parameter is the physical parameter of the data transmission channel, the second action parameter is the number of conversions for each data transmission; or when the first action parameter is the number of conversions for each data transmission, the second action parameter is the physical parameter of the data transmission channel; Using the data transmission channel and the chip enable signal after the second action parameter is adjusted, determining to read data in the target storage chip; Performing data consistency analysis on the read data and the original data to obtain a second analysis result; When it is determined according to the second analysis result that the data are consistent, determining a window lower limit of a parameter value corresponding to the second action parameter; Continue to adjust the second action parameter according to the preset action parameter value adjustment amount corresponding to the second action parameter until it is determined that the read data is inconsistent with the original data, and determine the window upper limit of the parameter value corresponding to the second action parameter; Determining an initial parameter value corresponding to the second action parameter according to the window upper limit and the window lower limit; The operation is stopped after the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission are determined respectively.

6. The method according to claim 1 or 2, characterized in that: The state parameter set also includes an operating temperature and a power supply voltage of the SSD; and the method further includes: Obtaining the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage; According to the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage, the initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission are adjusted respectively.

7. The method according to claim 6, characterized in that The initial values ​​of the physical parameters of the data transmission channel and the initial values ​​of the number of conversions of each data transmission are adjusted according to the operating temperature of the SSD at the last moment, the power supply voltage at the last moment, the reference temperature and the reference voltage, respectively. For details, see the following formula: PHY new =PHY init +k T ·(T-T ref )+k v ·(V-V ref ) NTODT new =NTODT init +k T (T-T ref )+k v (V-V ref ) Among them, PHY new is the initial value of the physical parameters of the data transmission channel after adjustment, PHY init To adjust the initial value of the physical parameters of the data transmission channel before, k T and k v are the adjustment coefficients of PHY and NTODT for temperature and voltage, respectively. ref and V ref are the reference temperature and reference voltage values, respectively. T is the operating temperature of the SSD at the last moment, and V is the power supply voltage of the SSD at the last moment. NTODT new The initial value of the number of conversions per data transfer after adjustment, NTODT init The initial value for the number of conversions per data transfer before adjustment.

8. A device for determining an action parameter value, characterized in that: The device comprises: an acquisition module, used to acquire, within a current iteration loop, a state parameter set of the SSD at a previous moment, a state parameter set at a current moment, an immediate reward value corresponding to the previous moment, and a first action parameter value set corresponding to the state parameter set at the previous moment, wherein the state parameter set includes an operating frequency, a data transmission channel, a chip enable signal, and a target storage chip corresponding to the chip enable signal; the action parameter value set includes at least one of a physical parameter of the data transmission channel and a conversion number of each data transmission; the parameters in the state parameter set and the parameters in the action parameter value set correspond to multiple parameter values ​​respectively; when the parameters in the state parameter set and the parameters in the action parameter value set correspond to different parameter values ​​respectively, different state data combinations are constituted; A model building module, used to build a signal transmission parameter adjustment model according to a state parameter set of the SSD at a previous moment, an immediate reward value corresponding to the previous moment, a state parameter set at a current moment, and the first action parameter value set; a processing module, configured to input a first state data combination of the SSD at a previous moment, an immediate reward value corresponding to the previous moment, a second state data combination corresponding to the SSD at a current moment, and a set of different action parameter values ​​corresponding to the second state data combination into the signal transmission parameter adjustment model, to obtain an expected reward value of the SSD at the previous moment, wherein the first state data combination is any state data combination among different state data combinations corresponding to the SSD at the previous moment, and an operation frequency in the second state data combination is the same as an operation frequency in the first state data combination; An adjustment module, configured to adjust the first action parameter value set to generate a second action parameter value set when it is determined that a preset condition is not satisfied according to the expected reward value and the target reward value at the previous moment; A determination module, used to determine to read data in a target storage chip through a data transmission channel and a chip enable signal affected by the second action parameter value set; An analysis module, used for performing data consistency analysis on the read data and the original data to obtain a first analysis result, where the first analysis result is used to indicate the accuracy of the data read from the target storage chip through the data transmission channel; The processing module is also used to, when the accuracy of the read data cannot meet the condition for stopping iteration, use the first analysis result as a reward for feedback, so as to select a new set of action parameter values, until the data transmission channel is adjusted based on the final action parameter value set, and the accuracy of the read data reaches a preset threshold, then stop the iterative operation, and obtain the optimal set of action parameter values ​​corresponding to the first state data combination at the previous moment.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for determining the action parameter value according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for determining the action parameter value according to any one of claims 1 to 7.