Troubleshooting methods and related equipment for memory
By recording memory fault information in electronic devices and starting a backup operating system in case of serious problems, the problem of bad blocks and failures in storage devices during use is solved, user data loss is avoided, and the reliability and energy efficiency of the device are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2023-12-05
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies lack effective solutions to address bad blocks and failures in storage devices during the use of electronic devices, resulting in a high risk of user data loss.
When an electronic device detects a memory anomaly, it records fault information and boots a backup operating system in case of a serious problem, avoiding entering emergency download mode. By scanning critical and frequently used partitions, it identifies faults and migrates critical image files, ensuring a switch to a backup operating system in the event of severe memory damage.
It effectively avoids emergency download mode caused by serious memory problems, protects user data, reduces the number of scans and power consumption, and improves the reliability of electronic devices.
Smart Images

Figure CN120144374B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a fault handling method and related equipment for memory. Background Technology
[0002] Storage devices (or simply memory) serve as the carriers of all files in electronic devices, and their operational status plays a decisive role in the user experience. However, over a long period of use, storage devices gradually develop bad blocks or even fail completely due to varying usage environments (extreme high and low temperatures, drops). Once these problems occur, they will severely impact the user experience. Therefore, corresponding solutions need to be developed to avoid and prevent storage device failures (testing before shipment). For example, testing for bad blocks and failures in storage devices should be conducted before the electronic device leaves the factory. However, there is still no suitable solution for storage device problems that arise during user use. Summary of the Invention
[0003] This application provides a method and related apparatus for handling memory faults. According to this method, an electronic device can record relevant fault information and perform preliminary identification and judgment even when no serious problem is detected in its memory. Upon detection of a serious problem, it triggers entry into a backup operating system during subsequent boot processes and identifies the faulty partition based on the previously recorded fault information. Through this method, the electronic device can continuously detect and record relevant fault information even before a serious problem exists in the memory, enabling timely identification of serious memory problems and booting of a backup operating system, thereby avoiding entry into emergency download mode and preventing user data loss due to entering emergency download mode.
[0004] Firstly, this application provides a fault handling method for memory. This method can be applied to an electronic device equipped with memory. The memory partitions include critical partitions and frequently used partitions. The method may include: the electronic device can detect an operation that triggers its startup; in response to the startup operation, the electronic device can begin startup; if, in a first state, an anomaly detected by the electronic device satisfies the requirement to enter an emergency download mode, and in a second state prior to startup, the electronic device scans and confirms that a critical partition has an anomaly and the migration of a critical image file affected by the anomaly in the critical partition has failed, or scans and confirms that a frequently used partition has an anomaly and meets a first preset condition, then the electronic device can start a first operating system. Here, the first state refers to the state of the electronic device during the boot process, and the second state refers to the state of the kernel of the second operating system in the electronic device after startup.
[0005] In the solution provided in this application, after the electronic device boots normally (i.e., the kernel of the commonly used operating system in the electronic device boots), the electronic device can scan critical partitions. If the scan confirms that a critical partition is abnormal and the migration of the affected critical image file in the critical partition fails, and if the electronic device detects an anomaly upon the next boot and triggers an emergency download mode, the electronic device can boot a backup operating system instead of the commonly used operating system. Similarly, after the electronic device boots normally, it can also scan commonly used partitions. If the scan confirms that a commonly used partition is abnormal and meets a first preset condition, and if the electronic device detects an anomaly upon the next boot and triggers an emergency download mode, the electronic device can boot a backup operating system instead of the commonly used operating system. Through this method, the electronic device can detect the partitions corresponding to the memory after booting and record the corresponding fault information. Furthermore, once a serious problem is detected in the partition corresponding to the memory, the electronic device can determine that the memory is severely damaged. Thus, the electronic device can trigger the boot of the backup operating system upon the next boot to avoid the inability to recover after entering the emergency download mode, which would render the electronic device unusable and prevent user data loss.
[0006] In some embodiments of this application, the first operating system may refer to a backup operating system of the electronic device. The second operating system may be the commonly used operating system of the electronic device, that is, the operating system generally used by the electronic device. For example, the second operating system may be an advanced operating system (as shown in step S101).
[0007] It should be noted that the first operating system is a separate operating system designed to handle memory failures, and is different from the operating system originally used by the electronic device (i.e., the second operating system). Under normal circumstances (e.g., except in cases of severe memory damage), the electronic device uses the second operating system and does not use the first operating system.
[0008] In some embodiments of this application, the memory refers to UFS. Of course, the memory involved in this application may also be other types of memory, and this application does not limit it.
[0009] In some embodiments of this application, the first state refers to the boot phase, which can be understood as the stage in which the electronic device runs the bootloader. The second state can be the standby state mentioned later.
[0010] In some embodiments of this application, the first preset condition may be preset condition 1 mentioned below.
[0011] It should be noted that the abnormality detected by the electronic device in the first state satisfies the requirement to enter the emergency download mode. Specifically, this may include: the electronic device detecting an abnormality during startup and triggering the entry into the emergency download mode (as shown in step S112).
[0012] In conjunction with the first aspect, in one possible implementation, before the electronic device detects an operation that triggers its startup, the method may further include: in a second state, the electronic device can scan the critical partition, generate first fault information, and determine whether the critical partition is abnormal based on the first fault information; if the critical partition is abnormal, the electronic device can migrate the affected critical image file in the critical partition; if the migration of the critical image file fails, the electronic device can set the content corresponding to the first parameter to the first content; if the critical partition is not abnormal, the electronic device can scan the commonly used partition, generate second fault information, and determine whether the commonly used partition is abnormal based on the second fault information; if the commonly used partition is abnormal, and the abnormality of the commonly used partition meets a first preset condition, the electronic device can set the content corresponding to the first parameter to the first content. If, in the second state before startup, the electronic device scans and confirms that the critical partition is abnormal and the migration of the affected critical image file in the critical partition fails, or scans and confirms that the commonly used partition is abnormal and meets the first preset condition, then the electronic device starts the first operating system, specifically including: if the content corresponding to the first parameter is the first content, then the electronic device can start the first operating system.
[0013] In the solution provided in this application, after the electronic device boots normally (i.e., the kernel of the commonly used operating system in the electronic device boots), the electronic device can first scan the critical partition, and then scan the commonly used partition if no critical partition is found to be abnormal. If a critical partition is found to be abnormal, the electronic device can migrate the critical image file, and if the migration fails, it marks the memory as severely damaged by setting the content corresponding to the first parameter to the first content. If a commonly used partition is found to be abnormal and the abnormality meets the first preset condition, the electronic device can also mark the memory as severely damaged by setting the content corresponding to the first parameter to the first content. In this way, during subsequent boot processes, the electronic device can directly determine whether to boot a backup operating system based on the content corresponding to the first parameter. This method not only allows the electronic device to promptly confirm the extent of memory damage, ensuring that no severely damaged memory is overlooked, but also reduces the number of scans and simplifies the process to some extent, thereby saving energy consumption.
[0014] In some embodiments of this application, the first fault information can be fault information 2, and the second fault information can be fault information 3.
[0015] In some embodiments of this application, in the event of a failure to migrate a critical image file, or when an anomaly is detected in a frequently used partition and the anomaly meets a first preset condition, the electronic device may also record first information. The first information can be used to mark a serious problem (i.e., severe damage) in the memory of the electronic device. Thus, if an anomaly is detected during the subsequent startup process of the electronic device and an emergency download mode is triggered, the electronic device can determine whether to start a backup operating system based on whether the first information is recorded. Specifically, if the electronic device records the first information, the electronic device can start the operating system. It is understood that this application does not limit the specific content and form of the first information.
[0016] In conjunction with the first aspect, in one possible implementation, before the electronic device detects an operation that triggers its startup, the method may further include: in a second state, the electronic device can scan a critical partition and a frequently used partition, generate first fault information and second fault information respectively, and determine whether the critical partition and the frequently used partition are abnormal based on the first fault information and the second fault information respectively; if the critical partition is abnormal, the electronic device can migrate the affected critical image file in the critical partition; if the frequently used partition is abnormal, and the abnormality of the frequently used partition meets a first preset condition, or if the migration of the affected critical image file in the critical partition fails, the electronic device can set the content corresponding to the first parameter to the first content. If, in the second state before startup, the electronic device scans and confirms that the critical partition is abnormal and the migration of the affected critical image file in the critical partition fails, or scans and confirms that the frequently used partition is abnormal and meets the first preset condition, then the electronic device starts the first operating system, specifically including: if the content corresponding to the first parameter is the first content, then the electronic device can start the first operating system.
[0017] In the solution provided in this application, after the electronic device boots normally (i.e., the kernel of the commonly used operating system in the electronic device boots), the electronic device can scan critical partitions and commonly used partitions. If an anomaly is detected in a critical partition, the electronic device can migrate the critical image file, and if the migration fails, it marks the memory as severely damaged by setting the content corresponding to the first parameter to the first content. Similarly, if an anomaly is detected in a commonly used partition and the anomaly meets a first preset condition, the electronic device can also mark the memory as severely damaged by setting the content corresponding to the first parameter to the first content. In this way, during subsequent boot processes, the electronic device can directly determine whether a backup operating system needs to be booted based on the content corresponding to the first parameter. This method allows the electronic device to perform a more comprehensive scan of the partitions corresponding to the memory, helping it to confirm the memory's fault condition and obtain more comprehensive fault information.
[0018] It is understood that, under the above circumstances, this application does not restrict the order in which the electronic device scans critical partitions and frequently used partitions. In some embodiments of this application, the electronic device may scan the critical partitions first and then scan the frequently used partitions, but the electronic device does not determine whether to scan the frequently used partitions based on the scanning results of the critical partitions.
[0019] In conjunction with the first aspect, in one possible implementation, the electronic device determines that a critical partition has malfunctioned based on the first fault information, which may specifically include any one or more of the following: the electronic device determines that the number of failed addresses included in the first fault information is not less than a first threshold; the electronic device determines that the number of logical region access failures included in the first fault information is not less than a second threshold; the electronic device determines that the number of single access IO failures included in the first fault information is not less than a third threshold.
[0020] In the solution provided in this application, memory partition anomalies can include memory damage that does not conform to the typical wear and tear patterns, but exclude natural damage to the memory caused by user usage or other factors. The electronic device can determine whether a critical partition is experiencing anomalies based specifically on the number of failed addresses and / or the number of logical region access failures and / or the number of I / O failures included in the fault information. Through this method, the electronic device can promptly record anomalies that do not conform to the typical wear and tear patterns (i.e., the degree of memory damage does not match its service life and usage).
[0021] In some embodiments of this application, the first threshold may be threshold 1 mentioned below, the second threshold may be threshold 2 mentioned below, and the third threshold may be threshold 3 mentioned below.
[0022] In conjunction with the first aspect, in one possible implementation, the electronic device determines that a critical partition has malfunctioned based on the first fault information. Specifically, this may further include: the electronic device can determine the set of failed addresses included in the first fault information.
[0023] In the solution provided in this application, the critical partition anomaly corresponding to the memory may include a set of failed addresses. In this case, the electronic device will only migrate the critical image file if the first fault information meets certain requirements (specific conditions involving the first threshold, second threshold, or third threshold mentioned above). It is understood that compared to the case of scattered failed addresses, it is easier for the electronic device to migrate the critical image file when the failed addresses are set.
[0024] In conjunction with the first aspect, in one possible implementation, the set of failed addresses included in the first fault information may specifically include any one or more of the following: the ratio of the number of failed addresses belonging to the same sector included in the first fault information to the total number of failed addresses included in the first fault information is not less than a first ratio; the number of first-type failed addresses included in the first fault information is greater than a first value; the ratio of the number of first-type failed addresses included in the first fault information to the total number of failed addresses included in the first fault information is greater than a second ratio. Wherein, the difference between the number corresponding to the first-type failed address and the number corresponding to other failed addresses included in the first fault information is less than a second value.
[0025] In some embodiments of this application, the first ratio may be a%, the first value may be c, the second value may be b, and the second ratio may be d.
[0026] In some embodiments of this application, the difference between the number corresponding to the first type of failure address and the number corresponding to all other failure addresses included in the first fault information is less than the second value.
[0027] In some embodiments of this application, the difference between the number corresponding to the first type of failure address and the number corresponding to a subset of other failure addresses included in the first fault information (e.g., e% of all other failure addresses included in the first fault information) is less than a second value. It is understood that e can be set according to actual needs, and this application does not limit its specific value.
[0028] In conjunction with the first aspect, in one possible implementation, the electronic device determines that a commonly used partition has malfunctioned based on the second fault information. Specifically, this may include any one or more of the following: the electronic device determines that the number of failed addresses included in the second fault information is not less than a fourth threshold; the electronic device determines that the number of logical region access failures included in the second fault information is not less than a fifth threshold; the electronic device determines that the number of I / O failures included in the second fault information is not less than a sixth threshold.
[0029] In the solution provided in this application, memory partition anomalies can include memory damage that does not conform to the typical wear and tear patterns, but exclude natural damage to the memory caused by user usage or other factors. The electronic device can determine whether frequently used partitions are experiencing anomalies based specifically on the number of failed addresses and / or the number of logical region access failures and / or the number of I / O failures included in the fault information. Through this method, the electronic device can promptly record anomalies that do not conform to the typical wear and tear patterns (i.e., the degree of memory damage does not match its service life and usage).
[0030] In some embodiments of this application, the fourth threshold may be the threshold 4 mentioned below, the fifth threshold may be the threshold 5 mentioned below, and the sixth threshold may be the threshold 6 mentioned below.
[0031] In conjunction with the first aspect, in one possible implementation, the anomalies occurring in frequently used partitions satisfy a first preset condition, which may specifically include any one or more of the following: the number of failed addresses included in the second fault information is not less than a seventh threshold; the number of logical region access failures included in the second fault information is not less than an eighth threshold; and the number of IO failures included in the second fault information is not less than a ninth threshold. Wherein, the seventh threshold is greater than the fourth threshold, the eighth threshold is greater than the fifth threshold, and the ninth threshold is greater than the sixth threshold.
[0032] In the solution provided in this application, the electronic device can set different thresholds to determine the degree of anomaly in frequently used partitions. The electronic device can set a relatively smaller threshold to determine if an anomaly has occurred in a frequently used partition, while setting a relatively larger threshold to determine if the anomaly in a frequently used partition is more severe (i.e., meets the first preset condition). Through this method, the electronic device can not only record minor memory anomalies, but also promptly handle severe memory anomalies.
[0033] In some embodiments of this application, the seventh threshold may be the threshold 7 mentioned below, the eighth threshold may be the threshold 8 mentioned below, and the ninth threshold may be the threshold 9 mentioned below.
[0034] In conjunction with the first aspect, in one possible implementation, when the electronic device determines that a common partition is abnormal based on the second fault information, the method may further include: the electronic device may determine the fault level of the memory based on the second fault information; when the electronic device determines that the fault level of the memory meets the second preset condition, the electronic device may determine that the abnormality of the common partition meets the first preset condition.
[0035] In the solution provided in this application, the electronic device can be configured with different fault levels to determine the degree of anomaly in frequently used partitions. Through this method, the electronic device can not only record minor memory anomalies but also promptly handle serious abnormal memory anomalies.
[0036] In some embodiments of this application, the fault level of the memory meets a second preset condition, which may specifically include: the fault level of the memory (e.g., the UFS fault mentioned below) is among the top b% of fault levels.
[0037] In some embodiments of this application, the fault level of the memory meets a second preset condition, which may specifically include: the fault level corresponding to the memory is the most severe fault level.
[0038] In some embodiments of this application, the fault level of the memory meets a second preset condition, which may specifically include: the number of fault levels is greater than 2, and the fault level corresponding to the memory is one of the two most severe fault levels.
[0039] Of course, the fault level of the memory meeting the second preset condition may include other specific contents, and this application does not limit this.
[0040] In conjunction with the first aspect, in one possible implementation, the electronic device determines the fault level of the memory based on the second fault information. Specifically, this may include: the electronic device determining the fault level of the memory based on the number of failed addresses included in the second fault information; the electronic device determining the fault level of the memory based on the number of logical region access failures included in the second fault information; the electronic device determining the fault level of the memory based on the number of I / O failures included in the second fault information; the electronic device determining the fault level of the memory based on the concentration of failed addresses included in the second fault information; the electronic device determining the fault level of the memory based on the number of interaction failures between the memory and the processor in the electronic device involved in the second fault information; and the electronic device determining the fault level of the memory based on the ratio of the number of interaction failures between the memory and the processor involved in the second fault information to the total number of interactions between the memory and the processor.
[0041] In the solution provided in this application, the electronic device can specifically determine the fault level based on the failure address and / or logical area access and / or IO access involved in the second fault information.
[0042] In some embodiments of this application, the processor in the electronic device may be a System-on-a-Chip (SoC).
[0043] In conjunction with the first aspect, in one possible implementation, the electronic device scans frequently used partitions and generates second fault information. Specifically, this may include: the electronic device scanning frequently used partitions at a preset frequency, and after each scan of a frequently used partition, the electronic device generating corresponding second fault information.
[0044] In some embodiments of this application, the electronic device scans the frequently used partition not only once, but at a preset frequency, and generates corresponding fault information (i.e., second fault information) after each scan. Through this method, the electronic device can continuously detect anomalies in the frequently used partition and record the corresponding fault information in a timely manner, so that serious anomalies can be detected promptly.
[0045] In some embodiments of this application, the preset frequency can be a certain frequency mentioned below. It is understood that the preset frequency can be set according to actual needs, and this application does not impose specific limitations on it.
[0046] In conjunction with the first aspect, in one possible implementation, after the electronic device scans frequently used partitions, the method may further include: if an anomaly occurs in a frequently used partition, and the anomaly does not meet a first preset condition, the electronic device may save the second fault information obtained in this scan, and if an anomaly is determined to occur in the frequently used partition in the next scan, the electronic device may combine the second fault information obtained from historical scans of frequently used partitions with the second fault information obtained after the next scan of frequently used partitions to determine whether the anomaly in the frequently used partition meets the first preset condition. The second fault information obtained from historical scans of frequently used partitions includes the second fault information obtained in this scan.
[0047] In the solution provided in this application, if the electronic device determines that a frequently used partition is abnormal but the abnormality does not meet a first preset condition, the electronic device can save the corresponding fault information. During subsequent scanning, the saved fault information can be used to comprehensively determine whether the abnormality of the frequently used partition meets the first preset condition. Through this method, the electronic device can continuously detect abnormalities in frequently used partitions and promptly record the corresponding fault information, so that serious abnormalities can be detected in a timely manner.
[0048] In conjunction with the first aspect, in one possible implementation, the memory partition also includes a partition related to the startup of the electronic device. Before the electronic device detects an operation that triggers its startup, the method may further include: if the electronic device detected an anomaly during the boot phase of the previous startup process, the electronic device can generate third fault information and store the third fault information in the partition related to the startup of the electronic device.
[0049] In the solution provided in this application, the electronic device can also perform detection during the boot phase to generate fault information corresponding to the boot-related partitions. This allows the electronic device to perform more comprehensive detection on the partitions corresponding to the memory, thereby obtaining more complete fault information and facilitating subsequent identification of faulty partitions.
[0050] In some embodiments of this application, the third fault information may be fault information 1 mentioned below.
[0051] In conjunction with the first aspect, in one possible implementation, after the electronic device boots the first operating system, the method may further include: the electronic device can capture third fault information to determine the faulty partition of the memory, and perform a recovery operation on the faulty partition of the memory.
[0052] In the solution provided in this application, the electronic device can determine the faulty partition of the memory based on third fault information, and then perform a recovery operation on the faulty partition. Through this method, the electronic device can handle memory faults to a certain extent, thereby recovering some user data and avoiding the unusability of the electronic device and the loss of user data caused by entering emergency download mode.
[0053] In conjunction with the first aspect, in one possible implementation, after the electronic device starts the first operating system, the method may further include: the electronic device can capture first fault information and second fault information to determine the faulty partition of the memory, and perform a recovery operation on the faulty partition of the memory.
[0054] In the solution provided in this application, the electronic device can determine the faulty partition of the memory based on first fault information and second fault information, and then perform a recovery operation on the faulty partition. Through this method, the electronic device can handle memory faults to a certain extent, thereby recovering some user data and avoiding the unusability of the electronic device and the loss of user data caused by entering emergency download mode.
[0055] In conjunction with the first aspect, in one possible implementation, the first operating system could be a recovery partition that shields non-critical disk read / write operations, or a fastboot partition that shields non-critical disk read / write operations.
[0056] In the solution provided in this application, the electronic device can obtain a backup operating system based on either a recovery partition or a fastboot partition. The electronic device can shield the recovery partition from non-critical disk read / write operations to obtain a backup operating system. Similarly, the electronic device can shield the fastboot partition from non-critical disk read / write operations to obtain a backup operating system.
[0057] It should be noted that setting a backup operating system for an electronic device will not affect the original recovery partition and fastboot partition of the electronic device, that is, the recovery partition and fastboot partition in the commonly used operating system.
[0058] In conjunction with the first aspect, in one possible implementation, the anomaly detected by the electronic device in the first state satisfies the requirement to enter the emergency download mode, which may specifically include any one or more of the following: failure of interaction between the memory and processor in the electronic device; inability of the electronic device to read the image file in the memory; failure of the electronic device to load.
[0059] In the solution provided in this application, many reasons can cause an electronic device to enter emergency download mode, not limited to memory damage. Therefore, when an electronic device detects an anomaly, it needs to determine whether the entry into emergency download mode is caused by a reason related to memory damage. If the entry into emergency download mode is caused by a reason related to memory damage, the electronic device can further determine whether to start a backup operating system. If the entry into emergency download mode is not caused by a reason related to memory damage, the electronic device does not need to further determine whether to start a backup operating system.
[0060] In a second aspect, this application provides an electronic device comprising: one or more memories and one or more processors; the one or more memories being coupled to the one or more processors, the memories being used to store computer program code including computer instructions, the one or more processors calling the computer instructions to cause the electronic device to perform the method as described in the first aspect or any implementation thereof.
[0061] Thirdly, this application provides a computer storage medium. The computer storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect or any implementation thereof.
[0062] Fourthly, embodiments of this application provide a chip. This chip can be applied to an electronic device, and the chip includes one or more processors. The processors are used to invoke computer instructions to cause the electronic device to perform the methods described in the first aspect or any implementation thereof.
[0063] Fifthly, embodiments of this application provide a computer program product containing instructions. When the computer program product is run on an electronic device, it causes the electronic device to perform the method described in the first aspect or any implementation thereof.
[0064] It is understood that the electronic device provided in the second aspect, the computer storage medium provided in the third aspect, the chip provided in the fourth aspect, and the computer program product provided in the fifth aspect are all used to execute the method described in the first aspect or any implementation thereof. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects of any possible implementation of the first aspect, which will not be repeated here. Attached Figure Description
[0065] Figure 1 A flowchart illustrating the process of entering emergency download mode in the event of a memory failure, as provided in this application embodiment;
[0066] Figure 2AA flowchart illustrating a fault handling method for a memory provided in an embodiment of this application;
[0067] Figure 2B A flowchart illustrating yet another fault handling method for a memory provided in this application embodiment;
[0068] Figure 3 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;
[0069] Figure 4 This is a schematic diagram of the software structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0070] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0071] It should be understood that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0072] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.
[0073] Memory is a storage device used in modern information technology to store information. All information in an electronic device, including raw input data, programs, intermediate results, and final results, is stored in memory. It stores and retrieves information according to the location specified by the controller. Memory is essential for electronic devices to have a memory function and ensure normal operation. As mentioned above, memory may gradually develop bad blocks over time and with different usage conditions, which can seriously affect the use of electronic devices.
[0074] It is understood that the memory involved in this application may include Universal Flash Storage (UFS), and may also include other types of memory, which are not limited in this application. Among them, UFS is an internal file storage, the specific meaning of which can be found in relevant technical documents, and will not be elaborated here.
[0075] In some embodiments of this application, when an electronic device detects a memory abnormality, it can enter an emergency download mode and cannot actively exit. "Cannot actively exit" means that after entering emergency download mode, the electronic device only supports a limited number of command-line instructions enumerated once; even if the electronic device is powered off and restarted, it will still enter emergency download mode again. Once in emergency download mode, the electronic device will remain black and will not respond to user button presses.
[0076] Understandably, emergency download mode is a special state that is entered during the startup phase due to device malfunction or after the corresponding flag is manually enabled during the debugging phase.
[0077] Specifically, such as Figure 1 As shown, if an electronic device suffers physical damage, malfunctions, or shuts down due to lack of power, the user can trigger a restart. If a memory anomaly is detected during the restart process, the device can enter emergency download mode and remain in this mode. If the device is powered off and restarted, it will again detect a memory anomaly and re-enter emergency download mode.
[0078] After an electronic device enters emergency download mode, testers can capture dump logs using single-enumeration command-line commands to obtain partial information about the memory anomaly. It's understood that dump logs generally refer to the operating system kernel logs. When the operating system experiences an abnormal lockout / blue screen / crash, the system kernel process generates dump logs for diagnosing and analyzing system, application, and hardware failures.
[0079] It is worth noting that once an electronic device enters emergency download mode, it remains black, has no Android Debug Bridge (ADB) port, and does not respond to user button operations. Therefore, once an electronic device enters emergency download mode, the user can only replace the electronic device, and the user data originally stored on the electronic device cannot be retrieved.
[0080] Based on the above, this application provides a method and related device for handling memory faults. According to this method, an electronic device can continuously detect whether the memory is malfunctioning, and if so, determine the type of the malfunction. For memory malfunctions related to critical partitions, the electronic device can migrate the critical image file and set a flag if the migration fails. However, for other memory malfunctions related to frequently used partitions that meet certain conditions (e.g., preset condition 1), the electronic device can set a flag. After the electronic device restarts, if it determines that it is about to enter an emergency download state and the flag has been set, the electronic device can boot a backup operating system and analyze the relevant fault information of this memory malfunction. Through this method, users can continue to use the electronic device even if the memory malfunctions, avoiding data loss if it directly enters emergency download mode.
[0081] The following uses UFS as an example to introduce a fault handling method for memory provided in the embodiments of this application.
[0082] Please see Figure 2A , Figure 2A A flowchart illustrating a fault handling method for a memory provided in this application embodiment. The method may include, but is not limited to, the following steps:
[0083] S101: The electronic device receives a command to start the electronic device, and in response to the command, the electronic device starts up.
[0084] Users can trigger the electronic device to start by pressing and holding the corresponding button (e.g., pressing and holding the power button, volume down button, and power button). Correspondingly, the electronic device receives the command to start and, in response, begins the startup process.
[0085] During the boot process of an electronic device, the device's kernel starts. Specifically, in response to the command to boot the electronic device, the device is powered on and enters the boot phase to prepare for loading the kernel. Once the device is ready to load the kernel, it can start the kernel. Powering on means energizing the electronic device.
[0086] In some embodiments of this application, the electronic device enters the boot phase, which may specifically include the electronic device running a bootloader. It should be noted that the boot phase of the electronic device may include different specific processes depending on the chip platform used.
[0087] As we can understand, the bootloader is the program that boots the operating system (e.g., Android) before it starts. Its main functions include checking random access memory (RAM) and initializing parameters. Electronic devices can use the bootloader to initialize and load their hardware (e.g., memory, central processing unit, etc.) into RAM, thereby establishing a memory space mapping and preparing for the loading of the kernel.
[0088] It should be noted that during the execution of step S101, the kernel booted during the electronic device's startup process is the kernel of a High Level Operating System (HLOS). This can be understood as including operating systems such as Linux and Android. For example, if the electronic device's operating system is Android (which can be simply referred to as the Android system), the electronic device can boot the Android operating system kernel.
[0089] S102: During the startup phase, the electronic device detects whether any abnormalities have occurred.
[0090] During the startup phase of an electronic device, the device can detect any abnormal conditions, i.e., determine whether the device meets the conditions for normal startup. If the device does not meet the requirements for normal startup, it can determine that an abnormal condition has occurred. In some embodiments of this application, the conditions for normal startup may include, but are not limited to, the fact that the core components of the electronic device are functioning normally. These core components may include, but are not limited to, UFS, System on Chip (SOC), and Double Data Rate Synchronous Dynamic Random Access Memory (DDR).
[0091] In some embodiments of this application, the electronic device can detect whether any abnormal situation occurs during the boot-up phase.
[0092] In some embodiments of this application, the electronic device detects whether an abnormal situation has occurred, which may specifically include one or more of the following: the electronic device detects whether the image file (which may be referred to as the image) is complete, the electronic device detects whether the device is damaged, and the electronic device detects whether there are problems with the interaction between UFS and SOC. Of course, the electronic device may also detect other aspects related to UFS, and this application does not limit this.
[0093] As we understand it, an image file is the physical carrier of an operating system; all the logic of system operation is reflected in the operating system's image file. Electronic devices can detect the integrity of image files, specifically by checking whether the image files involved in UFS are complete. The image file mentioned here can refer to a UFS file.
[0094] As can be understood, a System-on-a-Chip (SoC) refers to a system or product formed by combining multiple integrated circuits with specific functions on a single chip, including a complete hardware system and its embedded software. For example, an SoC can integrate a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), memory, a modem, a navigation and positioning module, and a multimedia module.
[0095] Understandably, if an electronic device detects an incomplete image file, it can determine that an anomaly has occurred. Similarly, if an electronic device detects a damaged component, it can determine that an anomaly has occurred. Likewise, if an electronic device detects a problem with the interaction between the UFS and the SOC (e.g., the UFS is not responding to messages sent by the SOC, or the SOC is not responding to messages sent by the UFS, etc.), it can determine that an anomaly has occurred.
[0096] In some embodiments of this application, during the startup phase of the electronic device, the electronic device checks for any abnormal conditions only once.
[0097] S103: When the electronic device detects an abnormal situation, the electronic device will store the generated fault information in the partition related to the startup of the electronic device.
[0098] When an electronic device detects an anomaly, it can record the corresponding fault information and store it in a partition related to the device's startup. In other words, after detecting an anomaly, the electronic device can store the generated fault information in a file within the corresponding storage area. It can be understood that the partition related to the electronic device's startup mentioned in this application specifically refers to the partition within the UFS partition related to the electronic device's startup.
[0099] For ease of description, this application refers to the fault information generated when an abnormality is detected during the boot process of an electronic device as fault information 1. That is, the fault information stored in the partition related to the boot of the electronic device is fault information 1.
[0100] As can be understood, a partition is a storage area in an electronic device. In some embodiments of this application, the partitions involved can refer to logical partitions created by dividing a physical hard drive. Each partition appears as an independent hard drive, with its own file system, folders, and files. Different types of data and operating systems can be stored in different partitions, which allows for better organization and management of files, as well as better protection of data security.
[0101] It is understood that partitions related to the booting of electronic devices can be configured according to actual needs, and this application does not impose any restrictions on this. For example, partitions related to the booting of electronic devices may include a recovery partition. It is understood that a recovery partition is a recovery partition containing a simple Linux system for restoring and maintaining the phone, and can also be used to erase, reboot, and perform other operations on other partitions. As another example, partitions related to the booting of electronic devices may include an xbl partition. It is understood that the xbl partition is one of the important partitions for system booting.
[0102] In some embodiments of this application, the partition associated with the startup of the electronic device is a specific partition. For example, the partition associated with the startup of the electronic device can be an xbl partition.
[0103] In some embodiments of this application, the partition related to the startup of the electronic device may include multiple specific partitions. For example, the partition related to the startup of the electronic device may include a recovery partition and an xbl partition. The recovery partition is used to record fault information caused by multiple startup failures, and the xbl partition is used to record fault information generated during the startup process.
[0104] In some embodiments of this application, after an electronic device fails to boot multiple times, fault information 1 can be generated and stored in the recovery partition and the xbl partition. In this case, the fault information 1 stored in the recovery partition and the xbl partition can be different.
[0105] In some embodiments of this application, the electronic device may determine which specific partition, which is included in the partition related to the startup of the electronic device, to store fault information 1, based on the timing of the detection of an anomaly.
[0106] In some embodiments of this application, the electronic device can record fault information 1 in a log and output the log containing fault information 1 to a partition related to the startup of the electronic device. That is, the electronic device can use the log to mark abnormal situations. In one possible implementation, the electronic device can create an additional log to record fault information 1, instead of recording fault information 1 in an already generated log.
[0107] It is understood that the fault information 1 may also be stored in other forms in a partition related to the startup of the electronic device, and the present application does not limit this.
[0108] S104: In standby mode, the electronic device scans the critical partition to determine if the critical partition is abnormal.
[0109] It is understood that the standby state involved in this application refers to the state after the kernel (e.g., the HLOS kernel mentioned in step S101) in the electronic device has started. In the standby state, the electronic device can first scan one or more specific critical partitions and record the data points. The recorded data points are then summarized to obtain the corresponding fault information. Based on this fault information, it is determined whether the critical partition is abnormal, that is, whether the abnormality detected during the startup phase is specifically a UFS fault affecting the critical image file. If the electronic device determines that the critical partition is abnormal, it can proceed to step S105. If the electronic device determines that the critical partition is not abnormal, it can proceed to step S108.
[0110] As we understand it, a critical partition refers to a partition that plays a crucial role in the operation of the operating system. Critical partitions can include vendor partitions and system partitions, among others. The vendor partition stores system files and drivers provided by the vendor. The system partition is the core partition of the Android system, containing all the operating system files and applications. Users can install and uninstall applications, update the system, and so on, within this partition. Of course, critical partitions can also include other partitions; this application does not impose specific limitations on this.
[0111] It is understood that the tracking information may include relevant records of the scanned partition, including the scan time and scan results. In some embodiments of this application, the tracking information recorded by the electronic device when scanning the critical partition may specifically include: a log obtained by the electronic device when scanning the critical partition, which records the scan time and scan results.
[0112] For ease of description, this application refers to the fault information generated by the electronic device scanning the critical partition in standby mode as fault information 2.
[0113] In some embodiments of this application, fault information 2 may involve one or more of the following: failed address, logical region access status, and single access (Input / Output, IO) status. Here, the failed address refers to the address corresponding to the data that failed to be read. Read failure may include IO exceptions (e.g., IO timeout), unresponsive read operations, etc. Logical region refers to the detection logic of a UFS driver. IO refers to the interaction with the storage medium.
[0114] Of course, fault information 2 may also involve other content, and this application does not impose specific restrictions on it.
[0115] In some embodiments of this application, fault information 2 may specifically include one or more of the following: the number of failed addresses, the number of logical region access failures, and the number of I / O failures. Logical region access failures may include logical region access timeouts, access failures, etc. I / O failures may include access timeouts, access failures, etc.
[0116] In some embodiments of this application, the electronic device determines that the critical partition is abnormal based on the fault information 2, which may specifically include one or more of the following: the electronic device determines that the number of failed addresses included in the fault information 2 is not less than a threshold 1; the electronic device determines that the number of logical area access failures included in the fault information 2 is not less than a threshold 2; the electronic device determines that the number of IO failures included in the fault information 2 is not less than a threshold 3.
[0117] It is understood that thresholds 1, 2, and 3 are positive integers, and their specific values can be set according to actual needs; this application does not impose any restrictions on this. Thresholds 1, 2, and 3 can be the same or different. In some embodiments of this application, thresholds 1, 2, and 3 can be determined based on factors such as the device's service life, model, and the device's own lifespan. That is, when the service life of the electronic device changes, thresholds 1, 2, and 3 can be determined accordingly based on the changed service life. For example, if the electronic device has been used for 1 year, threshold 1 is 2; while if the electronic device has been used for 5 years, threshold 1 is 6.
[0118] In some embodiments of this application, the electronic device may scan critical partitions within a corresponding process after the corresponding process is started. It is understood that this corresponding process can be configured according to actual needs, and this application does not impose specific limitations on it.
[0119] S105: Electronic devices migrate critical image files that are affected in critical partitions.
[0120] If an electronic device determines that a critical partition is faulty, specifically a UFS failure affecting critical image files, the device can migrate the affected critical image files from that partition. Specifically, the device can migrate these affected critical image files to other storage areas to ensure the normal operation of core functions such as startup. This other storage area can be understood as the backup partition corresponding to UFS. The backup partition is a pre-defined partition created to prevent UFS failures and is different from the original partition corresponding to UFS (including partitions related to electronic device startup, critical partitions, and frequently used partitions).
[0121] As is understood, a critical image file refers to an image file located in a critical partition that affects the boot process of an electronic device. These critical partitions can include the vendor partition, the boot partition, and the system partition. The primary task of the boot partition is to provide all the files needed during system startup. The specific functions of the vendor partition, boot partition, and system partition can be further explained in relevant technical documentation; this application will not elaborate on them here.
[0122] In some embodiments of this application, the electronic device can further determine whether the failure addresses included in the fault information 2 are clustered. In one possible implementation, if the electronic device determines that the critical partition is abnormal according to the specific implementation method provided in step S104, and the failure addresses included in the fault information 2 are clustered, the electronic device can execute step S105. In yet another possible implementation, regardless of whether the failure addresses included in the fault information 2 are clustered, the electronic device will execute step S105 if it determines that the critical partition is abnormal according to the specific implementation method provided in step S104.
[0123] In some embodiments of this application, the electronic device may also determine whether the critical partition is abnormal by combining the set of failure addresses involved in the fault information 2.
[0124] In some embodiments of this application, the electronic device determines whether the failure addresses included in the fault information 2 are clustered. Specifically, the electronic device can determine whether the failure addresses included in the fault information 2 within the same sector are not less than a%. If the failure addresses included in the fault information 2 within the same sector are not less than a%, then the electronic device can determine that the failure addresses are clustered.
[0125] It is understood that a sector refers to a storage area smaller than a partition. It is understood that 'a' is less than 100 and greater than 0. 'a' can be set according to actual needs; this application does not restrict its specific value. For example, 'a' can be 70. Another example is that 'a' can be 85.
[0126] In some embodiments of this application, the electronic device determines whether the failure addresses included in the fault information 2 are concentrated. Specifically, the electronic device can determine that the number of failure addresses whose corresponding numbers in the fault information 2 are less than b and whose corresponding numbers are less than b is greater than c.
[0127] In some embodiments of this application, the electronic device determines whether the failure addresses included in the fault information 2 are concentrated. Specifically, the electronic device can determine that the number of failure addresses whose corresponding numbers in the fault information 2 are less than b, and the ratio of the number of failure addresses included in the fault information 2 to the total number of failure addresses included in the fault information 2 is greater than d.
[0128] It is understood that c is a positive integer, and b, c and d can be set according to actual needs. This application does not restrict their specific values.
[0129] For example, b can be 100, c can be 3, and the failure addresses included in fault information 2 can include 135842, 165212, 135840, 135846, and 135870. Among the above failure addresses, the failure addresses whose corresponding numbers differ from the numbers corresponding to other failure addresses by less than 100 include 135842, 135840, 135846, and 135870. If this number is greater than 3, the electronic device can determine the set of failure addresses included in fault information 2.
[0130] S106: The electronic device determines whether the critical image files affected in the critical partition have been successfully migrated.
[0131] After migrating the affected critical image files in the critical partition, the electronic device can determine whether the migration was successful. If the electronic device successfully migrates the affected critical image files in the critical partition, it can proceed to step S107; if the electronic device fails to migrate the affected critical image files in the critical partition, it can proceed to step S110.
[0132] S107: Electronic equipment continues to operate.
[0133] If the electronic device successfully migrates the affected critical image files from the critical partition, it indicates that the electronic device is working normally, and it can continue to operate as usual.
[0134] S108: In standby mode, the electronic device scans frequently used partitions to determine if they are abnormal.
[0135] In standby mode, if the electronic device determines that the critical partition is normal, it can continue scanning one or more specific commonly used partitions and recording the data points. The recorded data points are then summarized to obtain corresponding fault information. Based on this fault information, it is determined whether the critical partition is abnormal, i.e., whether the abnormality detected during the startup phase is a UFS fault other than affecting the critical image file. If the electronic device determines that the commonly used partition is abnormal, it can proceed to step S109; if it determines that the commonly used partition is normal, it can proceed to step S107.
[0136] As is understood, frequently used partitions refer to those that the operating system relies on for operation. Frequently used partitions can include data partitions, original design manufacturer (ODM) partitions, and product partitions. Data partitions can be used to store configuration files and data required for system operation. ODM partitions are used by original design manufacturers to customize their own board support packages, i.e., they contain the original design manufacturer's customization of the system-on-a-chip vendor's board support package. Product partitions are an extension of the system partition and can be used to store system components and themes. The specific functions of data partitions, ODM partitions, and product partitions can be further explained in relevant technical documents; this application will not elaborate on them here.
[0137] In some embodiments of this application, the logging information recorded by the electronic device scanning the frequently used partitions may specifically include: a log obtained by the electronic device scanning the frequently used partitions, which records the scanning time and the scanning results.
[0138] For ease of description, this application refers to the fault information generated by the electronic device scanning the commonly used partitions in standby mode as fault information 3.
[0139] In some embodiments of this application, fault information 3 may involve one or more of the following: failed address, logical region access status, and I / O status. Of course, fault information 3 may also involve other content, and this application does not impose specific limitations on it.
[0140] In some embodiments of this application, fault information 3 may specifically include one or more of the following: the number of failed addresses, the number of logical region access failures, and the number of I / O failures. Of course, fault information 3 may also include other content, and this application does not limit this.
[0141] In some embodiments of this application, the electronic device determines that the commonly used partition is abnormal based on the fault information 3, which may specifically include one or more of the following: the electronic device determines that the number of failed addresses included in the fault information 3 is not less than a threshold 4; the electronic device determines that the number of logical area access failures included in the fault information 3 is not less than a threshold 5; the electronic device determines that the number of IO failures included in the fault information 3 is not less than a threshold 6.
[0142] It is understood that thresholds 4, 5, and 6 are positive integers, and their specific values can be set according to actual needs; this application does not impose any restrictions on this. Thresholds 4, 5, and 6 can be the same or different. In some embodiments of this application, thresholds 4, 5, and 6 can be determined based on factors such as the device's service life, model, and the device's own lifespan. That is, when the service life of the electronic device changes, thresholds 4, 5, and 6 can be determined accordingly based on the changed service life. For example, if the electronic device has been used for 1 year, threshold 4 is 3; while if the electronic device has been used for 5 years, threshold 4 is 7.
[0143] It is understood that threshold 1 and threshold 4 can be the same or different, threshold 2 and threshold 5 can be the same or different, and threshold 3 and threshold 6 can be the same or different. In some embodiments of this application, threshold 4 is greater than threshold 1, threshold 5 is greater than threshold 2, and threshold 6 is greater than threshold 3.
[0144] In some embodiments of this application, the electronic device may skip step S108 and directly execute step S109. Specifically, if the electronic device determines that the critical partition is not abnormal, the electronic device may directly execute step S109, that is, the electronic device may directly determine whether the abnormality of the commonly used partition meets the preset condition 1.
[0145] In some embodiments of this application, after the electronic device starts up, it can execute step S104 once, and then execute step S108 at a certain frequency, such as executing step S108 three times a week. After the electronic device starts up again, it can execute step S104 once more, and then execute step S108 at a certain frequency. That is to say, in the standby state after each startup, the electronic device can scan the critical partition once, and then scan the frequently used partition at a certain frequency. That is, the critical partition is scanned only once, while the frequently used partition can be scanned multiple times.
[0146] S109: The electronic device determines whether the abnormal situation of the commonly used partition meets the preset condition 1.
[0147] If the electronic device determines that the critical partition is normal, but the frequently used partition is abnormal, the electronic device can further determine whether the abnormality of the frequently used partition meets preset condition 1. If the electronic device determines that the abnormality of the frequently used partition meets preset condition 1, the electronic device can execute step S110; if the electronic device determines that the abnormality of the frequently used partition does not meet preset condition 1, the electronic device can execute step S107. It can be understood that if the abnormality of the frequently used partition meets preset condition 1, it indicates that the UFS failure is relatively serious.
[0148] In some embodiments of this application, the electronic device determines that the abnormal situation of the commonly used partition meets the preset condition 1, which may specifically include: the electronic device determines that the number of failed addresses included in the fault information 3 is not less than the threshold 7.
[0149] In some embodiments of this application, the electronic device determines that the abnormal situation of the commonly used partition meets the preset condition 1, which may specifically include: the electronic device determines that the number of PA access failures included in the fault information 3 is not less than the threshold 8.
[0150] In some embodiments of this application, the electronic device determines that the abnormal situation of the commonly used partition meets the preset condition 1, which may specifically include: the electronic device determines that the number of IO failures included in the fault information 3 is not less than the threshold 9.
[0151] It is understood that thresholds 7, 8, and 9 are positive integers, and their specific values can be set according to actual needs; this application does not impose any restrictions on this. Thresholds 7, 8, and 9 can be the same or different. In some embodiments of this application, thresholds 7, 8, and 9 can be determined based on factors such as the device's service life, model, and the device's own lifespan. That is, when the service life of the electronic device changes, thresholds 7, 8, and 9 can be determined accordingly based on the changed service life. For example, if the electronic device has been used for 1 year, threshold 4 is 10, while if the electronic device has been used for 5 years, threshold 4 is 30.
[0152] It should be noted that threshold 7 is greater than threshold 4, threshold 8 is greater than threshold 5, and threshold 9 is greater than threshold 6.
[0153] In some embodiments of this application, the electronic device determines that the abnormal situation of a frequently used partition meets preset condition 1. Specifically, this may include: the electronic device can determine that the number of failed interactions between the UFS and SOC involved in fault information 3 is not less than K. If the number of failed interactions between the UFS and SOC involved in fault information 3 is not less than K, the electronic device can determine that the UFS fault is relatively serious. It is understood that K is a positive integer, and its specific value can be set according to actual needs; this application does not impose any restrictions on it.
[0154] In some embodiments of this application, the electronic device determines that the abnormal situation of a frequently used partition meets preset condition 1. Specifically, this may include: the electronic device determining whether the ratio of the number of failed interactions between the UFS and SOC involved in fault information 3 to the total number of interactions is not less than L. If the ratio of the number of failed interactions between the UFS and SOC involved in fault information 3 to the total number of interactions between the UFS and SOC involved in fault information 3 is not less than L, then the electronic device can determine that the UFS fault is relatively serious. It is understood that L is not less than 0 and not greater than 1, and its specific value can be set according to actual needs. This application does not impose any restrictions on this.
[0155] It is understood that the interaction between UFS and SOC mentioned in this application may include, but is not limited to: UFS reading relevant information from SOC, SOC reading relevant information from UFS, and message sending and receiving between UFS and SOC.
[0156] In some embodiments of this application, the electronic device determines that the abnormal situation of the commonly used partition meets the preset condition 1, which may specifically include: whether the failure addresses included in the fault information 3 are concentrated. It is understood that the specific implementation method adopted by the electronic device to determine whether the failure addresses included in the fault information 3 are concentrated can refer to the above (such as the specific judgment method for whether the failure addresses included in the fault information 2 are concentrated in step S105), and will not be repeated here.
[0157] In some embodiments of this application, the electronic device can determine the fault level corresponding to the current UFS fault based on the fault information 3.
[0158] In one possible implementation, if the fault level corresponding to this UFS fault is among the top b% of fault levels, the electronic device can determine that this UFS fault is serious, that is, the electronic device can determine that the anomaly of the commonly used partition meets the preset condition 1.
[0159] It is understood that b is less than 100 and greater than 0. b can be set according to actual needs, and this application does not limit its specific value. For example, b can be 30, and the fault level can include four levels: Level 1, Level 2, Level 3, and Level 4. Among these four fault levels, Level 1 represents the most severe UFS fault, belonging to the highest fault level, while the severity of UFS faults represented by Levels 2, 3, and 4 decreases in that order. In this case, only when the fault level corresponding to the UFS fault is in the top 30% of fault levels is the UFS fault considered relatively severe. In this case, the abnormal situation of the commonly used partition meets the preset condition 1. Since 30% * 4 = 1.2, the abnormal situation of the commonly used partition only meets the preset condition 1 when the fault level corresponding to the UFS fault is Level 1.
[0160] In one possible implementation, if the fault level corresponding to this UFS fault is the most severe fault level, the electronic device can determine that the anomaly of the commonly used partition meets the preset condition 1.
[0161] For example, the fault levels can include a first level, a second level, and a third level, with the severity of the UFS fault decreasing sequentially. If the fault level corresponding to this UFS fault is the first level, the electronic device can determine that the anomaly of the commonly used partition meets preset condition 1.
[0162] In one possible implementation, if the number of fault levels is greater than 2, and the fault level corresponding to this UFS fault is one of the two most severe fault levels, then the electronic device can determine that the anomaly of the commonly used partition meets the preset condition 1.
[0163] For example, the fault levels can include Level 1, Level 2, Level 3, Level 4, and Level 5, with the severity of the UFS fault decreasing sequentially. If the fault level corresponding to this UFS fault is Level 1 or Level 2, the electronic device can determine that the anomaly of the commonly used partition meets preset condition 1.
[0164] In some embodiments of this application, the electronic device may determine the fault level corresponding to the UFS fault based on one or more of the following:
[0165] (1) The number of failed addresses included in fault information 3;
[0166] (2) Fault information 3 includes the number of logical area access failures;
[0167] (3) Fault information 3 includes the number of IO failures;
[0168] (4) The number of failed interactions between UFS and SOC involved in fault information 3;
[0169] (5) The ratio of the number of failed interactions between UFS and SOC involved in fault information 3 to the total number of interactions;
[0170] (6) Fault information 3 includes the concentration of failure addresses.
[0171] It should be noted that if the abnormality of the frequently used partition does not meet the preset condition 1, the electronic device needs to continue monitoring the abnormality of the frequently used partition and continue collecting relevant fault information of UFS while continuing to work. In some embodiments of this application, if the electronic device determines that the frequently used partition is abnormal, but its abnormality does not meet the preset condition 1, the electronic device can accumulate the fault information 3 generated after multiple startups, and then determine whether the abnormality of the frequently used partition meets the preset condition 1 based on the accumulated fault information 3.
[0172] S110: Electronic device start flag.
[0173] If the electronic device fails to migrate the affected critical image file in the critical partition, or if the electronic device determines that the abnormal situation of the commonly used partition does not meet preset condition 1, the electronic device can set a flag. In some embodiments of this application, setting a flag may specifically include: the electronic device setting a flag and setting the flag to 1. It can be understood that a flag can also be called a status bit, referring to using a variable to record information about a call without affecting other information.
[0174] In some embodiments of this application, if the electronic device fails to migrate the critical image file affected in the critical partition, or if the electronic device determines that the abnormal situation of the commonly used partition does not meet the preset condition 1, the electronic device can set the first parameter to the first content.
[0175] It is understood that the first parameter can be set according to actual needs, and this application does not restrict the specific form of the first parameter (e.g., letter, string, number, etc.). For example, the first parameter can be represented by the letter 'x'. It is also understood that the first content can be set according to actual needs, and this application does not impose specific restrictions on its specific form (e.g., letter, string, number, etc.) and content. For example, the first content can be '1'. As another example, the first content can be 'true'.
[0176] S111: The electronic device receives a command to start the electronic device, and in response to the command, the electronic device starts up.
[0177] It is understood that after the electronic device receives the instruction to start the electronic device again, the electronic device can start in response to the instruction to start the electronic device, as can be seen in the relevant description of step S101.
[0178] S112: If an abnormality is detected during the startup of the electronic device and it is about to enter the emergency download mode, the electronic device determines whether the flag bit has been set.
[0179] During the boot process of an electronic device, if the device detects an anomaly during the boot phase and triggers an imminent entry into emergency download mode, the device determines whether a flag has been set. For example, if the electronic device detects an anomaly during the boot phase and triggers an imminent entry into emergency download mode, and the flag is set to 1, the device can determine that the flag has been set.
[0180] It is understandable that an anomaly is detected during the startup process of an electronic device, triggering an emergency download mode. This may include one or more of the following: failure of UFS and SOC interaction, inability to read the image file in UFS, or loading failure.
[0181] In some embodiments of this application, if an abnormality is detected during the startup process of the electronic device and an emergency download mode is triggered, the electronic device determines whether the first parameter is specifically the first content.
[0182] S113: Electronic devices start the backup operating system.
[0183] If an anomaly is detected during the startup process of an electronic device, triggering an emergency download mode, and the electronic device confirms that the flag has been set, then the electronic device can start the backup operating system.
[0184] In some embodiments of this application, the backup operating system can be a recovery partition that shields non-critical disk read / write operations. It is an alternative operating system, different from the recovery partition in the original operating system (e.g., Android system) used by the electronic device.
[0185] In some embodiments of this application, the backup operating system can be a fastboot partition that shields non-critical disk read / write operations. It is an alternative operating system, different from the fastboot partition in the original operating system (e.g., Android) used by the electronic device.
[0186] Understandably, after an electronic device boots up a backup operating system, it can support some maintenance and testing, such as capturing fault information 1, fault information 2, and fault information 3 to determine the faulty partition in UFS, or skipping the faulty partition to perform recovery operations.
[0187] In some embodiments of this application, the electronic device may skip steps S101-S103 and execute only the logic of steps S104-S113.
[0188] The following describes another fault handling method for memory provided by an embodiment of this application.
[0189] Please see Figure 2B , Figure 2B A flowchart illustrating a fault handling method for a memory provided in an embodiment of this application. Figure 2B For details on the implementation of each step (steps S101-S113), please refer to the above text. Figure 2A Related descriptions. (and) Figure 2A The method shown is different, such as Figure 2B As shown, after the electronic device executes step S103, it can continue to execute steps S104 and S108, instead of as... Figure 2AAs shown, step S108 is executed only after step S104 has been performed and it has been determined that there are no abnormalities in the critical partition.
[0190] In some embodiments of this application, the electronic device may first perform step S104 and then perform step S108.
[0191] In some embodiments of this application, after the electronic device starts up, it can execute step S104 once and step S108 at a certain frequency, such as executing step S108 three times a week. After the electronic device starts up again, it can execute step S104 once more and execute step S108 at a certain frequency. That is to say, in the standby state after each startup, the electronic device can scan the critical partition once and can scan the frequently used partition at a certain frequency. That is, the critical partition is scanned only once, while the frequently used partition can be scanned multiple times.
[0192] The apparatus involved in the embodiments of this application is described below.
[0193] Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.
[0194] like Figure 3 As shown, the electronic device may include: a processor, a mobile communication module, a wireless communication module, an audio module, a speaker, a receiver, a microphone, a headphone jack, internal memory, an external memory interface, a sensor module, a Subscriber Identity Module (SIM) card slot, a display screen, a camera, buttons, and a Universal Serial Bus (USB) interface, etc. The sensor module may include: a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer, a proximity sensor, a near-light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0195] The structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In some embodiments of this application, the electronic device may include more components than illustrated. For example, other types of sensors. Other examples include a charging management module, a power management module, a battery, a motor, indicators, etc. In some embodiments of this application, the electronic device may include fewer components than illustrated, or combine some components, or split some components, or arrange different components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware. The interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device.
[0196] A processor may include one or more processing units, such as a System-on-a-Chip (SoC), Application Processor (AP), Modem Processor, Graphics Processing Unit (GPU), Image Signal Processor (ISP), Controller, Codec (e.g., video codec), Digital Signal Processor (DSP), Baseband Processor, and / or Neural-network Processing Unit (NPU). The codec is used to transform a signal or a data stream. The processor may also include memory for storing instructions and data.
[0197] Electronic devices can implement display functions through GPUs, displays, and application processors. A GPU is a microprocessor for image processing, connecting the display and the application processor. GPUs are used to perform mathematical and geometric calculations for graphics rendering. A processor may include one or more GPUs, which execute program instructions to generate or modify display information. A display is used to display images, videos, etc. In some embodiments, an electronic device may include one or more displays.
[0198] A camera is used to capture still images or videos. An ISP (Image Signal Processor) processes the data returned by the camera. A camera may include a lens and an image sensor. An image sensor is a light-sensitive element. Light is transmitted through the lens to the image sensor, where the light signal is converted into an electrical signal. The image sensor then transmits this electrical signal to the ISP for processing, transforming it into a visible image.
[0199] It is understood that internal memory may include UFS. In some embodiments of this application, internal memory may include one or more RAMs and one or more non-volatile memory (NVMs). Random access memory can be directly read and written by the processor and can be used to store executable programs (e.g., machine instructions) of the operating system or other running programs, as well as user and application data. Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into random access memory for direct read and write operations by the processor.
[0200] In this embodiment, the code implementing the method described in this embodiment can be stored in non-volatile memory. After the electronic device is powered on, the electronic device can load the executable code stored in the non-volatile memory into random access memory.
[0201] External memory interfaces can be used to connect to external non-volatile memory, thereby expanding the storage capacity of electronic devices.
[0202] Electronic devices can implement audio functions through audio modules, speakers, receivers, microphones, headphone jacks, and application processors.
[0203] The audio module converts digital audio information into analog audio signals for output, and also converts analog audio input into digital audio signals. The speaker, also called a "horn," converts audio electrical signals into sound signals. The receiver, also called a "handset," converts audio electrical signals into sound signals. The microphone, also called a "microphone" or "voice transducer," converts sound signals into electrical signals. The headphone jack is used to connect wired headphones.
[0204] A touch sensor, also known as a "touch device," is a component of an electronic device. Touch sensors can be mounted on a display screen, and the touch sensor and display screen together form a touchscreen, also called a "touchscreen." The touch sensor detects touch operations applied to or near it. It then transmits the detected touch operation to an application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen. In some embodiments, the touch sensor may also be located on the surface of the electronic device, in a different position than the display screen.
[0205] It is understandable that the meaning and function of other sensors can be found in relevant technical documents, and will not be explained in detail here.
[0206] Electronic devices can communicate with other devices through mobile communication modules and wireless communication modules.
[0207] Mobile communication modules can provide solutions for wireless communication applications, including 2G / 3G / 4G / 5G, in electronic devices.
[0208] Wireless communication modules can provide solutions for wireless communication applications in electronic devices, including Wireless Local Area Networks (WLANs) (e.g., Wireless Fidelity networks, commonly known as Wi-Fi networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR) technologies. A wireless communication module can be one or more devices integrating at least one communication processing module. The wireless communication module can receive electromagnetic waves via antenna 2 or convert signals into electromagnetic waves for radiation.
[0209] The software system of an electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. Among these, a layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. This application uses the layered architecture of the Android system as an example to illustrate the software structure of an electronic device.
[0210] Figure 4 This is a schematic diagram of the software structure of an electronic device provided in an embodiment of this application.
[0211] like Figure 4 As shown, the software framework of an electronic device may include an application layer, an application framework layer, system libraries, a runtime, a hardware abstraction layer (HAL), and a kernel layer.
[0212] The application layer can include a series of application packages. Generally, the application layer can include system applications and third-party applications. System applications are a series of core applications that come with the Android system. For example, system applications may include applications such as camera, music, calendar, SMS, WLAN, calling, and gallery. Third-party applications are applications that users can choose to install. In some embodiments of this application, users can download and install third-party applications from app stores.
[0213] The application framework layer provides an Application Programming Interface (API) and programming framework for applications in the application layer. The application framework layer includes predefined functions. It may also include a series of system services. System services are modular components focused on specific functionalities. The functionality provided by the application framework API allows communication with these system services to access the underlying hardware.
[0214] like Figure 4 As shown, the application framework layer can include a content provider, window manager, resource manager, phone manager, notification manager, and view system. The window manager manages window programs. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture the screen. The content provider stores and retrieves data, making this data accessible to applications. This data can include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc. The view system includes visual controls, such as controls for displaying text and controls for displaying images. The view system can be used to build applications. The display interface can consist of one or more views. For example, a display interface including a text notification icon can include a view for displaying text and a view for displaying images. The phone manager provides communication functions for electronic devices, such as managing call status (including connection, hang-up, etc.). The resource manager provides various resources to applications, such as localized strings, icons, images, layout files, video files, etc. The notification manager allows applications to display notification information in the status bar. It can be used to convey informational messages and can disappear automatically after a short pause without user interaction. For example, the notification manager is used to notify of download completion, message alerts, etc. The notification manager can also display notifications as icons or scrolling text in the system's top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting alert sounds, causing electronic devices to vibrate, and flashing indicator lights.
[0215] System libraries can include multiple functional modules. Examples include: Surface Manager, Media Libraries, 3D graphics processing libraries (e.g., OpenGL ES), and 2D graphics engines (e.g., SGL). The specific meanings and functions of these modules can be found in relevant technical documentation and will not be elaborated upon here.
[0216] The runtime is responsible for system scheduling and management. The runtime includes core libraries and virtual machines.
[0217] The Hardware Abstraction Layer (HAL) is an interface layer located between the operating system kernel and upper-level software, its purpose being to abstract hardware. The HAL is an abstract interface for device kernel drivers, providing application programming interfaces (APIs) for accessing the underlying device to higher-level Java API frameworks. The HAL can provide a standard interface that exposes device hardware functionality to higher-level Java API frameworks. The HAL contains multiple library modules, each implementing an interface for a specific type of hardware component. When the system framework layer API requests access to the portable device's hardware, the operating system loads the library module for that hardware component.
[0218] The kernel layer is the layer between hardware and software. It is the foundation of the Android system. The kernel layer is responsible for hardware drivers, networking, power, system security, and memory management. As an intermediary between hardware and software, the kernel layer's role is to pass application requests to the hardware. The kernel layer can include display drivers, camera drivers, audio drivers, and sensor drivers, among others.
[0219] It should be noted that the application provides Figure 4 The illustrated software architecture diagram of the electronic device is merely an example and does not limit the specific module divisions within different layers of the Android system. For details, please refer to the descriptions of the Android system software architecture in conventional technical documents. Furthermore, the display method provided in this application can also be implemented on other operating systems, which will not be listed here.
[0220] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A fault handling method for memory, characterized in that, The method is applied to an electronic device equipped with memory, wherein the memory partitions include critical partitions and frequently used partitions, and the method includes: The electronic device detects an operation that triggers it to start, and in response to the operation, the electronic device begins to start. If the electronic device detects an anomaly in the first state that satisfies the need to enter the emergency download mode, and if the electronic device, in the second state before startup, scans and confirms that the critical partition is abnormal and the migration of the critical image file affected in the critical partition fails, or scans and confirms that the commonly used partition is abnormal and meets the first preset condition, then the electronic device starts the first operating system. The first state refers to the state when the electronic device is in the boot phase, and the second state refers to the state after the kernel of the second operating system in the electronic device has started.
2. The method as described in claim 1, characterized in that, Before the electronic device detects an operation that triggers the electronic device to start, the method further includes: In the second state, the electronic device scans the critical partition, generates first fault information, and determines whether the critical partition is abnormal based on the first fault information; In the event of an anomaly in the critical partition, the electronic device migrates the critical image files affected in the critical partition; If the electronic device fails to migrate the critical image file, the electronic device will set the content corresponding to the first parameter to the first content. If no abnormality is found in the critical partition, the electronic device scans the commonly used partition, generates second fault information, and determines whether the commonly used partition is abnormal based on the second fault information. If an anomaly occurs in the commonly used partition, and the anomaly in the commonly used partition meets the first preset condition, the electronic device sets the content corresponding to the first parameter to the first content. If, in the second state prior to startup, the electronic device scans and confirms that the critical partition is abnormal and the migration of the critical image file affected by the critical partition fails, or scans and confirms that the commonly used partition is abnormal and meets the first preset condition, then the electronic device starts the first operating system. Specifically, if the content corresponding to the first parameter is the first content, then the electronic device starts the first operating system.
3. The method as described in claim 1, characterized in that, Before the electronic device detects an operation that triggers the electronic device to start, the method further includes: In the second state, the electronic device scans the critical partition and the commonly used partition, generates first fault information and second fault information respectively, and determines whether the critical partition and the commonly used partition are abnormal based on the first fault information and the second fault information respectively; In the event of an anomaly in the critical partition, the electronic device migrates the critical image files affected in the critical partition; If an anomaly occurs in the commonly used partition and the anomaly in the commonly used partition meets the first preset condition, or if the migration of the critical image file fails, the electronic device sets the content corresponding to the first parameter to the first content. If, in the second state prior to startup, the electronic device scans and confirms that the critical partition is abnormal and the migration of the critical image file affected by the critical partition fails, or scans and confirms that the commonly used partition is abnormal and meets the first preset condition, then the electronic device starts the first operating system. Specifically, if the content corresponding to the first parameter is the first content, then the electronic device starts the first operating system.
4. The method as described in claim 2 or 3, characterized in that, The electronic device determines that the critical partition is abnormal based on the first fault information, specifically including any one or more of the following: The electronic device determines that the number of failed addresses included in the first fault information is not less than a first threshold. The electronic device determines that the number of failed accesses to the logical area included in the first fault information is not less than a second threshold. The electronic device determines that the number of single access IO failures included in the first fault information is not less than a third threshold.
5. The method as described in claim 4, characterized in that, The electronic device determines that the critical partition is abnormal based on the first fault information, and specifically further includes: the electronic device determining the set of failed addresses included in the first fault information.
6. The method as described in claim 5, characterized in that, The first fault information includes the set of failed addresses, specifically including any one or more of the following: The ratio of the number of failed addresses belonging to the same sector included in the first fault information to the total number of failed addresses included in the first fault information is not less than a first ratio. The number of first-type failure addresses included in the first fault information is greater than a first value; the difference between the number corresponding to the first-type failure address and the number corresponding to other failure addresses included in the first fault information is less than a second value; The ratio of the number of first-type failed addresses included in the first fault information to the total number of all failed addresses included in the first fault information is greater than the second ratio.
7. The method as described in claim 2 or 3, characterized in that, The electronic device determines that the commonly used partition is abnormal based on the second fault information, specifically including any one or more of the following: The electronic device determines that the number of failed addresses included in the second fault information is not less than a fourth threshold. The electronic device determines that the number of failed accesses to the logical area included in the second fault information is not less than the fifth threshold. The electronic device determines that the number of I / O failures included in the second fault information is not less than a sixth threshold.
8. The method as described in claim 7, characterized in that, The anomaly occurring in the commonly used partition satisfies the first preset condition, specifically including any one or more of the following: The number of failed addresses included in the second fault information is not less than the seventh threshold; The second fault information includes a logical area access failure count that is not less than the eighth threshold; The second fault information includes an I / O failure count that is not less than the ninth threshold; Wherein, the seventh threshold is greater than the fourth threshold, the eighth threshold is greater than the fifth threshold, and the ninth threshold is greater than the sixth threshold.
9. The method as described in claim 2 or 3, characterized in that, When the electronic device determines that the commonly used partition is malfunctioning based on the second fault information, the method further includes: The electronic device determines the fault level of the memory based on the second fault information; If the electronic device determines that the fault level of the memory meets the second preset condition, the electronic device also determines that the anomaly occurring in the commonly used partition meets the first preset condition.
10. The method as described in claim 9, characterized in that, The electronic device determines the fault level of the memory based on the second fault information, specifically including any one or more of the following: The electronic device determines the fault level of the memory based on the number of failed addresses included in the second fault information; The electronic device determines the fault level of the memory based on the number of logical region access failures included in the second fault information; The electronic device determines the fault level of the memory based on the number of I / O failures included in the second fault information; The electronic device determines the fault level of the memory based on the set of failure addresses included in the second fault information; The electronic device determines the fault level of the memory based on the number of interaction failures between the memory and the processor in the electronic device as per the second fault information. The electronic device determines the fault level of the memory based on the ratio of the number of failed interactions between the memory and the processor as described in the second fault information to the total number of interactions between the memory and the processor.
11. The method according to any one of claims 2, 3, 5, 6, 8, and 10, characterized in that, The electronic device scans the frequently used partitions to generate the second fault information, specifically including: the electronic device scans the frequently used partitions at a preset frequency, and after each scan of the frequently used partitions, the electronic device generates corresponding second fault information.
12. The method as described in claim 11, characterized in that, After the electronic device scans the frequently used partitions, the method further includes: If an anomaly occurs in the frequently used partition, and the anomaly does not meet the first preset condition, the electronic device saves the second fault information obtained in this scan. If the anomaly is determined in the next scan, the electronic device combines the second fault information obtained from the previous scan of the frequently used partition with the second fault information obtained after the next scan of the frequently used partition to determine whether the anomaly of the frequently used partition meets the first preset condition. The second fault information obtained from the previous scan of the frequently used partition includes the second fault information obtained in this scan.
13. The method according to any one of claims 1, 2, 3, 5, 6, 8, 10, and 12, characterized in that, The partition corresponding to the memory also includes a partition related to the startup of the electronic device; Before the electronic device detects an operation that triggers the electronic device to start, the method further includes: If the electronic device detects an abnormality during the boot phase of the previous startup process, the electronic device generates third fault information and stores the third fault information in the partition related to the startup of the electronic device.
14. The method as described in claim 13, characterized in that, After the electronic device boots the first operating system, the method further includes: The electronic device captures the third fault information to determine the faulty partition of the memory, and performs a recovery operation on the faulty partition of the memory.
15. The method according to any one of claims 2, 3, 5, 6, 8, 10, and 12, characterized in that, After the electronic device boots the first operating system, the method further includes: The electronic device captures the first fault information and the second fault information to determine the faulty partition of the memory, and performs a recovery operation on the faulty partition of the memory.
16. The method according to any one of claims 1, 2, 3, 5, 6, 8, 10, 12, and 14, characterized in that, The first operating system is either a recovery partition that blocks non-critical disk read / write operations, or a fastboot partition that blocks non-critical disk read / write operations.
17. The method according to any one of claims 1, 2, 3, 5, 6, 8, 10, 12, and 14, characterized in that, The abnormality detected by the electronic device in the first state satisfies the requirement to enter the emergency download mode, specifically including any one or more of the following: The memory and the processor in the electronic device failed to interact; The electronic device cannot read the image file in the memory; The electronic device failed to load.
18. An electronic device comprising one or more memories and one or more processors, characterized in that, The memory is used to store a computer program; the processor is used to invoke the computer program to cause the electronic device to perform the method of any one of claims 1-17.
19. A computer storage medium, characterized in that, include: Computer instructions; when the computer instructions are executed on an electronic device, causing the electronic device to perform the method of any one of claims 1-17.