A method and device for automatically detecting memory failure
The memory fault light signal is obtained through the photoresistor and passed to the BMC and OS, which realizes timely recording and independent transmission of memory fault information, solves the problem that BMC and OS are difficult to capture fault information, and improves the reliability and integrity of memory fault detection.
Patent Information
- Application Number
- CN202211260613.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-10-14
AI Technical Summary
In the prior art, after the memory failure light is on inside the server, it is difficult for the BMC and OS to capture the fault information in time, resulting in missing memory failures, especially when the fault light only flashes, it is more likely to be ignored.
The memory fault light light illumination signal is obtained through the photoresistor and passed it to the BMC and OS. The BMC and OS collect and log record the memory fault data, including SEL log, IDLE log, black box log, message log, dmesg log, mce log and dmidecode-t memory information compression and storage.
Ensure that memory fault information is recorded in the photosensitive recording document, which is convenient for testing and developer analysis, prevent omissions, and the two links are independently transmitted, ensuring that information transmission is not interrupted, and improving the timeliness and accuracy of fault information.
Smart Images

Figure CN115827367B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer detection technology, and in particular to a method and device for automatically detecting memory failures, a computer device, and a storage medium. Background Art
[0002] In recent years, the application scope of servers in our lives has become increasingly wider, and our requirements for servers have become higher and higher, which has promoted the rapid development of servers.
[0003] As the demand for servers increases, the server level is also required to be higher and higher, which indirectly improves the quality of servers. The quality of servers is inseparable from the timely detection of problems and causes when problems occur.
[0004] The memory is located inside the server. When a memory failure occurs, the memory fault indicator will light up. However, since the memory fault indicator is also inside the server, it cannot be detected visually. Instead, it can be detected through memory alarm information in the BMC log or OS log. This can lead to missed memory failures. Furthermore, if the memory fault indicator only lights up briefly and then turns off immediately, the BMC and OS may not detect the problem during this time, which can also lead to missed memory failures. Summary of the Invention
[0005] In order to solve the technical problems existing in the above-mentioned prior art, the present invention provides a method, device, computer equipment and storage medium for automatically detecting memory faults. When the memory fault light of the present invention lights up, logs are captured and alarms are issued in a timely manner; when the memory fault light only flashes once, the BMC and OS cannot capture fault information in a timely manner.
[0006] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0007] In a first aspect, in one embodiment provided by the present invention, a method for automatically detecting memory failure is provided, the method comprising the following steps:
[0008] Obtain the memory fault light on signal and pass it to the BMC and OS;
[0009] After receiving the fault light on signal, the BMC and OS respectively collect memory fault data.
[0010] As a further solution of the present invention, the method of obtaining the fault light on signal in the memory determines whether the fault light in the memory is on by using a photoresistor.
[0011] As a further solution of the present invention, the method of obtaining the fault light on signal in the memory determines that the fault light in the memory is on by using a photoresistor to obtain the signal.
[0012] As a further solution of the present invention, the collection of BMC memory fault data includes the following steps:
[0013] After receiving the memory fault signal from the photoresistor, the BMC checks whether it has recorded the memory fault information.
[0014] If not, the BMC will trigger the BMC to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor;
[0015] If yes, the photosensitive record file will check whether there is a corresponding memory fault record and time. If yes, the photosensitive record file will no longer save the memory fault record and time. If no, the photosensitive record file will retain the fault record and time in the photosensitive record file.
[0016] The BMC automatically collects logs.
[0017] As a further solution of the present invention, the logs automatically collected by the BMC include: SEL logs, IDLE logs, black box logs and one-key collection logs.
[0018] As a further solution of the present invention, the BMC automatically collecting logs further includes: compressing the collected logs, saving them in a photosensitive recording document in the form of a compressed package, and naming the compressed package as "memory failure time_BMClog".
[0019] As a further solution of the present invention, the collection of OS memory fault data includes the following steps:
[0020] After receiving the memory fault signal from the photoresistor, the OS will also perform a self-check to check whether the BMC has recorded the memory fault information.
[0021] If not, the BMC will trigger the OS to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor;
[0022] If yes, the photosensitive record file will check whether there is a corresponding memory fault record and time. If yes, the photosensitive record file will no longer save the memory fault record and time. If no, the photosensitive record file will retain the fault record and time in the photosensitive record file.
[0023] OS automatically collects logs.
[0024] As a further solution of the present invention, the logs automatically collected by the OS include: message logs, dmesg logs, mce logs and "dmidecode -t memory" memory related information.
[0025] As a further solution of the present invention, the automatic collection of logs by the OS further includes: compressing the collected logs, saving them in a photosensitive recording document in the form of a compressed package, and naming them with the current "memory failure time_OSlog".
[0026] In a second aspect, in another embodiment provided by the present invention, a device for automatically detecting memory failure is provided, the device comprising: a fault light signal acquisition module, a BMC and an OS;
[0027] The fault light signal acquisition module acquires the fault light lighting signal of the memory and transmits the fault light lighting signal of the memory to the BMC and the OS;
[0028] The BMC is used to collect memory fault data after receiving a fault light on signal;
[0029] The OS is used to collect memory fault data after receiving a fault light lighting signal.
[0030] As a further solution of the present invention, the fault light signal acquisition module includes photoresistors arranged on both sides of the memory; the photoresistors are arranged close to the memory fault lights.
[0031] In a third aspect, in another embodiment provided by the present invention, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method for automatically detecting memory failures when loading and executing the computer program.
[0032] In a fourth aspect, in another embodiment provided by the present invention, a storage medium is provided, which stores a computer program, and when the computer program is loaded and executed by a processor, the steps of the method for automatically detecting memory failures are implemented.
[0033] The technical solution provided by the present invention has the following beneficial effects:
[0034] The method, device, computer equipment, and storage medium for automatically detecting memory faults provided by the present invention can obtain the memory fault light lighting signal and save the memory fault light lighting signal in a photosensitive recording document. Memory fault information and logs are recorded therein, which facilitates analysis and query by testers and developers when a memory fault problem occurs and is reviewed. The BMC and OS respectively collect memory fault data to prevent the BMC and OS from failing to record memory faults and thus omitting memory fault information. The memory fault information is transmitted to the BMC and OS respectively. The two links are independent of each other and operate simultaneously. If one link fails, it does not affect the information transmission of the other link. This ensures that the transmission of memory fault information will not be interrupted due to a certain factor.
[0035] These and other aspects of the present invention will become more readily apparent in the following description of the embodiments. It should be understood that the above general description and the following detailed description are merely exemplary and explanatory and are not intended to limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 A flowchart of a method for automatically detecting memory failures according to an embodiment of the present invention;
[0038] Figure 2 The specific process of S20 in the method for automatically detecting memory failure according to an embodiment of the present invention is as follows: Figure 1 ;
[0039] Figure 3 The specific process of S30 in the method for automatically detecting memory failure according to an embodiment of the present invention is as follows: Figure 2 ;
[0040] Figure 4 A structural block diagram of an automatic memory fault detection device according to an embodiment of the present invention;
[0041] Figure 5 A block diagram of the BMC structure in an automatic memory fault detection device according to an embodiment of the present invention;
[0042] Figure 6 This is a block diagram of the OS structure in an automatic memory failure detection device according to an embodiment of the present invention.
[0043] In the figure: fault light signal acquisition module-100, BMC-200, judgment unit 1-201, collection unit 1-202, OS-300, judgment unit 2-301, collection unit 2-302. DETAILED DESCRIPTION
[0044] Various embodiments and / or forms are described below with reference to the accompanying drawings. In the following description, a number of specific details are disclosed for the purpose of explanation in order to understand one or more forms as a whole. However, those skilled in the art will understand that these forms can also be implemented without the specific details. Specific examples of one or more forms will be described in detail in the following description and drawings. However, these forms are only examples, and some of the various methods in the principles of the various forms can be utilized, and the descriptions set forth are intended to include all forms and their equivalents. Specifically, the terms "embodiment", "example", "form", "illustration", etc. used in this specification may be interpreted as any form or design described being better or having advantages over other forms or designs.
[0045] In addition, various forms and features may be embodied by systems including one or more devices, terminals, servers, equipment, components and / or modules. It should be understood and appreciated that various systems may include additional devices, terminals, servers, equipment, components and / or modules and / or may not include all of the devices, terminals, servers, equipment, components, modules, etc. shown in the figures.
[0046] As used in this specification, the terms "computer program," "component," "module," "system," and the like are used interchangeably and refer to computer-related entities, hardware, firmware, software, a combination of software and hardware, or the execution of software. For example, a component can be, but is not limited to, a process executed on a processor, a processor, an object, an execution thread, a program, and / or a computer. For example, it can be an application executed on a computing device and / or all components of a computing device. One or more components can be installed within a processor and / or execution thread. A component can be localized on one computer. A component can also be distributed between two or more computers.
[0047] Furthermore, these components may be executed by various computer-readable media that store various data structures internally. These components may communicate through local and / or remote processing based on signals having one or more data packets (e.g., data sent by a component interacting with other components on a local system or distributed system, and signals transmitted between other systems via a network such as the Internet).
[0048] Hereinafter, regardless of the reference numerals in the drawings, identical or similar components will be assigned the same reference numerals, and repeated descriptions thereof will be omitted. Furthermore, when describing the embodiments disclosed in this specification, if a detailed description of a known technique is judged to obscure the gist of the present invention, such detailed description will be omitted. Furthermore, the drawings are provided solely to facilitate understanding of the embodiments disclosed in this specification, and the technical concepts disclosed in this specification are not limited to the drawings.
[0049] The terms used in this specification are intended to illustrate the embodiments and are not intended to limit the present invention. Unless otherwise specified, the singular in this specification includes the plural. The use of "comprises" and / or "comprising" in this specification does not exclude the presence or addition of one or more other components in addition to the components mentioned.
[0050] Terms such as "first" and "second" can be used to describe various elements or components, but the elements or components are not limited to these terms. These terms are used to distinguish one element or component from other elements or components. Therefore, the first element or component mentioned below can also be the second element or component within the technical concept of the present invention.
[0051] Unless otherwise defined, all terms (including technical and scientific terms) used in this specification shall have the meanings commonly understood by those skilled in the art to which this invention belongs. In addition, terms defined in commonly used dictionaries should not be interpreted as idealistic or excessive unless otherwise clearly defined.
[0052] Furthermore, the term "or" is intended to be inclusive, not exclusive. That is, unless otherwise specified or the context ambiguously states, "X employs A or B" implies one of the naturally occurring alternatives. For example, when X employs A or; X employs B; or X employs A and B, "X employs A or B" can refer to any of the above. Furthermore, it should be understood that the term "and / or" as used in this specification refers to all possible combinations of one or more of the listed items.
[0053] Additionally, the terms "information" and "data" as used in this specification are generally used interchangeably.
[0054] The suffixes "module" and "unit" used in the following description for the constituent elements are given or used interchangeably for the convenience of writing the description, and do not have different meanings or functions.
[0055] Specifically, the embodiments of the present invention are further described below with reference to the accompanying drawings.
[0056] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of a method for automatically detecting memory failures provided by an embodiment of the present invention. Figure 1 As shown, the method for automatically detecting memory failure includes steps S10 to S20.
[0057] S10: Obtain a memory fault light on signal, and transmit the memory fault light on signal to a BMC (Baseboard Management Controller) and an OS (operating system).
[0058] In an embodiment of the present invention, the memory fault light on signal is obtained by determining whether the memory fault light is on through a photoresistor. The memory fault light is red when it is on.
[0059] In the embodiment of the present invention, the photoresistor should be selected to have good sensitivity and be sensitive to red light signals.
[0060] The photoresistor is installed near the memory on both sides of the CPU to be tested. The photoresistor should be set close to the memory fault light so that the photoresistor can better receive the signal when the memory fault light turns on.
[0061] S20 , after receiving the fault light on signal, the BMC (baseboard management controller) and OS (operating system) respectively collect memory fault data.
[0062] The present invention obtains the memory fault light lighting signal and can save the memory fault light lighting signal in a photosensitive recording file. Memory fault information and logs are recorded in the document, making it easy for testers and developers to analyze and query when a memory fault occurs and is reviewed. The BMC and OS separately collect memory fault data to prevent the BMC and OS from failing to record memory faults and thus missing memory fault information. The memory fault information is transmitted to the BMC and OS separately. The two links are independent of each other and operate simultaneously. If one link fails, it does not affect the information transmission of the other link. This ensures that the transmission of memory fault information will not be interrupted due to a certain factor.
[0063] See also Figure 2 , Figure 2 This is the specific process of S20 in the method for automatically detecting memory failure provided by an embodiment of the present invention. Figure 1 The collection of BMC memory fault data includes the following steps:
[0064] S2011: After receiving the memory fault signal from the photoresistor, the BMC will first perform a self-check to check whether the BMC has recorded the memory fault information.
[0065] If not, the BMC will trigger the BMC to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor;
[0066] If there is, the optical log file will check whether there is a corresponding memory failure record and time. If so, the optical log file will no longer save the memory failure record and time. If not, the optical log file will retain the failure record and time in the optical log file. Since the OS also retains memory failure records and time in the optical log file, it is necessary to check whether it is the same information to avoid duplicate records.
[0067] S2012, BMC automatically collects logs.
[0068] The logs automatically collected by S2012 and BMC include: SEL logs, IDLE logs, black box logs and one-key collection logs.
[0069] The BMC automatically collects logs in step S2012, further comprising compressing the collected logs, storing the compressed package in a photosensitive record file, and naming the compressed package "memory failure time_BMClog." The memory failure time in the compressed package name corresponds to the time of the memory failure record in the photosensitive record file.
[0070] See also Figure 3 , Figure 3 This is the specific process of S20 in the method for automatically detecting memory failure provided by an embodiment of the present invention. Figure 2 The collection of OS memory fault data includes the following steps:
[0071] S2021. After receiving the memory fault signal from the photoresistor, the OS will also perform a self-check to check whether the BMC has recorded the memory fault information.
[0072] If not, the BMC will trigger the OS to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor;
[0073] If there is, the photosensitive record file will check whether there is a corresponding memory fault record and time. If so, the photosensitive record file will no longer save the memory fault record and time. If not, the photosensitive record file will retain the fault record and time in the photosensitive record file. Since the BMC also retains the memory fault record and time in the photosensitive record file, it is necessary to check whether it is the same information to avoid duplicate records.
[0074] S2022. The OS automatically collects logs.
[0075] The logs in the S2012, OS automatic collection logs include: message logs, dmesg logs, mce logs and "dmidecode -t memory" memory related information.
[0076] The step S2012 of automatically collecting OS logs further includes compressing the collected logs, storing the compressed package in a photosensitive record file, and naming the compressed package "memory failure time_OSlog." The memory failure time in the compressed package name corresponds to the time of the memory failure record in the photosensitive record file.
[0077] In the present invention, the photoresistor transmits the signal of the memory fault light lighting up to the BMC and OS at the same time, and the two links do not affect each other, thus ensuring that the memory fault information will not rely on only one link. If one link fails to transmit information, the other link will also record the memory fault information. If the BMC fails, the OS will record the memory fault information. If the OS fails, the BMC will record the memory fault information. The memory fault information will be transmitted to the BMC and OS respectively, and the BMC and OS will first perform self-test operations. It can ensure that the BMC and OS information will only report information once for this memory fault, and there will be no duplication. Prevent duplicate information from disrupting the analysis and judgment of memory fault problems by testers and developers. The photosensitive recording document will save the memory fault information and logs, making it convenient for testers and developers to find the required fault information, analyze and query it. It saves time and manpower.
[0078] It should be understood that, although the above is described in a certain order, these steps are not necessarily performed in sequence according to the above order. Unless there is clear explanation in this article, the execution of these steps does not have strict order restriction, and these steps can be performed in other orders. Moreover, a part of the steps of the present embodiment may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.
[0079] In one embodiment, see Figure 4 As shown, an embodiment of the present invention further provides an automatic memory fault detection device, which includes a fault light signal acquisition module 100, a BMC 200 and an OS 300.
[0080] The fault light signal acquisition module 100 acquires the fault light lighting signal from the memory and transmits the fault light lighting signal from the memory to the BMC 200 and the OS 300 .
[0081] In an embodiment of the present invention, the fault light signal acquisition module 100 includes photoresistors arranged on both sides of a memory.
[0082] The photoresistor should be arranged close to the memory fault light so that the photoresistor can better receive the signal that the memory fault light is lighting up.
[0083] The BMC 200 is used to collect memory fault data after receiving a fault light on signal.
[0084] See also Figure 5 , Figure 5 1 is a structural block diagram of a BMC 200 in an automatic memory fault detection device provided by an embodiment of the present invention. The BMC 200 includes a judgment unit 201 and a collection unit 202.
[0085] The judgment unit 1 201 is used for the BMC to perform a self-check after receiving the memory fault signal sent back by the photoresistor to check whether the BMC has recorded the memory fault information just now;
[0086] If not, the BMC will trigger the BMC to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor;
[0087] If there is, the optical log file will check whether there is a corresponding memory failure record and time. If so, the optical log file will no longer save the memory failure record and time. If not, the optical log file will retain the failure record and time in the optical log file. Since the OS also retains memory failure records and time in the optical log file, it is necessary to check whether it is the same information to avoid duplicate records.
[0088] The collecting unit 202 is used to automatically collect logs.
[0089] The logs include: SEL logs, IDLE logs, black box logs and one-click collection logs.
[0090] The collecting unit 1 202 is further configured to compress the collected logs and save them in a photosensitive record file as a compressed package, which is named "memory failure time_BMClog". The memory failure time in the compressed package name corresponds to the time of the memory failure record in the photosensitive record file.
[0091] The OS300 is used to collect the internal fault data after receiving the fault light on signal.
[0092] See also Figure 6 , Figure 6 3 is a structural block diagram of an OS 300 in an automatic memory fault detection device provided by an embodiment of the present invention. The OS 300 includes a second judgment unit 301 and a second collection unit 302.
[0093] The second judgment unit 301 is used for the OS to perform a self-check after receiving the memory fault signal sent back by the photoresistor to check whether the BMC has recorded the memory fault information just now;
[0094] If not, the OS will trigger the BMC to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor;
[0095] If there is, the photosensitive record file will check whether there is a corresponding memory fault record and time. If so, the photosensitive record file will no longer save the memory fault record and time. If not, the photosensitive record file will retain the fault record and time in the photosensitive record file. Since the BMC also retains the memory fault record and time in the photosensitive record file, it is necessary to check whether it is the same information to avoid duplicate records.
[0096] In an embodiment of the present invention, the second collecting unit 302 is used to automatically collect logs.
[0097] In an embodiment of the present invention, the logs collected by the second collecting unit 302 include: message logs, dmesg logs, mce logs and "dmidecode -t memory" memory related information.
[0098] In an embodiment of the present invention, the second collecting unit 302 is further configured to compress the collected logs and save them in a photosensitive record file as a compressed package, which is named "Memory Failure Time_OSlog". The memory failure time in the compressed package name corresponds to the time of the memory failure record in the photosensitive record file.
[0099] The present invention obtains the memory fault light lighting signal and can save the memory fault light lighting signal in a photosensitive recording document. The memory fault information and logs will be recorded therein, which is convenient for testers and developers to analyze and query when a memory fault problem occurs and is reviewed; the BMC and OS collect memory fault data separately to prevent the BMC and OS from not recording the memory fault and missing the memory fault information; the memory fault information is transmitted to the BMC and OS separately. The two links are independent of each other and are carried out simultaneously. If one link fails, it will not affect the information transmission of the other link. It ensures that the transmission of memory fault information will not be interrupted due to a certain factor.
[0100] In one embodiment, a computer device is also provided in an embodiment of the present invention, comprising at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for automatically detecting memory failures, and the processor implements the steps in the above method embodiment when executing the instructions.
[0101] The computer devices include user devices and network devices. User devices include, but are not limited to, computers, smartphones, PDAs, and the like; network devices include, but are not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers. Cloud computing is a type of distributed computing that consists of a group of loosely coupled computers forming a super virtual computer. The computer devices can operate independently to implement the present invention, or they can connect to a network and implement the present invention through interactive operations with other computer devices in the network. The networks in which the computer devices reside include, but are not limited to, the Internet, wide area networks, metropolitan area networks, local area networks, VPN networks, and the like.
[0102] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0103] In one embodiment of the present invention, a storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiment are implemented.
[0104] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes in the above-described method embodiments. Any reference to memory, storage, database, or other media used in the various embodiments provided herein can include at least one of non-volatile and volatile memory.
[0105] Finally, it should be noted that the computer-readable storage medium (e.g., memory) herein may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory. By way of example and not limitation, the non-volatile memory may include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which may act as an external cache memory. By way of example and not limitation, RAM may be obtained in a variety of forms, such as synchronous RAM (DRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The storage devices of the disclosed aspects are intended to include, but are not limited to, these and other suitable types of memory.
[0106] The various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure herein may be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP, and / or any other such configuration.
[0107] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
[0108] The above are exemplary embodiments disclosed in the present invention, but it should be noted that various changes and modifications may be made without departing from the scope of the embodiments disclosed in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. In addition, although the elements disclosed in the embodiments of the present invention may be described or required in individual form, they may also be understood as multiple unless expressly limited to the singular.
[0109] It should be understood that, as used herein, the singular form "a" or "an" is intended to include the plural form as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" refers to any and all possible combinations of one or more of the items listed in association. The serial numbers of the embodiments disclosed in the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0110] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the disclosure of the embodiments of the present invention (including the claims) is limited to these examples. Within the spirit of the embodiments of the present invention, the technical features of the above embodiments or different embodiments may be combined, and there are many other variations of different aspects of the above embodiments of the present invention, which are not provided in detail for the sake of simplicity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the scope of protection of the embodiments of the present invention.
Claims
1. A method for automatically detecting memory failures, characterized in that: The method includes: Obtaining a memory fault light on signal, wherein the obtaining of the memory fault light on signal comprises determining whether the memory fault light is on through a photoresistor, obtaining a signal, and transmitting the memory fault light on signal to the BMC and the OS; After receiving the signal that the fault light is on, the BMC and OS respectively collect memory fault data. After receiving the memory fault signal transmitted by the photoresistor, the BMC checks whether the BMC has recorded the memory fault information just now. If not, the BMC will trigger the BMC to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor; If yes, the photosensitive record file will check whether there is a corresponding memory fault record and time. If yes, the photosensitive record file will no longer save the memory fault record and time. If no, the photosensitive record file will retain the fault record and time in the photosensitive record file. BMC automatically collects logs; After receiving the memory fault signal from the photoresistor, the OS checks whether the BMC has recorded the memory fault information. If not, the BMC will trigger the OS to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor; If yes, the photosensitive record file will check whether there is a corresponding memory fault record and time. If yes, the photosensitive record file will no longer save the memory fault record and time. If no, the photosensitive record file will retain the fault record and time in the photosensitive record file. The OS automatically collects logs; If the BMC fails, the OS will record the memory failure information. If the OS fails, the BMC will record the memory failure information.
2. The method for automatically detecting memory failure according to claim 1, wherein: The logs automatically collected by the BMC include: SEL logs, IDLE logs, black box logs, and one-click collection logs.
3. The method for automatically detecting memory failure according to claim 1, wherein: The BMC automatically collects logs further comprising: compressing the collected logs, saving the compressed logs in a photosensitive record file in the form of a compressed package, and naming the compressed logs as "memory failure time_BMClog".
4. The method for automatically detecting memory failure according to claim 1, wherein: The logs automatically collected by the OS include: message logs, dmesg logs, mce logs, and "dmidecode-tmemory" memory-related information.
5. The method for automatically detecting memory failure according to claim 1, wherein: The automatic collection of logs by the OS further includes: compressing the collected logs, saving them in a photosensitive recording document in the form of a compressed package, and naming them with the current "memory failure time_OSlog".
6. A device for automatically detecting memory failure, characterized in that: The device includes: a fault light signal acquisition module, a BMC and an OS; The fault light signal acquisition module acquires a memory fault light on signal, wherein the memory fault light on signal is acquired by determining whether the memory fault light is on through a photoresistor, and transmits the memory fault light on signal to the BMC and the OS; The BMC is used to collect memory fault data after receiving a fault light on signal; The OS is used to collect memory fault data after receiving a signal that a fault light is on. After the BMC receives the memory fault signal transmitted by the photoresistor, it checks whether the BMC has recorded the memory fault information just now. If not, the BMC will trigger the BMC to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor; If yes, the photosensitive record file will check whether there is a corresponding memory fault record and time. If yes, the photosensitive record file will no longer save the memory fault record and time. If no, the photosensitive record file will retain the fault record and time in the photosensitive record file. BMC automatically collects logs; After receiving the memory fault signal from the photoresistor, the OS checks whether the BMC has recorded the memory fault information. If not, the BMC will trigger the OS to generate a memory alarm information log based on the memory fault signal transmitted by the photoresistor; If yes, the photosensitive record file will check whether there is a corresponding memory fault record and time. If yes, the photosensitive record file will no longer save the memory fault record and time. If no, the photosensitive record file will retain the fault record and time in the photosensitive record file. The OS automatically collects logs; If the BMC fails, the OS will record the memory failure information. If the OS fails, the BMC will record the memory failure information.
7. The automatic memory failure detection device according to claim 6, wherein: The fault light signal acquisition module includes photoresistors arranged on both sides of the memory; the photoresistors should be arranged close to the memory fault light.
Citation Information
Patent Citations
Server information prompting method
CN108599972A
Memory fault recording method and device
CN113961478A