Information collection system, method, device, medium and program product

Through the forced interruption and independent network card transmission mechanism, the problem of information loss caused by DDR power failure is solved, and timely and effective storage and data collection of system crash site information are achieved, supporting fault location.

CN120704937APending Publication Date: 2025-09-26LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511213890.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

DDR is prone to power failure due to abnormal reasons, which may cause system crash and loss of key on-site information, increasing the difficulty of locating system failures and their causes.

Method used

It adopts forced interruption, independent information transmission function of the network card and secondary information storage mechanism. When the detection circuit detects the information storage condition, it outputs a forced interrupt signal to the processor. The processor saves the on-site information to the target storage medium and transmits it to the network card. The network card sends it to the information collection end based on the direct memory access technology.

Benefits of technology

It achieves effective storage of key on-site information of system crashes at the information collection end, provides a data basis for locating target equipment failures and causes, and avoids information transmission failures caused by restart operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704937A_ABST
    Figure CN120704937A_ABST
Patent Text Reader

Abstract

The invention discloses an information collection system, method and device, a medium and a program product in the technical field of computers. According to the invention, an information collection terminal can collect real-time field information in at least one target device; in addition, by means of forced interruption, an independent information transmission function of a network card and secondary storage of the information, effective storage of key site information of system crash at an information collection end is realized, and a data basis can be provided for positioning target equipment faults and reasons.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an information collection system, method, device, medium and program product. Background Art

[0002] Currently, critical on-site information about system crashes is stored in DDR (Double Data Rate), but DDR is prone to power failures due to various abnormal reasons, resulting in the loss of critical on-site information and making it more difficult to locate system failures and their causes.

[0003] Therefore, how to achieve effective storage of key on-site information of system crash is a problem that those skilled in the art need to solve. Summary of the Invention

[0004] In view of this, an object of the present invention is to provide an information collection system, method, device, medium and program product to achieve effective storage of key on-site information of system crashes.

[0005] In the first aspect, the present invention provides an information collection system, comprising: an information collection terminal and at least one target device; the target device comprises: a detection circuit, a processor, a target storage medium and a network card supporting direct memory access technology; the detection circuit is used to: output a forced interrupt signal to the processor when an information saving condition is detected; the processor is used to: if it is confirmed through the status register of the detection circuit that the forced interrupt signal comes from the detection circuit, and it is confirmed through the identification register of the detection circuit that the detection circuit is accessible, then obtain sensor data and a current timestamp from the detection circuit as field information; save the field information to a reserved area in the target storage medium, and transmit the field information to the network card; the network card is used to: send the field information to the information collection terminal based on direct memory access technology; the information collection terminal is used to: store the field information to a storage location corresponding to the target device.

[0006] In the second aspect, the present invention provides an information collection method, which is applied to any target device in an information collection system, the target device including: a detection circuit, a processor, a target storage medium and a network card supporting direct memory access technology; the target device is communicatively connected to the information collection end; the method includes: using the detection circuit to output a forced interrupt signal to the processor when an information saving condition is detected; using the processor to confirm that the forced interrupt signal comes from the detection circuit through the status register of the detection circuit, and confirm that the detection circuit is accessible through the identification register of the detection circuit, then obtaining sensor data and a current timestamp from the detection circuit as field information; saving the field information to a reserved area in the target storage medium, and transmitting the field information to the network card; using the network card based on direct memory access technology to send the field information to the information collection end, so that the information collection end stores the field information in the storage location corresponding to the target device.

[0007] In a third aspect, the present invention provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the aforementioned disclosed information collection method.

[0008] In a fourth aspect, the present invention provides a non-volatile storage medium for storing a computer program, wherein the computer program implements the aforementioned disclosed information collection method when executed by a processor.

[0009] In a fifth aspect, the present invention provides a computer program product, comprising a computer program / instruction, which implements the steps of the aforementioned disclosed information collection method when executed by a processor.

[0010] The beneficial effects of the present invention are as follows: the information collection terminal can collect real-time field information from at least one target device; and, when any target device detects an information storage condition using a detection circuit, it outputs a forced interrupt signal to the processor, so that the processor can stop any current operation and immediately respond to the forced interrupt signal, thereby saving the field information to a reserved area in the target storage medium; in order to achieve timely and effective storage of the field information, the processor also transmits the field information to the network card, so that the network card sends the field information to the information collection terminal based on direct memory access technology; wherein, in the process of the network card transmitting the field information to the information collection terminal, the processor-related programs are decoupled, and the information transmission can be quickly completed through direct memory access technology to avoid the subsequent restart operation of the target device causing the failure of information transmission. This solution uses forced interrupts, the independent information transmission function of the network card, and the secondary storage of information to achieve effective storage of key field information of system crashes at the information collection terminal, and can provide a data basis for locating target device failures and causes.

[0011] Correspondingly, the information collection method, device, medium and program product provided by the present invention also have the above technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0013] Figure 1 This is a schematic diagram of an information collection system disclosed in the present invention; Figure 2 This is a flow chart of an information collection method disclosed in the present invention; Figure 3 This is a schematic diagram of another information collection system disclosed in the present invention; Figure 4 This is a management flow chart for DDR disclosed in the present invention; Figure 5 A flowchart related to configuration for information collection during a system startup process disclosed in the present invention; Figure 6 This is a flow chart of information collection implemented in an interrupt handling program disclosed in the present invention; Figure 7 A flowchart of data transmission and reception preparation between a target device and an information collection server disclosed in the present invention; Figure 8 A server structure diagram provided by the present invention; Figure 9 This is a terminal structure diagram provided by the present invention. DETAILED DESCRIPTION

[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0015] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0016] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0017] Currently, critical system crash scene information is stored in DDR (Double Data Rate) synchronous dynamic random access memory. However, DDR is prone to power failure due to various abnormal reasons, resulting in the loss of critical system crash scene information, making it more difficult to locate system failures and their causes. Therefore, the present invention provides an information collection solution that utilizes forced interrupts, independent information transmission capabilities of network cards, and secondary information storage to effectively store critical system crash scene information at the information collection terminal, providing a data foundation for locating target device failures and their causes.

[0018] See also Figure 1 As shown, an embodiment of the present invention discloses an information collection system, including: an information collection terminal and at least one target device; the target device includes: a detection circuit, a processor, a target storage medium (such as DDR) and a network card supporting direct memory access technology.

[0019] It should be noted that a network card that supports direct memory access technology can be an RDMA network card. The process of sending data using an RDMA network card includes: ① Writing data to a cache (a block of memory requested by the application). ② The network card driver controls the RDMA network card to send the data packet. As can be seen, when using an RDMA network card, even if there are problems with the operating system's network protocol stack (such as code logic errors or memory leaks that prevent cache requests), data can still be successfully sent. Therefore, it can be considered that RDMA network cards have relatively independent data transmission capabilities.

[0020] In this embodiment, the detection circuit is configured to output a forced interrupt signal to the processor upon detecting an information storage condition. The processor is configured to obtain sensor data and a current timestamp from the detection circuit as field information if the detection circuit's status register confirms that the forced interrupt signal originates from the detection circuit and the detection circuit's identification register confirms that the detection circuit is accessible. The processor is configured to save the field information to a reserved area in a target storage medium and transmit the field information to a network card. The sensor data obtained from the detection circuit may include temperature data, voltage data, current data, and the like. The network card is configured to send the field information to an information collection terminal based on direct memory access technology. The direct memory access technology may be RDMA (Remote Direct Memory Access) or DMA (Data Memory Access). The information collection terminal is configured to store the field information in a storage location corresponding to the target device. The target device can then be restarted to recover from the abnormality.

[0021] In one embodiment, a processor includes an interrupt controller; the interrupt controller includes a status register and a flag register; the status register is used to indicate whether a forced interrupt signal from a detection circuit has been received; and the flag register is used to indicate whether the detection circuit is accessible. For example, the status register may use a value of 0 to indicate that a forced interrupt signal from the detection circuit has been received, and a value of 1 to indicate that a forced interrupt signal from the detection circuit has not been received. The flag register may use a value of 0 to indicate that the detection circuit is accessible, and a value of 1 to indicate that the detection circuit is inaccessible.

[0022] In this embodiment, the detection circuit can be implemented based on a complex programmable logic device (CPLD) or a field-programmable gate array (FPGA). This allows the processor to save field information using a hardware interrupt, thereby increasing the priority of this operation. Furthermore, hardware interrupts are immune to software interference, making field information saving more efficient and reliable. In one embodiment, the detection circuit includes a watchdog timer and an interrupt module. The watchdog timer is configured to trigger an information saving condition, thereby triggering the interrupt module, if a preset operation times out. The interrupt module is configured to output a forced interrupt signal to the processor. For example, during normal operation of a target device, a fixed value is periodically written to the watchdog timer. If the write operation times out due to a program freeze, abnormal shutdown, abnormal restart, or other reasons, the preset operation (i.e., the write operation to the watchdog timer) is considered to have timed out, triggering the information saving condition and triggering the interrupt module to generate and issue a forced interrupt signal. Thus, the watchdog timer is configured to determine that the preset operation has timed out if the fixed register value is not written within the predetermined period. The target storage medium also stores the watchdog driver and related data.

[0023] Accordingly, in one embodiment, the detection circuit includes: a power management module; a watchdog is used to: send an alarm notification message to the power management module if the preset operation times out; the power management module is used to: respond to the alarm notification message, perform a power-off operation on the processor and the target storage medium to recover from faults caused by program freeze, abnormal shutdown, abnormal restart, and other reasons.

[0024] It's important to note that the detection circuitry external to the processor not only passively responds to address access signals from the processor but can also proactively send signals to the processor at critical moments (such as when a fault occurs or when certain tasks are completed). These signals, called interrupts, interrupt the processor's ongoing work, requiring it to immediately address the event reported by the detection circuitry. The processor includes an interrupt controller. Programs running on the CPU core (typically drivers within the operating system) can configure the interrupt controller via the CPU's internal bus to enable (also called enable) or disable (also called disable) interrupts from certain devices or modules, selectively passing interrupts to the processor. Upon receiving an interrupt, the processor jumps to a pre-configured address and executes the interrupt handler. In addition to forwarding external interrupts to the processor, the interrupt controller can also forward interrupts from devices connected to certain I / O pins.

[0025] Processors typically have at least one pin that receives emergency interrupt signals. Because these interrupt signals cannot be disabled through software instructions, they are also called non-maskable interrupts (NMIs). Some CPU models can support connecting interrupt signals from external devices or controllers to the processor's non-maskable interrupt pins by configuring an interrupt controller.

[0026] Therefore, in this embodiment, the interrupt module is connected to the non-maskable interrupt pin of the processor to enable transmission of the forced interrupt signal to the processor. In other words, upon receiving the forced interrupt signal, the processor must immediately process the forced interrupt signal and cannot discard or mask it.

[0027] It should be noted that the field information can include information such as the temperature at the time of the fault and program logs. Therefore, the detection circuit can include a temperature detection module; the temperature detection module is used to read the temperature sensor data in the target device through the bus controller and transmit the temperature sensor data to the processor; the processor is used to merge the temperature sensor data into the field information. In addition, the field information can also include the values ​​of various status registers at the time of the interrupt, the call stack of the currently executing process or thread, etc.

[0028] To enable the target storage medium to properly store on-site information, the processor can configure and manage the target storage medium. In one embodiment, the processor is configured to: if the target device is started, detect whether the target storage medium has a self-refresh function; if so, cancel the self-refresh function; detect the available storage space of the target storage medium, select a reserved area in the available storage space, and record the address of the reserved area; and mask the address of the reserved area from the operating system running on the processor. If it is detected that the target storage medium does not have a self-refresh function, initialize the target storage medium, detect the available storage space of the target storage medium, select a reserved area in the available storage space, and record the address of the reserved area; and mask the address of the reserved area from the operating system running on the processor.

[0029] It should be noted that after receiving field information, the information collection terminal should record the current time (adding a timestamp to each received field information from a target device). This facilitates comparison of field information stored for the same target device at different times. This facilitates addressing scenarios where the local time of a problematic target device is incorrect (including complete unavailable time, inaccurate time, or incorrect time recorded in the field information). This is particularly useful in scenarios where multiple target devices experience problems within a relatively close timeframe due to similar causes or mutual impact, such as in distributed computing scenarios. This facilitates analyzing the order in which the problems occurred on each target device, thereby facilitating root cause identification. To facilitate distinguishing and comparing different field information, the information collection terminal stamps the field information with a receipt timestamp and stores the field information with the receipt timestamp in the corresponding storage location for the target device. Accordingly, the information collection terminal is configured to: compare field information corresponding to target devices with different receipt timestamps and record first distinguishing information determined by the comparison; and compare field information from different target devices with the same receipt timestamp and record second distinguishing information determined by the comparison. This first distinguishing information allows for detailed information on failures of the same target device at different times, facilitating lifecycle management and maintenance of the entire target device. The second distinguishing information can be used to determine the fault details of different target devices at the same or similar time, which is beneficial for the management and maintenance of common faults of different target devices.

[0030] In this embodiment, the information collection terminal is capable of collecting real-time field information from at least one target device; and, when any target device detects an information storage condition using a detection circuit, it outputs a forced interrupt signal to the processor, so that the processor can stop any current operation and immediately respond to the forced interrupt signal, thereby saving the field information to a reserved area in the target storage medium; in order to achieve timely and effective storage of the field information, the processor also transmits the field information to the network card, so that the network card sends the field information to the information collection terminal based on direct memory access technology; wherein, in the process of the network card transmitting the field information to the information collection terminal, the processor-related programs are decoupled, and the information transmission can be quickly completed through direct memory access technology to avoid the subsequent restart operation of the target device causing the failure of information transmission. This solution uses forced interrupts, the independent information transmission function of the network card, and the secondary storage of information to achieve effective storage of key field information of system crashes at the information collection terminal, and can provide a data basis for locating target device failures and causes.

[0031] An information collection method provided by an embodiment of the present invention is introduced below. The information collection method described below can be referenced with other embodiments described herein.

[0032] An embodiment of the present invention discloses an information collection method, which is applied to any target device in an information collection system. The target device includes: a detection circuit, a processor, a target storage medium and a network card supporting direct memory access technology; the target device is communicatively connected to an information collection terminal.

[0033] See also Figure 2 As shown, the information collection method includes: S201 : When an information saving condition is detected by a detection circuit, a forced interrupt signal is output to a processor.

[0034] S202. Use the processor to confirm through the status register of the detection circuit that the forced interrupt signal comes from the detection circuit, and confirm through the identification register of the detection circuit that the detection circuit is accessible, then obtain sensor data and the current timestamp from the detection circuit as field information; save the field information to a reserved area in the target storage medium, and transmit the field information to the network card.

[0035] S203: Using the network card based on direct memory access technology, the on-site information is sent to the information collection terminal, so that the information collection terminal stores the on-site information in a storage location corresponding to the target device.

[0036] In one embodiment, the detection circuit includes: a watchdog and an interrupt module; the watchdog is used to: trigger the information saving condition to trigger the interrupt module if the preset operation execution times out; the interrupt module is used to: output a forced interrupt signal to the processor.

[0037] In one embodiment, the detection circuit includes: a power management module; a watchdog is used to: send an alarm notification message to the power management module if a preset operation times out; the power management module is used to: perform a power-off operation on the processor and the target storage medium in response to the alarm notification message.

[0038] In one embodiment, the target storage medium is used to store a watchdog driver and related data.

[0039] In one embodiment, the watchdog is used to: if the fixed register value is not written according to a predetermined period, then confirm that the preset operation has timed out.

[0040] In one embodiment, the interrupt module is connected to a non-maskable interrupt pin of the processor.

[0041] In one embodiment, the detection circuit includes: a temperature detection module; the temperature detection module is used to: read temperature sensor data in the target device through a bus controller; transmit the temperature sensor data to a processor; the processor is used to: merge the temperature sensor data into field information.

[0042] In one embodiment, the processor is used to: if the target device is started, detect whether the target storage medium has a self-refresh function; if so, cancel the self-refresh function; detect the available storage space of the target storage medium, select a reserved area in the available storage space, and record the address of the reserved area.

[0043] In one embodiment, the processor is configured to: if it is detected that the target storage medium does not have a self-refresh function, initialize the target storage medium, detect available storage space of the target storage medium, select a reserved area in the available storage space, and record an address of the reserved area.

[0044] In one embodiment, the information collecting terminal is used to: mark the scene information with a receiving timestamp, and store the scene information marked with the receiving timestamp in a storage location corresponding to the target device.

[0045] In one embodiment, the information collection terminal is configured to: compare on-site information corresponding to target devices with different reception timestamps and record first distinguishing information determined by the comparison; and compare on-site information corresponding to different target devices with the same reception timestamp and record second distinguishing information determined by the comparison. The first distinguishing information allows for checking fault details of the same target device at different times, facilitating lifecycle management and maintenance of the entire target device. The second distinguishing information allows for determining fault details of different target devices at the same or similar times, facilitating management and maintenance of common faults across different target devices.

[0046] In one embodiment, the detection circuit is implemented based on a complex programmable logic device or a field programmable gate array.

[0047] Among them, for more specific working processes of each module and unit in this embodiment, reference can be made to the corresponding contents disclosed in other embodiments, and no further details will be given here.

[0048] It can be seen that this embodiment provides an information collection device, which uses forced interruption, independent information transmission function of the network card and secondary storage of information to achieve effective storage of key on-site information of system crashes at the information collection end, and can provide a data basis for locating target device failures and causes.

[0049] See Figure 3 , Figure 3 An information collection system is provided, including a local computer (i.e., a target device) and an information collection server (i.e., an information collection terminal). The local computer includes: a detection circuit implemented based on CPLD, a processor, a DDR, and an RDMA network card, which increases the possibility of saving sufficient on-site information when the system crashes.

[0050] Figure 3The system shown is suitable for interrupts triggered by system restart, shutdown, etc., and for saving corresponding on-site information. Specifically, an address segment (i.e., a reserved area) is reserved in the host memory, i.e., the DDR. Neither the operating system nor ordinary applications are aware of the existence of this memory segment. The local computer is connected to the information collection server via an RDMA network card. When any problem occurs within the local computer, the program can first save the on-site information to the reserved address segment in the DDR after collecting it, and then back up a copy to the information collection server via the network. The system requires a CPLD chip (an FPGA chip is also acceptable, but considering the cost-effectiveness, a CPLD is generally chosen). The CPLD chip logically implements a watchdog, power management module, timer, I2C controller, timed temperature reading module, and interrupt alarm module.

[0051] Among them, the temperature reading module will periodically (for example, every second) read the temperature sensors inside or next to other major chips in the local computer through the I2C controller to obtain the temperature of each chip. This temperature information will be transmitted to the processor and incorporated into the field information that needs to be saved.

[0052] For the watchdog timer, if the software does not trigger the watchdog timer (i.e., write a fixed value to a watchdog register) within 30 seconds, the watchdog timer will send an interrupt to the CPU. If the software does not trigger the watchdog timer within 1 minute, the watchdog timer will notify the power management module to power off and restart the CPU (power off and then power on).

[0053] The power management module controls the power supply to other chips. Typically, this module automatically powers other chips when the board is powered on. However, the CPU program can also proactively power certain chips on and off (if the CPU is powered off, the CPLD must immediately power it back on automatically, effectively restarting the system). This can be used to save power or to reboot the chips.

[0054] The power management module itself can record operations such as powering on and off the chip, and powering off and restarting the CPU. These operations include: the user pressing the restart or power button; software controlling the power management module to power off and restart the CPU or a specific chip; and the watchdog timer triggering a CPU power off and restart. Each operation is assigned a specific number. The CPLD stores the number of the last such operation in a register. The CPU reads this register to obtain the number and obtain the corresponding operation record. If no record is found, it means that the entire system (including the CPLD) is being powered on for the first time after a power outage. These records are invaluable when investigating the cause of a system reboot.

[0055] The CPLD will send an interrupt alarm to the CPU in the following situations: the temperature of a chip exceeds a threshold (configurable by software) or the watchdog is not kicked within 30 seconds.

[0056] When the system starts, the software will configure the current wall time into the timing module of the CPLD. After that, the timing module will automatically count and save the current time.

[0057] The software configures the interrupt controller so that interrupts from the CPLD are sent to the processor core. If the processor core supports non-maskable interrupts, the software also configures the interrupt controller to route interrupts to the processor core's non-maskable interrupt pin. The processor core runs the interrupt handler, whose code resides in the DDR. This handler is responsible for writing the current state information to a reserved address segment in the DDR, controlling the RDMA network card to copy the current state information to the information collection server, and finally restarting the system.

[0058] See Figure 4 During the system startup process, DDR management includes checking whether the DDR self-refresh is enabled. If so, it cancels the DDR self-refresh; if not, it initializes the DDR. The system then proceeds to the following steps: checking the total DDR size, reserving an address in the DDR, and booting the operating system to determine the address. This process is part of the code for system bootloaders such as BIOS or U-boot.

[0059] Specifically, in system boot programs such as BIOS or U-boot, a portion of the DDR address space is reserved for storing context information through methods such as the device tree and ACPI tables, and only the remaining DDR address space is passed to the operating system. This prevents the operating system from automatically using the reserved address space during startup or normal operation, but engineers can still write programs to actively access this address space.

[0060] See Figure 5 Assuming that processor core 0 handles the interrupt signal set in this embodiment, after the operating system starts and loads a series of drivers, the processor configures the interrupt controller to prepare for receiving non-maskable interrupt signals; sets the address of the interrupt handler to the entry address of core 0's non-maskable interrupt; starts a thread bound to core 0 and kicks the watchdog timer every 10 seconds; creates a queue pair (QP) for the network card's RDMA communication and binds it to the QP of the information collection server; creates an MR (Memory Region, a memory area accessible via RDMA) for the reserved address of the DDR so that the RDMA network card can access this address; and then performs other initialization operations.

[0061] It's important to note that all code and data for the watchdog driver reside in the DDR. As the primary program for storing context information, it can store all necessary context information as long as the DDR and operating system are operating normally (at least, the watchdog driver can still execute). Its initialization process configures the interrupt controller, forwarding interrupts from the CPLD to core 0 for processing. It also sets the address of the interrupt handler in this driver to the interrupt entry point address of core 0, ensuring that core 0 immediately jumps to the interrupt handler upon receiving an interrupt. Furthermore, to support RDMA communication within the interrupt handling process, the initialization process performs preparatory work, such as creating and binding a Query Passive Processor (QP) and creating a Memory Module (MR).

[0062] See Figure 6 After receiving an interrupt, processor core 0 executes the interrupt handler, which includes: reading the interrupt controller's status register; determining whether the interrupt originated from the CPLD; if so, reading the relevant CPLD registers to detect whether it was a high-temperature alarm or a watchdog alarm; then reading the temperature and current time from the CPLD and analyzing the call stack of the currently running thread in processor core 0; saving the temperature, time, and call stack information as context information to a reserved address segment; then determining whether the RDMA network card is available by reading the RDMA network card's identification register; if available, the RDMA network card is informed of the reserved address segment, reads the context information from it, and copies the context information to the information collection server; then sets the DDR to self-refresh mode and reboots the device. This demonstrates that critical context information is saved in the reserved address segment of the DDR and, while the RDMA network card is still accessible, backs up a copy of the context information to the remote information collection server. Even if the computer fails to reboot, the context information can still be retrieved from the server for problem analysis. If the system subsequently experiences a problem and successfully reboots, engineers can initiate an investigation by directly accessing the reserved address segment in the DDR and reading the previously saved context information. Since the system has recovered, all software and hardware modules in the system are now functioning normally. Accessing the reserved address segment in the DDR includes: mapping the reserved address segment in the DDR so that the CPU can read it; reading and displaying the data in the reserved address segment in the DDR; reading and displaying the operation records (restart / shutdown, etc.) in the CPLD power management module. Furthermore, a command option is provided for engineers to delete the original data after collecting sufficient information to prevent it from affecting the next on-site information storage. If the system fails to restart successfully after a problem occurs, engineers can check the on-site information on the information collection server.

[0063] See Figure 7The process of sending on-site data to the remote server includes two steps: RDMA_WRITE and POLL_CQ. The first is to write data to the remote end, and the second is to wait for the completion of the first step. In order to avoid network problems that cause RDMA_WRITE to fail, and thus POLL_CQ to be delayed and the system to be unable to restart, a timeout mechanism can be added to the POLL_CQ process. Because sometimes it may happen that the network card successfully transmits the information to the information collection server, but the restart process is stuck; to achieve a successful restart, a timeout can be set for the first step. If it times out, it will no longer wait for the second step, and will force a restart operation. That is: the network card sets a timer for the process of writing on-site information to the information collection server based on direct memory access technology. If the timer times out, the restart operation is forced to ensure that the target device can work normally after restarting. Figure 7 The data cache requested by the server before registering the MR is the memory address segment where the on-site information will be written in the event of a problem on the current computer. For the server, it can request a large data cache at once and then divide it into segments, each of which is used by different remote computers.

[0064] exist Figure 7 In the example, after the computer drives the kick-dog thread, it first calls ibv_alloc_pd to allocate a PD (Protection Domain) for it, then calls ibv_reg_mr to register an MR for the reserved DDR address, allowing access to that address via RDMA. It then calls ibv_create_cq to create a CQ (Completion Queue) and obtain the Completion Queue Number (CQN). It then calls ibv_create_qp to create a QP (Queue Pair) and obtain the QPN (Queue Pair Number). It then establishes socket communication with the peer to exchange information such as the QPN and GIO. It then calls ibv_modify_qp to inform the hardware of the peer's QPN and set the QP to the ready state. Correspondingly, the server also calls ibv_alloc_pd to allocate a PD, apply for and register an MR, call ibv_create_cq to create a CQ and obtain the corresponding CQN, and call ibv_create_qp to create a QP and obtain the corresponding QPN. After both parties exchange the information required for communication through socket, they each set QP to the ready state to wait for the transmission of field information; during the waiting period, both parties can continue other work processes. Figure 7 In the , GID (Global Identifier) ​​is a global identifier that can be used to identify a device.

[0065] It can be seen that in this embodiment, even if the operating system network protocol stack fails and / or the system cannot be restarted successfully, field information that can be analyzed can still be collected. Among them, the temperature reading is performed by the module in the CPLD, which has high stability and is independent of the CPU. Even if the software fails, it will not affect the temperature acquisition. The power management module can be used to actively control the power supply of other chips, providing power saving and the function of restarting the chip to allow the problematic chip to run again. The watchdog can prevent the system from being stuck for a long time. The power management module can record the process of CPU power outage and restart. The use of non-maskable interrupts ensures that when a problem occurs, even if a normal interrupt is being processed, the processor core will immediately execute the specified interrupt handler.

[0066] The following describes an electronic device provided by an embodiment of the present invention, and the electronic device described below can be referenced with other embodiments described herein. The electronic device in this embodiment can be a target device, a related functional module in the target device, an information collection terminal, etc.

[0067] An embodiment of the present invention discloses an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the method disclosed in any of the above embodiments.

[0068] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: upon detecting an information storage condition, the detection circuit outputs a forced interrupt signal to the processor. The processor confirms, through a status register of the detection circuit, that the forced interrupt signal originates from the detection circuit and confirms, through an identification register of the detection circuit, that the detection circuit is accessible, and then obtains sensor data and a current timestamp from the detection circuit as field information. The field information is stored in a reserved area of ​​a target storage medium and transmitted to a network card. The field information is then sent to an information collection terminal using the network card's direct memory access technology.

[0069] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: storing the on-site information in a storage location corresponding to the target device.

[0070] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: if the preset operation execution times out, an information saving condition is triggered to trigger an interrupt module.

[0071] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: outputting a forced interrupt signal to the processor.

[0072] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: if the preset operation execution times out, an alarm notification message is sent to the power management module.

[0073] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: in response to the alarm notification message, powering off the processor and the target storage medium.

[0074] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: storing a watchdog driver and related data.

[0075] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: if the fixed register value is not written according to the predetermined period, it is determined that the preset operation execution has timed out.

[0076] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: reading the temperature sensor data in the target device through the bus controller; and transmitting the temperature sensor data to the processor.

[0077] In this embodiment, when the processor executes the computer program stored in the memory, the following steps may be specifically implemented: merging the temperature sensor data into the field information.

[0078] In this embodiment, when the processor executes the computer program stored in the memory, the following steps can be specifically implemented: if the target device is started, it detects whether the target storage medium has a self-refresh function; if so, cancels the self-refresh function; detects the available storage space of the target storage medium, selects a reserved area in the available storage space, and records the address of the reserved area.

[0079] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: if it is detected that the target storage medium does not have a self-refresh function, the target storage medium is initialized, the available storage space of the target storage medium is detected, and a reserved area is selected in the available storage space, and the address of the reserved area is recorded.

[0080] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: marking the scene information with a reception timestamp, and storing the scene information marked with the reception timestamp in a storage location corresponding to the target device.

[0081] In this embodiment, when the processor executes the computer program stored in the memory, it can specifically implement the following steps: compare the field information of different receiving timestamps corresponding to the same target device, and record the first difference information determined by the comparison; compare the field information of the same receiving timestamps of different target devices, and record the second difference information determined by the comparison.

[0082] Furthermore, an embodiment of the present invention also provides an electronic device. Figure 8 The server shown can also be Figure 9 The terminal shown. Figure 8 and Figure 9 Each of the diagrams is a structural diagram of an electronic device according to an exemplary embodiment, and the contents in the diagrams should not be considered as limiting the scope of application of the present invention.

[0083] Figure 8 This is a schematic diagram of the structure of a server provided in an embodiment of the present invention. The server may include: at least one processor, at least one memory, a power supply, a communication interface, an input / output interface, and a communication bus. The memory is used to store a computer program, which is loaded and executed by the processor to implement the relevant steps of information collection disclosed in any of the aforementioned embodiments.

[0084] In this embodiment, the power supply is used to provide operating voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices. The communication protocol it follows is any communication protocol that can be applied to the technical solution of the present invention and is not specifically limited here; the input and output interface is used to obtain external input data or output data to the outside world. The specific interface type can be selected according to specific application needs and is not specifically limited here.

[0085] In addition, the memory as a carrier for resource storage can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon include operating system, computer programs and data, etc. The storage method can be temporary storage or permanent storage.

[0086] The operating system is used to manage and control the hardware devices and computer programs on the server, enabling the processor to operate and process data in the memory. It can be Windows Server, NetWare, Unix, Linux, etc. In addition to computer programs capable of performing the information collection method disclosed in any of the aforementioned embodiments, computer programs can also include computer programs capable of performing other specific tasks. Data can include data such as application update information and other data such as application developer information.

[0087] Figure 9 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present invention. The terminal may specifically include but is not limited to a smart phone, a tablet computer, a laptop computer or a desktop computer.

[0088] Generally, the terminal in this embodiment includes: a processor and a memory.

[0089] The processor may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor may be implemented in at least one of the following hardware forms: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). The processor may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content required to be displayed on the display. In some embodiments, the processor may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0090] The memory may include one or more computer non-volatile storage media, which may be non-transitory. The memory may also include high-speed random access memory, and non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory is used to store at least the following computer program, wherein, after the computer program is loaded and executed by the processor, it can implement the relevant steps in the information collection method performed by the terminal side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory may also include an operating system and data, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system may include Windows, Unix, Linux, etc. The data may include but is not limited to update information of the application.

[0091] In some embodiments, the terminal may further include a display screen, an input and output interface, a communication interface, a sensor, a power supply, and a communication bus.

[0092] Those skilled in the art will understand that Figure 9 The structure shown in the figure does not constitute a limitation to the terminal, and may include more or fewer components than shown in the figure.

[0093] A non-volatile storage medium provided in an embodiment of the present invention is introduced below. The non-volatile storage medium described below can be referenced with other embodiments described herein.

[0094] A non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the information collection method disclosed in the aforementioned embodiment. The non-volatile storage medium is a computer-readable non-volatile storage medium that, as a carrier for resource storage, may be a read-only memory, random access memory, magnetic disk, or optical disk. The resources stored thereon include an operating system, computer programs, and data, and the storage method may be either temporary or permanent.

[0095] A computer program product provided by an embodiment of the present invention is introduced below. The computer program product described below can be referenced with other embodiments described herein.

[0096] A computer program product includes a computer program / instruction, which implements the steps of the aforementioned disclosed information collection method when executed by a processor.

[0097] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium, where the non-volatile computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments are implemented.

[0098] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0099] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of non-volatile storage medium known in the art.

[0100] Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only applicable to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. An information collection system, characterized in that: include: An information collection terminal and at least one target device; The target device includes: a detection circuit, a processor, a target storage medium and a network card supporting direct memory access technology; The detection circuit is used to: output a forced interrupt signal to the processor when an information storage condition is detected; The processor is configured to: if it is confirmed through a status register of the detection circuit that the forced interrupt signal comes from the detection circuit and it is confirmed through an identification register of the detection circuit that the detection circuit is accessible, obtain sensor data and a current timestamp from the detection circuit as field information; in response to the forced interrupt signal, save the field information to a reserved area in the target storage medium, and transmit the field information to the network card; The network card is used to: send the field information to the information collection terminal based on the memory direct access technology; The information collecting terminal is used to store the on-site information in a storage location corresponding to the target device.

2. The information collection system according to claim 1, characterized in that The detection circuit includes: a watchdog and an interrupt module; the interrupt module is connected to the non-maskable interrupt pin of the processor; The watchdog is configured to: trigger the information saving condition if the preset operation is timed out, thereby triggering the interrupt module; The interrupt module is configured to output the forced interrupt signal to the processor.

3. The information collection system according to claim 2, characterized in that The detection circuit includes: a power management module; The watchdog is configured to: send an alarm notification message to the power management module if the preset operation execution times out; The power management module is configured to: in response to the alarm notification message, perform a power-off operation on the processor and the target storage medium.

4. The information collection system according to claim 2, characterized in that The target storage medium is used to store the watchdog driver and related data.

5. The information collection system according to claim 2, characterized in that The watchdog is used to: if the fixed register value is not written according to the predetermined period, confirm that the preset operation execution has timed out.

6. The information collection system according to claim 1, characterized in that The processor includes: an interrupt controller; the interrupt controller includes: the status register and the identification register; The status register is used to: mark whether the forced interrupt signal from the detection circuit is received; The identification register is used to mark whether the detection circuit is accessible.

7. The information collection system according to claim 1, characterized in that The detection circuit includes: a temperature detection module; The temperature detection module is used to: read the temperature sensor data in the target device through the bus controller; and transmit the temperature sensor data to the processor; The processor is configured to merge the temperature sensor data into the field information.

8. The information collection system according to claim 1, characterized in that The processor is used to: if the target device is started, detect whether the target storage medium is provided with a self-refresh function; if so, cancel the self-refresh function; detect the available storage space of the target storage medium, select the reserved area in the available storage space, and record the address of the reserved area; and shield the address of the reserved area from the operating system running on the processor.

9. The information collection system according to claim 8, characterized in that: The processor is configured to: if it is detected that the target storage medium does not have a self-refresh function, initialize the target storage medium, detect available storage space of the target storage medium, select the reserved area in the available storage space, and record the address of the reserved area; The address of the reserved area is masked from the operating system running on the processor.

10. The information collection system according to any one of claims 1 to 9, characterized in that: The information collecting terminal is used to mark a receiving timestamp for the on-site information, and store the on-site information marked with the receiving timestamp in a storage location corresponding to the target device.

11. The information collection system according to claim 10, characterized in that: The information collecting terminal is used to: compare the field information of different receiving timestamps corresponding to the target device and record the first difference information determined by the comparison; compare the field information of the same receiving timestamps of different target devices and record the second difference information determined by the comparison.

12. An information collection method, characterized in that: Applicable to any target device in an information collection system, the target device comprising: a detection circuit, a processor, a target storage medium and a network card supporting direct memory access technology; the target device is communicatively connected to an information collection terminal; The method includes: utilizing the detection circuit to output a forced interrupt signal to the processor when an information saving condition is detected; Confirming, by the processor, through a status register of the detection circuit, that the forced interrupt signal comes from the detection circuit and confirming, through an identification register of the detection circuit, that the detection circuit is accessible, then acquiring sensor data and a current timestamp from the detection circuit as field information; saving the field information to a reserved area in the target storage medium in response to the forced interrupt signal, and transmitting the field information to the network card; The network card is used to send the on-site information to the information collection terminal based on the memory direct access technology, so that the information collection terminal stores the on-site information in a storage location corresponding to the target device.

13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to execute the computer program to implement the method according to claim 12.

14. A non-volatile storage medium, characterized in that: Used for storing a computer program, wherein the computer program implements the method according to claim 12 when executed by a processor.

15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method of claim 12 is implemented.

Citation Information

Patent Citations

  • Data transmission method and device and storage medium

    CN117453582A

  • System crash information storage method and device and computer system

    CN118132386A

  • Vehicle fault file distributed uploading method, device, equipment and medium

    CN119854284A

  • Content-aware anomaly detection and diagnosis

    US20180131560A1