Chip initialization method and device, electronic equipment and storage medium

By decoupling kernel registration and hardware initialization operations in a multi-chip system and utilizing asynchronous task queues for concurrent execution, the problem of excessively long initialization time in multi-chip systems is solved, achieving efficient system startup and resource utilization.

CN120848965APending Publication Date: 2025-10-28YIZHU TECH (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510991733.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Multi-chip systems, when non-uniform memory access is disabled or in a single-processor environment, take too long to initialize, resulting in a significant increase in system startup time and low processor resource utilization.

Method used

The operating system performs a kernel registration operation for each chip, establishes device visibility, marks the initialization status as incomplete, and then immediately returns a success status code. It then initiates the registration operation for the next chip without blocking, submits the hardware initialization operation to the asynchronous task queue for concurrent execution, and decouples the kernel registration and hardware initialization operations.

Benefits of technology

It significantly shortens the initialization time of multi-chip systems, improves processor resource utilization and system startup efficiency, and ensures the accuracy of resource allocation and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848965A_ABST
    Figure CN120848965A_ABST
Patent Text Reader

Abstract

The invention provides a chip initialization method and device, electronic equipment and a storage medium. The method comprises the steps that an operating system executes kernel registration operation for each chip, equipment visibility of the operating system to the chips is established, and initialization states of the chips are marked as uncompleted in a global state recording structure; after the kernel registration operation is completed, a success status code is immediately returned to the operating system; the operating system initiates kernel registration operation of the next chip in a non-blocking manner based on the success state code; submitting the hardware initialization operation of each chip to an asynchronous task queue for concurrent execution; and after the hardware initialization operation is completed, updating the initialization state of the corresponding chip in the global state recording structure to be ready. According to the method, the kernel registration operation and the hardware initialization operation are decoupled, so that the operating system immediately triggers the kernel registration operation of the next chip after completing the kernel registration operation, and meanwhile, the time-consuming hardware initialization operation is unloaded to the asynchronous task queue for concurrent execution, so that the total time consumption of a multi-chip system is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a chip initialization method, apparatus, electronic device, and storage medium. Background Technology

[0002] As the complexity of artificial intelligence models grows exponentially, single-chip computing power encounters bottlenecks such as power consumption, area, and memory limitations, leading modern AI computing systems to adopt multi-chip architectures to achieve breakthroughs in computing power. Multi-chip systems require initialization of all chips upon startup.

[0003] In existing technologies, the operating system triggers the chip initialization process by calling a unified initialization function interface. This interface places all initialization operations (such as device registration, resource allocation, firmware loading, and hardware self-test) into a single callback function for synchronous execution. The initialization function returns an operation status flag (0 for success, non-zero for exception). When the kernel receives the return value of the initialization function, it indicates that the entire initialization process has ended.

[0004] When a multi-chip system enables Non-Unified Memory Access (NUMA), the operating system distributes the initialization tasks of different chips to their respective physical processors for concurrent execution based on memory node affinity. The total initialization time is determined by the longest initialization time of a single chip. However, when NUMA is disabled in a multi-chip system or in a single-processor environment, the operating system must execute initialization sequentially. That is, the operating system must wait for the current chip to complete all initialization operations before starting the initialization process of the next chip. The initialization time of a single chip increases with the complexity of the application, and the system startup time accumulates linearly, resulting in a significant increase in system startup time. Summary of the Invention

[0005] This disclosure provides a chip initialization method, apparatus, electronic device, and storage medium that can effectively improve the initialization efficiency of multiple chip devices and shorten system startup time in a multi-chip system environment where Non-Unified Memory Access (NUMA) is disabled or the system is single-processor.

[0006] According to one aspect of this disclosure, a chip initialization method is provided, comprising: the operating system performing a kernel registration operation for each chip, establishing device visibility of the operating system to the chip, and marking the initialization state of the chip as incomplete in a global state record structure; immediately returning a success status code to the operating system after completing the kernel registration operation; the operating system initiating a kernel registration operation for the next chip without blocking based on the success status code; submitting the hardware initialization operations of each chip to an asynchronous task queue for concurrent execution; and updating the initialization state of the corresponding chip in the global state record structure to ready after the hardware initialization operation is completed.

[0007] Optionally, the global state record structure is a shared linked list maintained in the kernel space, and the device node corresponding to each chip contains a device identifier and an initialization state flag.

[0008] Optionally, the kernel registration operation includes: registering the device object with the operating system kernel and requesting a device metadata storage area; filling in device information and associating the device object with the metadata storage area; creating a device node in the global state record structure and setting the chip's initialization state to incomplete.

[0009] Optionally, the hardware initialization operation includes: requesting memory resources and initializing hardware resources; and performing a self-test operation after downloading firmware.

[0010] Optionally, the initialization of chip hardware resources includes configuring memory mapping space, setting computing unit operating parameters, and enabling the interrupt controller.

[0011] Optionally, the asynchronous task queue is implemented in either of the following ways: creating a dedicated kernel thread for each chip to perform hardware initialization operations; or encapsulating the hardware initialization operations into a task package and submitting it to a work queue maintained by the operating system.

[0012] Optionally, the chip initialization method further includes: when a user program requests access to chip resources, querying the status of the corresponding chip in the global status record structure; and allowing resource allocation when the chip status is ready.

[0013] Optionally, the success status code indicates that the kernel registration operation is complete, and is unrelated to the result of the hardware initialization operation.

[0014] According to one aspect of this disclosure, a chip initialization apparatus is provided, comprising: a kernel registration module, configured to allow the operating system to perform a kernel registration operation for each chip, establish device visibility of the operating system to the chip and mark the initialization state as incomplete in a global state record structure, and immediately return a success status code to the operating system after completing the kernel registration operation; a scheduling module, configured to allow the operating system to initiate a kernel registration operation for the next chip without blocking based on the success status code; a hardware initialization module, configured to submit the hardware initialization operations of each chip to an asynchronous task queue for concurrent execution; and a state synchronization module, configured to update the state of the corresponding chip in the global state record structure to ready after the hardware initialization operation is completed.

[0015] Optionally, the hardware initialization module includes: a concurrent execution unit for creating a dedicated kernel thread for each chip to perform hardware initialization operations; or a task dispatch unit for encapsulating the hardware initialization operations into task packages and submitting them to the work queue maintained by the operating system.

[0016] Optionally, the chip initialization device further includes: a resource arbitration module, used to query the status of the corresponding chip in the global status record structure when a user program requests access to chip resources; and to allow resource allocation when the chip status is ready.

[0017] According to one aspect of this disclosure, an electronic device is provided, the electronic device including a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for implementing connection communication between the processor and the memory, wherein the program is executed by the processor to implement the chip initialization method as described above.

[0018] According to one aspect of this disclosure, a computer-readable storage medium is provided that stores one or more programs, which can be executed by one or more processors to implement the chip initialization method described above.

[0019] This disclosure proposes a chip initialization method, apparatus, electronic device, and storage medium. The method includes: the operating system performing a kernel registration operation for each chip, establishing device visibility of the operating system to the chip, and marking the initialization state of the chip as incomplete in the global state record structure; immediately returning a success status code to the operating system after completing the kernel registration operation; the operating system initiating the kernel registration operation of the next chip without blocking based on the success status code; submitting the hardware initialization operations of each chip to an asynchronous task queue for concurrent execution; and updating the initialization state of the corresponding chip in the global state record structure to ready after the hardware initialization operation is completed. This disclosure decouples the kernel registration operation and the hardware initialization operation, enabling the operating system to immediately trigger the kernel registration operation of the next chip after completing the kernel registration operation, while unloading the time-consuming hardware initialization operation to an asynchronous task queue for concurrent execution, avoiding serial waiting, greatly compressing the overall initialization time of the multi-chip system, and thus significantly reducing the overall startup time of the multi-chip system.

[0020] Furthermore, after the kernel registration operation is completed, the system can drive the hardware initialization operations of multiple chips in parallel without occupying kernel scheduling resources. While maintaining the independent initialization resources of each chip, it makes full use of the idle computing power of the multi-core processor, improves the overall load efficiency of the processor, and effectively reduces the waste of processor resources when idle or waiting in queues.

[0021] Furthermore, since kernel registration only involves device visibility establishment and metadata allocation, subsequent asynchronous tasks can perform their own unique firmware download, self-test, and parameter configuration processes for different types of chips, achieving unified management and parallel initialization of heterogeneous computing chips under a single framework.

[0022] Furthermore, by utilizing a global state recording structure to track the status of each chip in real time, it ensures that user programs only allocate resources to ready chips when accessing them, thus avoiding resource conflicts or system anomalies caused by accessing unready chips and improving the reliability and stability of the system.

[0023] Furthermore, the global state tracking mechanism ensures that resource allocation accurately matches the chip's ready state, eliminates invalid polling overhead, and enables the host computing resources to operate at near full load and high efficiency during the system startup phase.

[0024] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0025] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0026] Figure 1 This is a system architecture diagram of the chip initialization method applied in the embodiments of this disclosure;

[0027] Figure 2 This is a main flowchart of a chip initialization method according to an embodiment of the present disclosure;

[0028] Figure 3 This is a flowchart of the main kernel registration operation in an embodiment of this disclosure;

[0029] Figure 4 This is a flowchart of the main hardware initialization operation of an embodiment of this disclosure;

[0030] Figure 5 It is a timing diagram for chip initialization in related technologies;

[0031] Figure 6 This is a timing diagram of chip initialization according to an embodiment of this disclosure;

[0032] Figure 7 This is a schematic diagram of the structure of a chip initialization device according to an embodiment of the present disclosure;

[0033] Figure 8 This is a schematic diagram of the structure of an electronic device proposed in one embodiment of the present disclosure. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.

[0035] Before providing a further detailed description of the embodiments of this disclosure, the terms and concepts used in these embodiments are explained, and they are subject to the following interpretations:

[0036] In related technologies, when a multi-chip system enables Non-Unified Memory Access (NUMA), the operating system distributes the initialization tasks of different chips to their respective physical processors for concurrent execution based on memory node affinity. The total initialization time is determined by the longest initialization time of a single chip. However, when NUMA is disabled in a multi-chip system, the operating system must execute initialization sequentially. That is, the operating system must wait for the current chip to complete all initialization operations before starting the initialization process of the next chip. The initialization time of a single chip increases with the complexity of the business, and the system startup time accumulates linearly, resulting in a significant increase in system startup time. During the period when the host processor is waiting for initialization to complete, it is in an idle state, and the system memory resources cannot be utilized by other computing tasks, leading to insufficient utilization of host computing resources.

[0037] Based on this, this disclosure proposes a chip initialization method, apparatus, electronic device, and storage medium. The method includes: the operating system performing a kernel registration operation for each chip, establishing device visibility of the operating system to the chip, and marking the initialization state of the chip as incomplete in the global state record structure; immediately returning a success status code to the operating system after completing the kernel registration operation; the operating system initiating the kernel registration operation of the next chip without blocking based on the success status code; submitting the hardware initialization operations of each chip to an asynchronous task queue for concurrent execution; and updating the initialization state of the corresponding chip in the global state record structure to ready after the hardware initialization operation is completed. This disclosure decouples the kernel registration operation and the hardware initialization operation, enabling the operating system to immediately trigger the kernel registration operation of the next chip after completing the kernel registration operation, while unloading the time-consuming hardware initialization operation to an asynchronous task queue for concurrent execution, avoiding serial waiting, greatly compressing the overall initialization time of the multi-chip system, and thus significantly reducing the overall startup time of the multi-chip system.

[0038] System architecture description applied in the embodiments of this disclosure

[0039] Figure 1 This is a system architecture diagram of the chip initialization method used in this embodiment of the disclosure, which includes: a host 110 and multiple computing chips 120 (e.g., computing chip 0 to computing chip 3).

[0040] The host 110, as the core processing unit of the computer system, integrates a central processing unit and an operating system, possessing the ability to uniformly schedule and manage the resources of the entire system. The host 110 establishes physical connections with multiple computing chips 120 via a high-speed serial bus (such as PCIe) and constructs a logical control channel to realize command transmission and data interaction. To support direct communication with the computing chips 120, the host is internally configured with a physical interface controller (such as a PCIe controller) for initiating bus transactions and driving device initialization and control operations during operation.

[0041] In addition, the host 110 is equipped with system memory to cache program instructions and intermediate computation data, and accesses the register address space of each computing chip 120 through a memory mapping mechanism, thereby achieving direct control over the underlying resources of the chips. The host 110 not only supports parallel collaborative processing with computing chips of the same type, but can also flexibly adapt to the mixed deployment and collaborative operation of different types of computing chips, thereby improving the system's computing performance and scalability in multiple scenarios.

[0042] The computing chip 120 is a processor suitable for executing deep learning algorithms, specifically including but not limited to a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), and a deep learning processing unit (DPU). Multiple computing chips 120 can be of the same type or different types. It should be noted that the number of computing chips 120 can also be other numbers, and this disclosure does not limit this.

[0043] Overall Implementation of the Chip Initialization Method of the Embodiments of this Disclosure

[0044] This disclosure provides a chip initialization method, applied to a chip initialization device, with reference to... Figure 2 The chip initialization methods include:

[0045] Step S210: The operating system performs a kernel registration operation for each chip, establishes the operating system's device visibility to the chip, and marks the chip's initialization state as incomplete in the global state record structure;

[0046] Step S220: After completing the kernel registration operation, immediately return a success status code to the operating system;

[0047] Step S230: The operating system initiates the kernel registration operation for the next chip without blocking based on the success status code;

[0048] Step S240: Submit the hardware initialization operations of each chip to the asynchronous task queue for concurrent execution;

[0049] Step S250: After the hardware initialization operation is completed, update the initialization state of the corresponding chip in the global state record structure to ready.

[0050] In step S210, during the power-on startup of the multi-chip system, the operating system scans and identifies the computing chip devices connected to the bus (such as PCIe) one by one. For each detected chip, a kernel registration operation is performed to complete the mapping and initialization of the chip from a hardware entity to a software-manageable object. This kernel registration operation is a lightweight kernel registration operation.

[0051] In some embodiments, see Figure 3 The kernel registration operation specifically includes:

[0052] Step S211: Register the device object with the operating system kernel and request a device metadata storage area;

[0053] Step S212: Fill in the device information and associate the device object with the metadata storage area;

[0054] Step S213: Create a device node in the global state record structure and set the chip's initialization state to incomplete.

[0055] In step S211, at the beginning of the chip initialization process, the operating system registers a device object representing the current chip with the kernel through the driver, giving the chip an independent "device identity" in the kernel space. This device object serves as the basic unit for the operating system to identify and manage the chip, and is used for subsequent driver binding and resource scheduling. Simultaneously, the operating system allocates an independent device metadata storage area for the chip to record relevant chip information, including but not limited to device type, physical address, driver binding status, and initialization status.

[0056] In step S212, after completing device object registration and metadata allocation, the operating system driver populates the chip's hardware attributes and status information, such as device vendor ID, bus address, supported operating modes, and whether the driver is loaded, and writes this information to the metadata storage area. Simultaneously, through pointer binding or handle reference, the metadata structure is associated with the device object, enabling the operating system and driver to quickly access the chip's detailed status data and control parameters through the device object.

[0057] In step S213, to achieve unified management of the initialization states of multiple chips, the operating system maintains a globally shared linked list in the kernel space as a state record structure. Whenever a new chip device is registered, the system creates a device node in this linked list. Each device node contains at least two key fields: a device identifier (used to uniquely identify the chip) and an initialization status flag (used to indicate the current chip's initialization progress). In this step, the initialization status is set to "Not Ready," indicating that the current chip has not yet completed hardware initialization and cannot provide services externally. Subsequent processes will perform asynchronous updates and access control based on this status flag.

[0058] Through the above steps, the system successfully completed the entire process of chip physical connection, kernel-level awareness, state initialization, and unified registration, laying the structural foundation and information path for subsequent asynchronous initialization and resource scheduling.

[0059] In step S220, after each chip completes its kernel registration operation, the driver does not need to wait for the chip itself to complete complex hardware initialization operations. Instead, it immediately returns a success status code (e.g., 0) to the operating system indicating "kernel registration successful." This success status code indicates the successful completion of the device object registration operation, meaning the current chip has completed device visibility establishment and meets the basic conditions for entering hardware initialization operations. It is worth noting that this success status code is unrelated to the chip's hardware initialization operation and does not wait for or depend on whether the chip loads firmware or completes self-test. Assuming the driver initialization function for chip device 0 is ai_chip_probe(), and the function returns 0 at the end, it indicates that chip device 0 has completed the kernel registration operation, even though the hardware initialization operation is still executing in the background. This non-blocking return mechanism breaks the dependency relationship in traditional serial initialization, providing real-time feedback signals for the operating system to schedule the initialization of the next chip.

[0060] In step S230, since the driver immediately returns a success status code, which is unrelated to the hardware initialization operation, the operating system does not need to wait for the current chip to complete the entire initialization process and can immediately initiate the kernel registration operation for the next chip. The operating system schedules each chip driver to perform kernel registration operations sequentially, executing the kernel registration operations of multiple chips serially through polling or queue scheduling, thus avoiding the bottleneck of serial waiting. If the system supports multi-core parallel scheduling, the kernel registration operations of multiple chips can also be executed in parallel, further shortening the kernel registration time. For example, after the operating system completes the kernel registration of computing chip 0, it immediately initiates the kernel registration of computing chip 1, and so on, thereby completing the kernel registration operations of all chips in a short time.

[0061] In step S240, after the kernel registration operation is completed, the chip-related hardware initialization operations are submitted as asynchronous tasks to the asynchronous task queue maintained by the system, and are executed in parallel by kernel threads or a work queue mechanism. The asynchronous task queue can be implemented using either a thread-based approach or a work queue approach. Specifically, in the thread-based approach: a dedicated kernel thread is created for each chip to perform hardware initialization operations. Specifically, a separate kernel thread (e.g., kthread_create()) is created for each chip, and this thread is responsible for performing the chip's hardware initialization operations, such as firmware loading, hardware resource configuration, and self-testing.

[0062] Work queue method: Hardware initialization operations (such as work_struct) are encapsulated into task packages and submitted to the work queue maintained by the operating system, and then asynchronously scheduled and executed by the kernel thread pool.

[0063] For example, the operating system creates a kernel thread (init_gup0_thread()) for computing chip 0. This thread sequentially executes tasks such as loading the firmware file, initializing the register mapping area, starting the interrupt controller, and initiating the self-test process. Similarly, the hardware initialization operations of computing chips 1 through 3 run in parallel without interfering with each other. This significantly improves the startup concurrency of the multi-chip system, especially ensuring high initialization efficiency when NUMA is disabled or core resources are limited.

[0064] In some embodiments, see Figure 4 The hardware initialization operation includes:

[0065] Step S241: Allocate memory resources and initialize hardware resources;

[0066] Step S242: After downloading the firmware, perform a self-test.

[0067] In step S241, the thread executing hardware initialization operations for each chip first allocates the memory resources required for the chip's operation and completes the initialization configuration of the chip's underlying hardware structure, such as configuring the memory mapping space, setting the computing unit operating parameters, and enabling the interrupt controller. Specifically, configuring the memory mapping space involves mapping the chip's physical register address space to the kernel virtual address space, facilitating the driver's read / write control of the chip's registers. Setting the computing unit operating parameters involves configuring core parameters, such as tensor core size, execution unit scheduling mode, DMA channel priority, and caching strategy, according to the chip model and firmware requirements. Enabling the interrupt controller involves initializing the interrupt control structure, configuring interrupt routes and priorities, and ensuring the effective operation of the exception notification and task completion callback mechanisms between the chip and the host. For example, computing chip 0 maps the BAR0 and BAR2 regions by reading its configuration space information. BAR0 is a mapped region in the host's physical memory space for computing chip 0, used to store its control registers, status registers, etc. BAR2 is another mapped region in the host's physical memory space for computing chip 0, used to map its DMA buffer, frame buffer, or other large data areas. Mapping BAR0 maps the physical address corresponding to BAR0 to the kernel virtual address space, allowing the chip to be started, its operating mode configured, and its running status queried by reading and writing these registers. Mapping BAR2 maps the physical address corresponding to BAR0 to the kernel virtual address space, allowing access to the DMA buffer, frame buffer, or other large data areas. Another example is setting the worker thread block size of the computing unit to 1024; enabling global interrupts and registering interrupt handlers to respond to hardware errors and computation completion signals.

[0068] In step S242, after initializing the hardware resources, the driver will download the firmware, and then perform a self-test. Specifically, the firmware image file is loaded from a system preset path and transferred to the chip's internal storage unit via DMA or register write. Then, after the firmware boots, a chip self-test is triggered, including but not limited to memory read / write verification, computing unit self-verification, temperature / voltage sensor status check, and communication link confirmation. If the self-test passes, the entire chip initialization process is complete. If an anomaly is detected, an error code will be reported and subsequent tasks will be terminated.

[0069] In step S250, after the hardware initialization operation of each chip is completed, the driver will update the global state record structure maintained in the kernel space. This global state record structure is usually implemented in the form of a linked list, hash table, or array, and is used to uniformly track the initialization progress of all computing chips in the system. Each chip corresponds to a unique device node at the time of registration, which includes a device identifier and initialization status flags.

[0070] Specifically, the operating system uses the unique device identifier created by the chip during the kernel registration phase to search for the device node corresponding to the current chip in the global state record. Then, it updates the initialization status flag of that device node from "incomplete" to "ready". This flag update operation can be implemented through atomic instructions or locking mechanisms to ensure the consistency and reliability of the state in a multi-threaded concurrent environment.

[0071] In some embodiments, the chip initialization method further includes:

[0072] The updated node information will be used for permission checks on subsequent user-level access requests. When an application requests access to a chip resource, the operating system will first check the chip's initialization status flag in the global state structure. Only when the initialization flag is "ready" will resource allocation, task scheduling, and other operations be allowed. If the initialization status flag is still "incomplete" or "failed," the access request will be rejected to prevent the chip from being misused in an unstable or uninitialized state.

[0073] In related technologies, see Figure 5 The operating system initiates the initialization of computing chip 0, waiting for the registration and hardware initialization phases of computing chip 0 to complete before initiating the initialization of computing chip 1. The initialization time for a single chip increases with business complexity, and the linear accumulation of system startup time significantly extends the overall system startup time. However, in this embodiment, see... Figure 6 The operating system initiates the initialization of computing chip 0. Once the kernel registration operation is complete, it returns a success status code. Then, the operating system immediately and non-blockingly initiates the initialization of computing chip 1. The hardware initialization operations of computing chip 0 and computing chip 1 are offloaded to the asynchronous task queue and executed concurrently, avoiding serial waiting and greatly compressing the overall initialization time of the multi-chip system, thereby significantly reducing the overall startup time of the multi-chip system.

[0074] The chip initialization method provided in this disclosure includes: the operating system performing a kernel registration operation for each chip, establishing device visibility of the operating system to the chip, and marking the initialization state of the chip as incomplete in the global state record structure; immediately returning a success status code to the operating system after completing the kernel registration operation; the operating system initiating the kernel registration operation of the next chip without blocking based on the success status code; submitting the hardware initialization operations of each chip to an asynchronous task queue for concurrent execution; and updating the initialization state of the corresponding chip in the global state record structure to ready after the hardware initialization operation is completed. This disclosure decouples the kernel registration operation and the hardware initialization operation, enabling the operating system to trigger the kernel registration operation of the next chip immediately after completing the kernel registration operation, while unloading the time-consuming hardware initialization operation to an asynchronous task queue for concurrent execution, avoiding serial waiting, greatly compressing the overall initialization time of the multi-chip system, and thus significantly reducing the overall startup time of the multi-chip system.

[0075] Furthermore, after the kernel registration operation is completed, the system can drive the hardware initialization operations of multiple chips in parallel without occupying kernel scheduling resources. While maintaining the independent initialization resources of each chip, it makes full use of the idle computing power of the multi-core processor, improves the overall load efficiency of the processor, and effectively reduces the waste of processor resources when idle or waiting in queues.

[0076] Furthermore, since kernel registration only involves device visibility establishment and metadata allocation, subsequent asynchronous tasks can perform their own unique firmware download, self-test, and parameter configuration processes for different types of chips, achieving unified management and parallel initialization of heterogeneous computing chips under a single framework.

[0077] Furthermore, by utilizing a global state recording structure to track the status of each chip in real time, it ensures that user programs only allocate resources to ready chips when accessing them, thus avoiding resource conflicts or system anomalies caused by accessing unready chips and improving the reliability and stability of the system.

[0078] Furthermore, the global state tracking mechanism ensures that resource allocation accurately matches the chip's ready state, eliminates invalid polling overhead, and enables the host computing resources to operate at near full load and high efficiency during the system startup phase.

[0079] Description of apparatus and devices according to embodiments of this disclosure

[0080] See Figure 7 This disclosure also provides a chip initialization device 700, including a kernel registration module 710, a scheduling module 720, a hardware initialization module 730, and a state synchronization module 740. Among them,

[0081] The kernel registration module 710 is used by the operating system to perform kernel registration operations for each chip, establish the operating system's device visibility to the chip, mark the initialization state as incomplete in the global state record structure, and immediately return a success status code to the operating system after completing the kernel registration operation.

[0082] The scheduling module 720 is used by the operating system to initiate the kernel registration operation of the next chip without blocking based on the success status code;

[0083] The hardware initialization module 730 is used to submit the hardware initialization operations of each chip to the asynchronous task queue for concurrent execution;

[0084] The state synchronization module 740 is used to update the state of the corresponding chip in the global state record structure to ready after the hardware initialization operation is completed.

[0085] In some embodiments, the hardware initialization module 730 includes a concurrent execution unit 731 or a task dispatch unit 732, wherein the concurrent execution unit 731 is used to create a dedicated kernel thread for each chip to perform hardware initialization operations; and the task dispatch unit 732 is used to encapsulate the hardware initialization operations into task packages and submit them to the work queue maintained by the operating system.

[0086] In some embodiments, the chip initialization device 700 further includes a resource arbitration module 750, which is used to query the status of the corresponding chip in the global status record structure when a user program requests access to chip resources; and to allow resource allocation when the chip status is ready.

[0087] The chip initialization apparatus 700 disclosed herein is used to execute the chip initialization method as described in the above embodiments. Its specific processing procedure is the same as that of the chip initialization method in the above embodiments, and will not be repeated here.

[0088] This disclosure also provides an electronic device 800, including:

[0089] At least one processor, and,

[0090] A memory that is communicatively connected to at least one processor; wherein,

[0091] The memory stores instructions that are executed by at least one processor to cause the at least one processor to perform the method as described in any of the above embodiments of this application when executing the instructions.

[0092] The following combination Figure 8 The hardware structure of the electronic device is described in detail. The electronic device includes: a processor 810, a memory 820, an input / output interface 830, a communication interface 840, and a bus 850.

[0093] The processor 810 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure.

[0094] The memory 820 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 820 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 820 and is called and executed by the processor 810 using the chip initialization method of the embodiments of this disclosure.

[0095] The input / output interface 830 is used to implement information input and output;

[0096] The communication interface 840 is used to enable communication and interaction between this device and other devices. Communication can be achieved via wired means (e.g., USB, Ethernet cable) or wireless means (e.g., mobile network, Wi-Fi, Bluetooth).

[0097] Bus 850 transmits information between various components of the device (e.g., processor 810, memory 820, input / output interface 830, and communication interface 840);

[0098] The processor 810, memory 820, input / output interface 830 and communication interface 840 are connected to each other within the device via bus 850.

[0099] This application also provides a computer-readable storage medium that stores one or more programs, which can be executed by one or more processors to implement the chip initialization method of the above embodiments, which will not be described again here.

[0100] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0101] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0102] It should be understood that in the description of the embodiments of this disclosure, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0103] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0104] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0105] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0106] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.

[0107] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A chip initialization method, characterized in that, include: The operating system performs kernel registration for each chip, establishes the operating system's device visibility to the chip, and marks the chip's initialization state as incomplete in the global state record structure; Immediately after completing the kernel registration operation, a success status code is returned to the operating system. The operating system initiates the kernel registration operation for the next chip without blocking based on the success status code; Submit the hardware initialization operations of each chip to the asynchronous task queue for concurrent execution; After the hardware initialization operation is completed, update the initialization status of the corresponding chip in the global state record structure to "ready".

2. The chip initialization method according to claim 1, characterized in that, The global state record structure is a shared linked list maintained in the kernel space. Each chip's corresponding device node contains a device identifier and an initialization state flag.

3. The chip initialization method according to claim 1, characterized in that, The kernel registration operation includes: Register the device object with the operating system kernel and request a device metadata storage area; Populate device information and associate device objects with the metadata storage area; Create a device node in the global state record structure and set the chip's initialization state to incomplete.

4. The chip initialization method according to claim 1, characterized in that, The hardware initialization operation includes: Allocate memory resources and initialize hardware resources; After downloading the firmware, perform a self-test.

5. The chip initialization method according to claim 4, characterized in that, The initialization of chip hardware resources includes configuring memory mapping space, setting computing unit operating parameters, and enabling the interrupt controller.

6. The chip initialization method according to claim 1, characterized in that, The asynchronous task queue is implemented in any of the following ways: Create a dedicated kernel thread for each chip to perform hardware initialization operations; or The hardware initialization operation is encapsulated as a task package and submitted to the operating system maintenance work queue.

7. The chip initialization method according to claim 1, characterized in that, Also includes: When a user program requests access to chip resources, it queries the status of the corresponding chip in the global status record structure. Resource allocation is allowed when the chip is in a ready state.

8. The chip initialization method according to claim 1, characterized in that, The success status code indicates that the kernel registration operation is complete and is unrelated to the result of the hardware initialization operation.

9. A chip initialization device, characterized in that, include: The kernel registration module is used by the operating system to perform kernel registration operations for each chip, establish the operating system's device visibility to the chip, mark the initialization state as incomplete in the global state record structure, and immediately return a success status code to the operating system after completing the kernel registration operation. The scheduling module is used by the operating system to initiate the kernel registration operation of the next chip without blocking based on the success status code; The hardware initialization module is used to submit the hardware initialization operations of each chip to the asynchronous task queue for concurrent execution; The state synchronization module is used to update the state of the corresponding chip in the global state record structure to "ready" after the hardware initialization operation is completed.

10. The chip initialization apparatus according to claim 9, characterized in that, The hardware initialization module includes: Concurrent execution units are used to create dedicated kernel threads for each chip to perform hardware initialization operations; or The task dispatch unit is used to encapsulate hardware initialization operations into task packages and submit them to the operating system maintenance work queue.

11. The chip initialization apparatus according to claim 9, characterized in that, Also includes: The resource arbitration module is used to query the status of the corresponding chip in the global status record structure when a user program requests access to chip resources. Resource allocation is allowed when the chip is in a ready state.

12. An electronic device, characterized in that, The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory. The program is executed by the processor to implement the chip initialization method as described in any one of claims 1 to 8.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, which can be executed by one or more processors to implement the chip initialization method as described in any one of claims 1 to 8.