A method and device for simulating and testing a multi-core chip
By establishing a node data structure in the main core and having the slave core poll the data structure, flexible binding and scheduling of slave core and functional programs is achieved, and the problem of inflexible binding of slave core and functional programs in the existing technology is solved, and the efficiency of multi-core chip simulation testing is improved.
Patent Information
- Application Number
- CN202010363204.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-30
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-04-30
AI Technical Summary
In the multi-core chip simulation test, the existing naked-to-turn platform is inflexible in the way of binding from cores to functional programs, resulting in low testing efficiency.
By establishing a node data structure in the main core, including the program entrance of the slave core bitmap and functional programs, polling the node data structure after the kernel is started, and when the corresponding flag bitmap in the slave core bitmap is set, the corresponding functional program is run by the kernel. The master core dispatches the slave core to run different functional programs by setting the slave core bitmap.
It realizes flexible binding and scheduling between slave core and functional programs, improves the efficiency of simulation testing, reduces memory consumption, and reduces the coupling between slave core and functional programs.
Smart Images

Figure CN111695314B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chip simulation testing, and in particular to a multi-core chip simulation testing method and device. Background Art
[0002] During the chip development process, in order to verify the functions of the chip, a large number of simulation tests are usually required. On the simulation platform, due to time and efficiency reasons, it is impossible to run a complete operating system (OS), so a new program is needed to replace the OS as the software platform for chip testing.
[0003] The chip simulation testing platform without an operating system (abbreviated as the bare-metal platform in the industry) is a solution widely used in the industry to simulate and test the logic, functions, parameters, etc. of the chip during the chip design stage. Its characteristics such as refinement and fast startup can meet the verification needs of most chips. In the simulation verification of the chip, the bare-metal platform software is mainly responsible for the initialization of the hardware environment, and the hardware functions are verified by the functional programs. Therefore, the bare-metal platform software and the functional programs are relatively independent. During the verification process, multiple functional programs are often required to verify different functional modules. If all the functional programs are compiled together with the bare-metal platform software, it will cause the bare-metal platform software file to be too large and the coupling between the functional programs and the bare-metal platform software to be too large; if each functional test program is compiled into an image with the bare-metal platform software separately, the simulation needs to be restarted every time the test function is switched, which will reduce the simulation efficiency.
[0004] The bare-metal platform software can run on a simulation device that simulates a real scenario or on an FPGA device. Since there is no operating system in the simulation environment, it does not support the creation and scheduling of processes and can only run in a single-task mode. When simulating and testing a chip with a multi-core processor through the bare-metal platform, it is necessary to bind the functional programs to different cores to achieve parallel operation of the bare-metal platform software and multiple functional programs. The common approach is to record the startup information of each core through a two-dimensional array before the slave core starts, and directly obtain the program entry and parameters from the array after the slave core starts. In this way, it is convenient and fast for the master core to maintain and the slave core to find the program entry information. For a 32-bit processor, each slave core in this solution contains a 4-byte functional program entry address, a 4-byte parameter count, and a 4-byte parameter vector. If the processor has n slave cores, the occupied space size is fixed at 12 * n bytes. It can be seen that when the existing bare-metal platform simulates and tests a multi-core chip, a fixed amount of memory is required to store the startup information, and each slave core can only specify one entry address and one set of parameters, resulting in relatively low test flexibility and efficiency. Summary of the Invention
[0005] In view of this, the present invention provides a method and apparatus for simulating and testing a multi-core chip, which are used to solve the technical problems such as inflexible binding method between slave cores and functional programs and low test efficiency in the chip simulation test environment.
[0006] Based on one aspect of the embodiments of the present invention, the present invention provides a method for simulating and testing a multi-core chip, and the method includes:
[0007] After the master core is started, one or more node data structures for binding and scheduling slave cores are established. The node data structure includes a slave core bitmap and a program entry of a functional program, and the slave core bitmap is used to indicate the slave cores bound to the functional program corresponding to this node data structure;
[0008] After the slave core is started, it polls the node data structure. When it is determined that the flag bit corresponding to this slave core in the slave core bitmap of the node data structure is set, this slave core runs the corresponding functional program based on the program entry of this node data structure.
[0009] Further, the master core schedules the slave core to run the functional program corresponding to the node data structure by setting the slave core bitmap in the node data structure, and each flag bit in the slave core bitmap corresponds to a slave core; after the slave core finishes running the functional program corresponding to the node data structure, it clears the flag bit corresponding to this slave core in the slave core bitmap.
[0010] Further, the method further includes: the master core sets the same boot program entry for all slave cores; after the slave core is started, it starts running from the boot program entry and executes the step of polling the node data structure through the boot program.
[0011] Further, the method further includes: the node data structure further includes a read lock and a write lock; when the master core or a certain slave core writes a lock on a certain node data structure, other cores cannot read or write; when the master core or a certain slave core reads a lock on a certain node data structure, other cores are allowed to read but not allowed to write.
[0012] Further, the one or more node data structures are organized in the form of a linked list or a tree data structure.
[0013] Based on another aspect of the embodiments of the present invention, the present invention provides a multi-core chip simulation test apparatus, and the simulation apparatus is used to simulate the chip hardware timing and logic to test and verify the functions of the multi-core chip. The apparatus includes: a master core, a plurality of slave cores, and a machine-readable storage medium;
[0014] The master core is used to establish one or more node data structures for binding and scheduling slave cores in the machine-readable storage medium. The node data structure includes a slave core bitmap and a program entry of a functional program, and the slave core bitmap is used to indicate the slave cores bound to the functional program corresponding to this node data structure;
[0015] The slave core is used to poll the node data structure. When it is determined that the flag bit corresponding to the slave core in the slave core bitmap of the node data structure is set, the slave core runs the corresponding functional program based on the program entry of the node data structure.
[0016] Furthermore, the master core is also used to set the slave core bitmap in the node data structure, and schedule the slave core to run the functional program corresponding to the node data structure by setting the slave core bitmap in the node data structure;
[0017] The slave core is also used to clear the flag bit in the slave core bitmap corresponding to itself in the node data structure. After the slave core runs the functional program corresponding to the node data structure, it clears the flag bit corresponding to the slave core in the slave core bitmap.
[0018] Furthermore, the machine-readable storage medium also includes a boot program entry; the master core is also used to set the same boot program entry for all slave cores; the slave core is also used to start running from the boot program entry after startup, and perform polling of the node data structure through the boot program.
[0019] Furthermore, the node data structure also includes a read-write mutex lock; the master core and the slave core are also used to lock the node data structure with a write lock when performing a write operation on the node data structure. When a certain node data structure is locked with a write lock, other cores cannot read or write; the master core and the slave core are also used to lock the node data structure with a read lock when performing a read operation on the node data structure. When a certain node data structure is locked with a read lock, other cores are allowed to read but not allowed to perform write operations.
[0020] Furthermore, the master core is also used to organize the one or more node data structures in the form of a linked list or tree data structure.
[0021] The present invention is used for simulating and testing a multi-core chip. The present invention establishes a node data structure for binding and scheduling slave cores by the master core. The node data structure includes a slave core bitmap and a program entry of a functional program. The slave core bitmap indicates the slave cores bound to the functional program corresponding to the node data structure. The slave core polls the node data structure. When it is determined that the flag bit corresponding to the slave core in the slave core bitmap is set, it runs the functional program corresponding to the node data structure. The present invention loads different functional programs through a bare-metal transfer platform, can realize the independent development of the bare-metal transfer platform program and the functional program, and can use one bare-metal transfer platform program to load multiple different functional programs for chip function testing during the simulation process, improving the efficiency of simulation testing. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present invention or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other accompanying drawings can also be obtained according to these accompanying drawings of the embodiments of the present invention.
[0023] Figure 1 Schematic flow diagram of a method for multi-core chip simulation testing provided by an embodiment of the present invention;
[0024] Figure 2 Schematic logical diagram for setting a unified boot program entry for multiple slave cores in an embodiment of the present invention;
[0025] Figure 3 Schematic diagram for implementing the binding of slave cores and functional programs through a node data structure in an embodiment of the present invention;
[0026] Figure 4 Schematic diagram for implementing the master core to schedule and execute different functional programs for slave cores through a node data structure in an embodiment of the present invention;
[0027] Figure 5 Schematic diagram of the device structure of a multi-core chip simulation testing provided by an embodiment of the present invention. Detailed implementation manners
[0028] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, rather than limiting the embodiments of the present invention. The singular forms of "a", "the", and "said" used in the embodiments of the present invention and the claims are also intended to include the plural forms, unless the context clearly indicates otherwise. The term "and / or" used in the present invention refers to any or all possible combinations including one or more of the associated listed items.
[0029] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, in addition, the word "if" may be interpreted as "when" or "while" or "in response to a determination".
[0030] In order to achieve flexible scheduling and binding of the hardware resources and functional programs in a multi-core chip in a bare-metal platform environment, the present invention provides a method for multi-core chip simulation testing. Figure 1Schematic flowchart of a method for multi-core chip simulation testing provided by an embodiment of the present invention. This method is applied to a simulation testing device that simulates the hardware timing and logic of a chip. The simulation testing device can be implemented by an FPGA or by a pure software approach on a device with an operating system. The entire simulation testing device can be regarded as the chip to be simulated and tested, and it can access an external random access memory to read and execute programs. Since the clock of the simulation testing device is very slow, running a complete operating system would take a lot of time. Therefore, the multi-core chip simulation testing method and device provided in this aspect are implemented based on a bare-metal platform, and the chip hardware timing and logic are simulated and tested in an environment without operating system support, that is, the functional program for verifying the chip hardware timing and logic does not have operating system support. Therefore, resources and functional programs in the chip to be simulated and tested cannot be scheduled in the same way as processes and resources are scheduled in an operating system environment. The method includes:
[0031] Step 101. After the main core starts, one or more node data structures for binding and scheduling slave cores are established. The node data structure includes a slave core bitmap and the program entry of the functional program. The slave core bitmap is used to indicate the slave cores bound to the functional program corresponding to this node data structure.
[0032] After a chip with multiple processors (abbreviated as cores) starts, generally, the first powered-on and started processor or core on the chip to be tested is called the main core, and the other later-started cores are called slave cores. After the main core and slave cores start normally, various hardware functions on the chip can be simulated and tested by loading functional software. After the main core starts, it will complete necessary hardware initialization and then wake up the slave cores. The program entry of the functional program to be run by the slave cores is specified by the main core. If a fixed program entry is bound to each slave core respectively, it will not only occupy extra storage space but also make the scheduling of functional programs very inflexible and unable to achieve flexible binding and scheduling of functional programs and slave cores. In the embodiment of the present invention, a node data structure is established for each functional program. The slave cores start executing from the same boot program entry after being woken up by the main core. This boot program will guide the slave cores to start polling the node data structure after startup. When a slave core finds that it is bound to a certain node data structure, it can obtain the corresponding functional program entry from this node data structure and thus run the functional program indicated by this functional program entry.
[0033] The data fields included in the node data structure include but are not limited to the slave core bitmap, functional program entry address, number of parameters, and parameter address. In the present invention, the functional program entry address, number of parameters, and parameter address are combined and abbreviated as the program entry.
[0034] Step 102. After booting from the core, poll the node data structure. When it is determined that the flag bit corresponding to this core in the core bitmap of the node data structure is set, this core runs the corresponding functional program based on the program entry of the node data structure.
[0035] The core bitmap in the node data structure is used by the core to determine whether it is bound to the functional program. The basis for judgment is to check whether the flag bit corresponding to this core is set. If it is set, the corresponding core can read the program entry in the node data structure and run the corresponding functional program based on the read program entry. If it is not set, it means that the corresponding core does not need to run the functional program corresponding to the node data structure.
[0036] In an embodiment of the present invention, in the scenario of simulating and testing a multi-core chip, a unified boot program is set for the cores. The boot program traverses the linked list structure composed of multiple node data structures. Each node data structure is mapped and bound to a functional program. The node data structure includes a core bitmap to reflect the binding relationship between the core and the functional program. When a core determines that it is bound to the corresponding application, it obtains the program entry of the functional program from the node data structure, and thus runs the corresponding functional program. As Figure 2 shown, this embodiment is a logical schematic diagram of setting a unified boot program entry for multiple cores. When the multi-core processor is powered on and starts up, after the main core (such as CPU0) starts first, it will set the same boot program entry for each awakened core, such as CPU1, CPU2,..., CPUn. After being awakened, the core will directly read the program instructions from the boot program entry address and start running.
[0037] Compared with the implementation method of fixedly binding the functional program entry for each core separately, in the case where a multi-core processor chip has many cores and these cores need to run multiple functional programs for simulation testing, the technical solution of this embodiment enables each core to independently determine whether it is bound to the functional program through a unified boot program, which can greatly improve the flexibility of scheduling cores to run different functional programs and can save a lot of storage space.
[0038] First, after the main core starts, it loads the functional program into memory and creates a node data structure for the first functional program. If multiple functional programs need to be loaded into memory, corresponding node data structures need to be created for each functional program respectively. The present invention does not limit the data structure type and organization method of the node data structure, which can be a linked list data structure, an array data structure, or a tree data structure, as long as it can store the information required by the invention and enable the slave core to traverse each node data structure. The entry address and parameters of the functional program are written into the node data structure, and the corresponding positions of the slave cores that need to run this program are set in the slave core bitmap. Multiple slave core corresponding positions can be set simultaneously. The meaning of "setting the bit" is to indicate that the slave core corresponding to this bit needs to run the functional program indicated by the program entry of this node data structure. Please refer to Figure 3 In the Block1 part of Figure 3 , after the main core CPU0 loads the functional program APP1 and creates a node data structure for APP1, it sets the bit corresponding to the slave core CPU that needs to run this functional program in the slave core bitmap, that is, the CPU bitmap. Each bit in the slave core bitmap corresponds to a slave core CPU.
[0039] Secondly, after the slave core starts, it automatically starts running from the boot program. In this embodiment, the function of the boot program is to guide the slave core to traverse the linked list composed of node data structures in sequence, and detect whether the bit corresponding to itself in the slave core bitmap in the node data structure is set in sequence. Please refer to Figure 3 In the Block2 part of Figure 3 , when the slave core CPU1 starts, during the process of traversing the node data structure linked list, it detects whether the bit corresponding to itself in the slave core bitmap in each node data structure is set. For example, CPU1 detects the 1st bit, CPU2 detects the 2nd bit, and so on.
[0040] If the slave core detects that the bit corresponding to itself in the bitmap is set, it obtains the entry address and parameters of the functional program in this node data structure and runs this functional program. Please refer to Figure 3 In the example of the Block3 part of Figure 3 , when the slave core CPU2 detects that the 2nd bit corresponding to itself in the slave core bitmap of the node data structure of APP1 is set, then CPU2 can obtain the program entry address and parameter information of APP1 from this node data structure and start running APP1.
[0041] Preferably, when the bit corresponding to itself in the bitmap of the slave core detected from the core is not set, the slave core continues to check the next node data structure. After all node data structures are detected, the slave core will traverse the linked list composed of node data structures again. The purpose of such cyclic detection is to improve the flexibility of scheduling the slave core to execute the functional program and reduce the coupling between the slave core and the functional program. Such a technical effect is reflected in that regardless of whether the slave core starts first or later, after the slave core starts, it only needs to traverse the linked list composed of node data structures to know which functional programs it should execute. The main core loading the functional program and the slave core running the functional program can be asynchronously processed. Therefore, this structure does not depend on the startup time of the slave core, making it more flexible and less error-prone for users in programming or operation. Please refer to Figure 3 In the Block3 part, when the slave core CPU1 detects that the bit corresponding to itself in the slave core bitmap in the node data structure of APP1 is not set during this round of traversal, CPU1 will traverse the subsequent nodes of the linked list and, after traversing the node data structure of APP1 next time, detect again whether the bit corresponding to itself is set.
[0042] Preferably, the main core can dynamically update, add or delete operations on the linked list as needed. For example, when APP1 is no longer needed to perform simulation tests on all slave core CPUs, the node data structure of APP1 can be deleted from the linked list. When APP1 is no longer needed to test CPU2, the bit corresponding to CPU2 in the slave core bitmap in the node data structure of APP1 can be updated from 1 to 0. When a new functional program needs to run, the main core can create a new node data structure in the linked list, such as Figure 3 In the example of Block4 part, the main CPU0 inserts a new node data structure into the linked list and writes the program entry of the new APP2 into it.
[0043] Preferably, please refer to Figure 4For example, after each functional program is completed by a slave core, an operation opposite to the setting operation is performed on the flag bit corresponding to the slave core in the slave core bitmap in the node data structure of the functional program, such as clearing or emptying the corresponding flag bit. In Block1, CPU1 detects that the slave core bitmap corresponding to its own APP1 is set, so it runs APP1. In Block2, when the slave core CPU1 finishes running APP1, CPU1 clears the flag bit in the slave core bitmap corresponding to its own APP1. In Block3, the master core can read the content of each node data structure when needed, and the master core can also obtain the execution status of the slave core. When the master core CPU0 detects that the flag bit of the slave core CPU1 in the slave core bitmap of APP1 is cleared, the master core can continue to schedule CPU1 to execute other functional programs. For example, the master core sets the flag bit of CPU1 in the slave core bitmap of APP2. In Block4, when the slave core CPU1 traverses to the node data structure of APP2, it will run APP2, and after running, it will also clear the flag bit in the slave core bitmap.
[0044] Preferably, in the embodiments of the present invention, to ensure the synchronization consistency and correctness of the content of the node data structure among all cores, a read-write mutex lock is added to the node data structure. When the master core or a certain slave core locks the write lock for a certain node data structure, other cores cannot read or write; when the master core or a certain slave core locks the read lock for a certain node data structure, other cores are allowed to read but not allowed to write. The write operation on the node data structure includes setting and clearing the setting of the slave core bitmap, writing to the program entry, adding and deleting the entire node data structure. The read-write mutex lock can prevent the processor from reading incorrect data, and can prevent data inconsistency and chaos when multiple cores perform concurrent read and write operations on the same node data structure. The purpose of adding the read-write mutex lock is to achieve the following purposes: when a core performs a write operation on a certain node data structure, it prevents other cores from accessing the node data structure and does not allow other cores to perform read and / or write operations on the node data structure at the same time; when no core performs a write operation on the node data structure, all cores can read the content in the node data structure at the same time; when a core is reading the content of the node data structure, it prevents other cores from performing a write operation on the node data structure.
[0045] Based on the embodiments of the present invention, not only can the memory consumption be reduced but also the efficiency of the simulation can be improved at the same time. For example, when there are n slave cores and k functional program entries that need to be executed, if the three fields of the functional program entry address, the number of parameters and the parameters occupy a total of 12 bytes, and the one-to-one binding of the slave core and the functional application is achieved by the traditional two-dimensional array method, then the occupied memory size is fixed to 12*n, while the memory size occupied by the solution of the present invention is (n / 8+12)*k. When the number of slave cores is large and there are not many functional programs to be executed, the memory consumption can be effectively reduced. Based on the embodiments of the present invention, the slave core can execute multiple functional programs in sequence, and the main core can change the slave core bitmap in the node data structure at any time to re-schedule the slave core, so there is no need to repeatedly restart the multi-core chip simulation test device during the simulation process, thereby improving the simulation efficiency.
[0046] When the chip is designed in the early stage, the hardware timing and logic of the chip are designed by hardware description language (HDL). In order to verify whether the design is correct, the hardware description language needs to be compiled and then run on a simulation device. At this time, the simulation device can process various input and output signals like a real chip, so it can also run programs by accessing random access memory like a real chip, but due to the clock, the operation will be very slow. The master core and slave core mentioned in the present invention refer to the master core and slave core of the chip designed by HDL.
[0047] Figure 5 A schematic diagram of the structure of a multi-core chip simulation test device provided for one embodiment of the present invention. The device 500 includes a multi-core chip simulation test device 520, a machine-readable storage medium 512, and a bus 513. In this embodiment, the multi-core chip simulation test device 520 runs on an operating system, and the functional program bound to the slave core inside the device is used to perform functional verification and testing on the chip, but there is no operating system support inside the device. In another embodiment of the present invention, the multi-core chip simulation test device 520 can also be implemented based on an independent FPGA, and the entire test device can be logically regarded as a multi-core chip under test, and the machine-readable storage medium 512 can also be implemented inside the test device 520.
[0048] The multi-core chip simulation test device 520 includes a master core 510 and multiple slave cores 511 . The master core 510 , the slave cores 511 and a machine-readable storage medium 512 are connected via a bus 513 .
[0049] The master core 510 may establish one or more node data structures in the machine-readable storage medium 512, wherein the node data structure includes a slave core bitmap and a program entry of a functional program, wherein the slave core bitmap is used to indicate a slave core 511 bound to a functional program corresponding to the node data structure.
[0050] The slave core 511 polls the node data structure. When it determines that the flag bit corresponding to the slave core 511 in the slave core bitmap of the node data structure is set, the slave core 511 runs the corresponding functional program based on the program entry of the node data structure.
[0051] The master core 510 is responsible for setting the slave core bitmap in the node data structure, and schedules the slave core 511 to run the functional program corresponding to the node data structure by setting the slave core bitmap in the node data structure. The slave core 511 is also responsible for clearing the flag bit in the slave core bitmap corresponding to itself in the node data structure. When the slave core 511 finishes running the functional program corresponding to the node data structure, it clears the flag bit corresponding to the slave core 511 in the slave core bitmap.
[0052] In an embodiment of the present invention, the machine-readable storage medium 512 includes a boot program entry. The master core 510 is also used to set the same boot program entry for all slave cores 511. The slave core 511 is also used to start running from the boot program entry after startup, and polls the node data structure through the boot program.
[0053] In an embodiment of the present invention, a read-write mutex lock is also set for each node data structure. When the master core 510 and the slave core 511 perform a write operation on the node data structure, a write lock is applied to the node data structure. When a write lock is applied to a certain node data structure, other cores cannot read or write. When the master core 510 and the slave core 511 perform a read operation on the node data structure, a read lock is applied to the node data structure. When a read lock is applied to a certain node data structure, other cores are allowed to read but not allowed to perform a write operation.
[0054] The above are only embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A method for simulating and testing a multi-core chip, characterized in that, The method includes: After the main core is started, one or more node data structures for binding and scheduling slave cores are established. The node data structure includes a slave core bitmap and the program entry of the functional program. The slave core bitmap is used to indicate the slave cores bound to the functional program corresponding to this node data structure; After the slave core is started, it polls the node data structure. When it determines that the flag bit corresponding to this slave core in the slave core bitmap of the node data structure is set, this slave core runs the corresponding functional program based on the program entry of this node data structure. When the flag bit corresponding to this slave core is not set, this slave core does not need to run the functional program corresponding to this node data structure.
2. The method according to claim 1, characterized in that, The main core schedules the slave core to run the functional program corresponding to the node data structure by setting the slave core bitmap in the node data structure. Each flag bit in the slave core bitmap corresponds to a slave core; After the slave core finishes running the functional program corresponding to the node data structure, it clears the flag bit corresponding to this slave core in the slave core bitmap.
3. The method according to claim 1, characterized in that, The method further includes: The main core sets the same boot program entry for all slave cores; After the slave core is started, it starts running from the boot program entry and executes the step of polling the node data structure through the boot program.
4. The method according to claim 1, characterized in that, The method further includes: The node data structure further includes a read lock and a write lock; When the main core or a certain slave core writes a lock on a certain node data structure, other cores cannot read or write; when the main core or a certain slave core reads a lock on a certain node data structure, other cores are allowed to read but not allowed to write.
5. The method according to claim 1, characterized in that, The one or more node data structures are organized in the form of a linked list or a tree data structure.
6. A device for simulating and testing a multi-core chip, characterized in that, The device includes: a main core, multiple slave cores, and a machine-readable storage medium; The main core is used to establish one or more node data structures for binding and scheduling slave cores in the machine-readable storage medium. The node data structure includes a slave core bitmap and the program entry of the functional program. The slave core bitmap is used to indicate the slave cores bound to the functional program corresponding to this node data structure; The slave core is used to poll the node data structure. When it determines that the flag bit corresponding to this slave core in the slave core bitmap of the node data structure is set, the slave core runs the corresponding functional program based on the program entry of this node data structure. When the flag bit corresponding to this slave core is not set, this slave core does not need to run the functional program corresponding to this node data structure.
7. The device according to claim 6, characterized in that, The main core is further used to set the slave core bitmap in the node data structure, and schedules the slave core to run the functional program corresponding to the node data structure by setting the slave core bitmap in the node data structure; The slave core is further used to clear the flag bit in the slave core bitmap corresponding to itself in the node data structure. After the slave core finishes running the functional program corresponding to the node data structure, it clears the flag bit corresponding to this slave core in the slave core bitmap.
8. The device according to claim 6, characterized in that, The machine-readable storage medium further includes a boot program entry; The main core is further used to set the same boot program entry for all slave cores; The slave core is further used to start running from the boot program entry after startup and execute the polling of the node data structure through the boot program.
9. The device according to claim 6, characterized in that, The node data structure further includes a read-write mutex lock; The master core and the slave cores are also used to write a write lock to the node data structure when performing a write operation on the node data structure. When a write lock is set on a certain node data structure, other cores cannot read or write it; The master core and the slave cores are also used to write a read lock to the node data structure when performing a read operation on the node data structure. When a read lock is set on a certain node data structure, other cores are allowed to read but not allowed to perform write operations.
10. The device according to claim 6, characterized in that, The master core is also used to organize the one or more node data structures in the form of a linked list or a tree data structure.