Data processing device, data processing method, and electronic device

CN122837751APending Publication Date: 2026-09-29CHENGDU KAIYUAN COMPUTING ECOLOGICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611329941.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-31
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,目前的写屏障机制实现需要通过软件执行数十条指令,导致执行周期较长,设备开销较大

Benefits of technology

本申请实施例提供一种数据处理装置、数据处理方法及电子设备,该数据处理装置包括处理器核心、运行于处理器核心之上的虚拟机,以及与处理器核心耦合的写屏障加速单元。虚拟机用于响应于目标对象的引用类型字段的赋值操作,驱动处理器核心执行卡表设置指令,该卡表设置指令包括地址信息,地址信息指示目标对象的对象地址。处理器核心用于执行卡表设置指令,以对对象地址执行空指针检查,在检查结果指示对象地址为非空的情况下,触发写屏障加速单元基于对象地址确定卡表中目标对象对应的目标卡片的卡片索引,基于卡片索引将卡表中目标卡片的脏位标识原子地设置为第一标识,该第一标识用于标识目标卡片中存在跨代引用,有效实现写屏障机制。本申请技术方案可以采用单条卡表设置指令代替相关技术中的数十条指令,以利用写屏障加速单元实现写屏障机制,从而有效缩短写屏障的执行周期,降低设备开销。并且,通过在触发写屏障加速单元执行脏位标识设置操作之间,对目标对象的对象地址进行空指针检查,以有效辨别无效的对象地址,从而可以有效避免写屏障加速单元因无效的对象地址出现脏位标识设置操作出错的问题,保证写屏障机制的运行稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837751A_ABST
    Figure CN122837751A_ABST
Patent Text Reader

Abstract

The application provides a data processing device, a data processing method and an electronic equipment, and relates to the technical field of computers. The data processing device comprises a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled with the processor core; the virtual machine drives the processor core to execute a card table setting instruction in response to an assignment operation of a reference type field of a target object, the card table setting instruction comprising address information, the address information indicating an object address of the target object; the processor core executes the card table setting instruction to perform a null pointer check on the object address, and in the case where the check result indicates that the object address is not null, triggers the write barrier acceleration unit to determine a card index of a target card corresponding to the target object in a card table based on the object address, and atomically sets a dirty bit of the target card in the card table to a first identifier based on the card index. The application can effectively shorten the execution period of the write barrier and reduce the device overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to, but is not limited to, the field of computer technology, and in particular to a data processing device, a data processing method, and an electronic device. Background Technology

[0002] Generational Garbage Collection (GCC) is an efficient automatic memory management mechanism in the computer field. It divides objects in memory into young generation and old generation according to the length of their life cycle, and uses different garbage collection algorithms for objects with different life cycles to improve garbage collection efficiency.

[0003] In generational garbage collection, a write barrier mechanism is typically required to record cross-generational object references from the old generation to the young generation, thus providing a basis for determining whether the young generation is alive. However, current write barrier mechanisms require the execution of dozens of instructions in software, resulting in long execution cycles and significant equipment overhead. Summary of the Invention

[0004] This application provides a data processing apparatus, a data processing method, and an electronic device to address, to some extent, the problems of long execution cycles and high equipment costs associated with current write barrier mechanisms.

[0005] The technical solution of this application embodiment is implemented as follows: This application provides a data processing device, which includes: a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core. The virtual machine is used to drive the processor core to execute a card table setting instruction in response to an assignment operation of a reference type field of the target object. The card table setting instruction includes address information, which indicates the object address of the target object. The processor core is used to execute the card table setting instruction to perform a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to a first identifier, which is used to identify that there is a cross-generational reference in the target card.

[0006] This application provides a data processing method applied to a data processing device, the data processing device including a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core; the method includes: In response to the assignment operation of the reference type field of the target object, the virtual machine drives the processor core to execute the card table setting instruction, which includes address information indicating the object address of the target object; The processor core executes the card table setting instruction to perform a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to a first identifier, which is used to identify that there is a cross-generational reference in the target card.

[0007] This application provides an electronic device, which includes any of the data processing devices described in this application.

[0008] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the data processing method provided in this application when executed by a processor.

[0009] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the data processing method provided in this application.

[0010] The embodiments of this application have the following beneficial effects: This application provides a data processing apparatus, a data processing method, and an electronic device. The data processing apparatus includes a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core. The virtual machine, in response to an assignment operation on a reference type field of a target object, drives the processor core to execute a card table setting instruction. This card table setting instruction includes address information indicating the object address of the target object. The processor core executes the card table setting instruction to perform a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to a first identifier. This first identifier is used to identify the existence of cross-generational references in the target card, effectively implementing the write barrier mechanism. This application's technical solution can use a single card table setting instruction to replace dozens of instructions in related technologies, utilizing the write barrier acceleration unit to implement the write barrier mechanism, thereby effectively shortening the execution cycle of the write barrier and reducing equipment overhead. Furthermore, by performing a null pointer check on the object address of the target object before the write barrier acceleration unit executes the dirty bit flag setting operation, invalid object addresses can be effectively identified. This effectively avoids the problem of the write barrier acceleration unit failing to set the dirty bit flag due to invalid object addresses, thus ensuring the operational stability of the write barrier mechanism. Attached Figure Description

[0011] Figure 1 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Figure 1 ; Figure 2 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Figure 2 ; Figure 3 This is a schematic diagram of the structure of a write barrier acceleration unit provided in an embodiment of this application. Figure 1 ; Figure 4 This is a schematic diagram of the structure of a write barrier acceleration unit provided in an embodiment of this application. Figure 2 ; Figure 5 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Figure 3 ; Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Figure 4 ; Figure 7 This is a schematic diagram of the encoding format of a card table setting instruction provided in an embodiment of this application; Figure 8 This is a schematic diagram of instruction pipeline processing of a data processing device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of a write barrier acceleration unit provided in an embodiment of this application. Figure 3 ; Figure 10 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 1 ; Figure 11 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 2 .

[0012] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] GCC is a highly efficient automatic memory management mechanism in the computer field. It divides objects in memory into young generation and old generation according to the length of their life cycle, and uses different garbage collection algorithms for objects with different life cycles to improve garbage collection efficiency.

[0015] In generational garbage collection, a write barrier mechanism is typically introduced. This mechanism inserts extra code after each assignment operation of a reference type field to record cross-generational object references from the old generation to the young generation, thus providing a basis for determining the survival of the young generation. Specifically, memory (or heap memory) is usually divided into multiple fixed-size cards (also called card pages). A card table records the state of each card; it can be an array recording whether cross-generational references exist for each card. The elements of this array are dirty bit flags, indicating whether a cross-generational reference exists within a card. When the application modifies a referenced object from the old generation to the young generation within a card, the write barrier code is triggered. The implementation logic of the write barrier code is as follows: it calculates the card page index of the old generation in the card table, records the physical address of the card page based on the base address of the card table and the card page index, reads the dirty bit flag corresponding to the card based on the physical address of the card page, sets the dirty bit flag of the card to 1, and then writes it back. A dirty bit flag with a value of 1 indicates the existence of a cross-generational reference within the card.

[0016] Clearly, implementing write barrier code requires the application to execute dozens of instructions, including conditional branches, memory accesses, card page indexing, and physical address calculations. This results in a long execution cycle and significant device overhead for write barriers. For example, on x86 / ARM architecture devices, a write barrier consumes approximately 20 to 50 clock cycles. Consequently, for object-intensive applications that frequently assign values ​​to reference type fields (such as applications that frequently update caches or perform batch operations), the overhead of write barriers can account for 10% to 20% of the overall application execution cycle, leading to substantial device costs.

[0017] Please refer to Figure 1 This illustration shows a schematic diagram of a data processing apparatus according to an embodiment of this application. Optionally, the data processing apparatus can be installed in electronic devices such as mobile phones, computers, servers, and wearable devices. For example, the data processing apparatus can be installed on the processor of an electronic device. Figure 1 As shown, the data processing device 10 includes: a processor core 101, a virtual machine running on the processor core 101, and a write barrier acceleration unit 102 coupled to the processor core 101.

[0018] The virtual machine is used to respond to the assignment operation of the reference type field of the target object, and drives the processor core 101 to execute the card table setting instruction. The card table setting instruction includes address information, which indicates the object address of the target object.

[0019] The processor core 101 is used to execute card table setting instructions to perform a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit 102 is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to the first identifier, which is used to identify that there is a cross-generation reference in the target card.

[0020] In some embodiments, the coupling of the write barrier acceleration unit 102 with the processor core 101 emphasizes the control relationship between the processor core 101 and the write barrier acceleration unit. The processor core 101 and the write barrier acceleration unit 102 can be two relatively independent hardware units that can communicate with each other via an interface. Alternatively, the write barrier acceleration unit 102 can be integrated onto the processor core 101. Figure 1 The example illustrates the integration of the write barrier acceleration unit 102 onto the processor core 101.

[0021] In some embodiments, the write barrier acceleration unit 102 can be a dedicated hardware circuit that, triggered by a dedicated hardware instruction such as a single card table setting instruction, determines the card index of the target card corresponding to the target object in the card table based on the object address, and atomically sets the dirty bit identifier of the target card in the card table to the first identifier based on the card index, so as to assist in the implementation of the write barrier mechanism.

[0022] In this embodiment, a null pointer check is used to check whether the object address of the target object is null to determine whether the target address is valid. In some embodiments, under the drive of the processor core 101, the processor core 101 can execute a card table setting instruction to parse the card table setting instruction to obtain the object address of the target object, and then determine whether the object address is 0 to perform a null pointer check on the object address. When the object address is 0 (i.e., object address == 0), indicating that the object address is not null, the processor core 101 triggers the write barrier acceleration unit 102 to determine the card index of the target card corresponding to the target object in the card table based on the object address, and atomically sets the dirty bit identifier of the target card in the card table to the first identifier based on the card index, so as to realize the setting of the dirty bit identifier of the target card corresponding to the target object in the card table. Further optionally, the processor core 101 can also be used to trigger null pointer exception handling when the check result indicates that the object address is null. In some embodiments, null pointer exception handling is used to trigger the virtual machine to throw a null pointer exception (NPE). For example, null pointer exception handling is used to be caught by the virtual machine's exception handler to throw an NPE. For example, processor core 101 can be a processor core based on the RISC-V instruction set framework. The RISC-V instruction set framework includes a standard Machine Trap Vector Table (MTVT) register, which stores the code base address of exception / interrupt handlers. During virtual machine initialization, the code base address of a Null PointerException (NPE) handler can be written to the MTVT register, ensuring that the virtual machine's exception handler correctly throws an NPE based on the code base address in the MTVT register, even when the target object's address is null.

[0023] Obviously, by performing a null pointer check on the object address of the target object before the write barrier acceleration unit 102 executes the dirty bit flag setting operation, invalid object addresses can be effectively identified. This can effectively prevent the write barrier acceleration unit 102 from erroneously executing or failing to execute the dirty bit flag setting operation due to invalid object addresses, thus ensuring the operational stability of the write barrier mechanism.

[0024] In some embodiments, the data processing apparatus may be a device running on the Java language, which is a general-purpose, object-oriented programming language. Correspondingly, alternatively, the virtual machine may refer to the JVM (Java Virtual Machine), a virtual computer that runs Java programs.

[0025] In this embodiment, the assignment operation of the reference type field of the target object refers to the process of storing the reference of an object (i.e., the object address) into the reference type field (a member variable) of the target object. For example, the assignment operation `obj.field = newObj` stores the object reference `newObj` into the reference type field `field` of the target object `obj`. Here, `newObj` corresponds to the young generation, and `obj` corresponds to the old generation.

[0026] In response to an assignment operation on a reference type field of the target object, the virtual machine triggers a write barrier mechanism, driving the processor core 101 to execute a card set instruction. This instruction, when the target object's address is invalid, controls the write barrier acceleration unit 102 to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit flag of the target card in the card table is atomically set to the first flag, thus implementing the dirty bit flag setting operation for the target card corresponding to the target object in the card table. This card set instruction can be referred to as a Card Set for Barrier (CST) instruction for write barriers.

[0027] In some embodiments, such as Figure 2 As shown, the card table is stored in memory 103. Under the trigger of the card table setting instruction executed by the processor core 101, the write barrier acceleration unit 102 performs a write operation on memory 103 to atomically set the dirty bit identifier of the target card corresponding to the target object in the card table in memory 103 to the first identifier.

[0028] Obviously, in this embodiment, the dirty bit flag setting operation of the card table in the write barrier mechanism can be integrated into the write barrier acceleration unit 102. The processor core 101 triggers the execution of the card table setting instruction during the assignment operation of the reference type field, thereby controlling the write barrier acceleration unit 102 to execute the dirty bit flag setting operation. This effectively compresses dozens of instructions into a single card table setting instruction, shortening the execution cycle of the write barrier and reducing device overhead. Furthermore, by performing a null pointer check on the object address of the target object before triggering the write barrier acceleration unit to execute the dirty bit flag setting operation, invalid object addresses can be effectively identified. This effectively avoids errors in the dirty bit flag setting operation caused by invalid object addresses, ensuring the operational stability of the write barrier mechanism.

[0029] In an optional embodiment of this application, the write barrier acceleration unit 102 can directly set the dirty bit identifier of the target card corresponding to the target object in the card table to the first identifier.

[0030] In some embodiments, the card table setup instruction may include address information indicating the object address of the target object. The processor core 101, driven by a virtual machine, can parse the card table setup instruction to obtain the object address of the target object, and perform a null pointer check on the object address. If the check result indicates that the object address is not null, the processor core 101 transmits the object address of the target object to the write barrier acceleration unit 102, and triggers the write barrier acceleration unit 102 to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the processor core 101 atomically sets the dirty bit identifier of the target card to the first identifier.

[0031] In one optional implementation, the write barrier acceleration unit 102 can determine the card index of the target card corresponding to the target object in the card table based on the object address of the target object and the card table mask, and calculate the identifier physical address of the target card based on the card index and the card table base address. Based on the identifier physical address, the dirty bit identifier of the target card is atomically set to the first identifier directly. Specifically, the write barrier acceleration unit 102 can determine the card index of the target card corresponding to the target object in the card table based on the object address of the target object, the card table mask of the card table, and the address shift bits, and calculate the identifier physical address of the target card based on the card index and the card table base address. Based on the identifier physical address, the dirty bit identifier of the target card is atomically set to the first identifier directly.

[0032] The card mask (card_mask) is determined based on the card size in the card list. For example, the card mask can be the difference between the card size and 1, that is, the card mask card_mask satisfies: card_mask = card_size - 1, where card_size represents the card size.

[0033] Furthermore, as mentioned earlier, under the write barrier mechanism, memory (or heap memory) is divided into multiple cards of fixed size. The card table is an array that records whether each card has cross-generational references, and it is stored in memory. Elements in the card table are the dirty bit identifiers of the cards, and the element index (i.e., the card index) is the offset of the element address relative to the card table base address. The element address is the physical address of the identifier, referring to the storage address of the element in memory, and the card table base address refers to the starting address of the card table in memory. Optionally, the card index can be the object address of the target object shifted right by the address shift number and then moduloed, i.e., the card index idx satisfies: idx = (addr >> card_shift) & card_mask. addr represents the object address of the target object; card_shift represents the address shift number; and card_mask represents the card table mask.

[0034] The physical address of a card's identifier refers to the memory address where the dirty bit identifier corresponding to the card is stored. Optionally, the physical address of a card's identifier can be the arithmetic sum of the card index and the card table base address, that is, the physical address of the card's identifier card_addr satisfies: card_addr = card_base + idx, where card_base represents the card table base address.

[0035] In some embodiments, after calculating the identifier physical address of the target card, the write barrier acceleration unit 102 can directly overwrite the dirty bit identifier of the target card stored at the identifier physical address with the first identifier. Alternatively, the write barrier acceleration unit 102 can read the dirty bit identifier of the target card based on the identifier physical address, and write it back after setting the dirty bit identifier to the first identifier. For example, the write barrier acceleration unit 102 can execute a write instruction on the identifier physical address to overwrite the dirty bit identifier of the target card stored at the identifier physical address with the first identifier. In this optional embodiment, the write barrier acceleration unit 102 can directly perform the dirty bit identifier setting operation, which effectively assists in implementing the write barrier mechanism with simple logic and a shorter execution path.

[0036] In optional embodiments of this application, such as Figure 3 As shown, the write barrier acceleration unit 102 includes: a base address parameter register 1021, a mask parameter register 1022, a calculation parameter register 1023, and an acceleration execution module 1024. The base address parameter register 1021 is used to store the card table base address. The mask parameter register 1022 is used to store the card table mask. The calculation parameter register 1023 is used to store the address shift bits.

[0037] The accelerated execution module 1024 is used to read the card table base address stored in the base address parameter register 1021, the card table mask stored in the mask parameter register 1022, and the address shift bits stored in the calculation parameter register 1023.

[0038] The accelerated execution module 1024 is used to determine the card index of the target card corresponding to the target object in the card table based on the object address of the target object, the card table mask and the address shift bits, calculate the identifier physical address of the target card based on the card index and the card table base address, and atomically set the dirty bit identifier of the target card to the first identifier based on the identifier physical address, so as to realize the dirty bit identifier setting of the target card corresponding to the target object in the card table.

[0039] In one optional implementation, the accelerated execution module 1024, triggered by the card table setting instruction executed by the processor core 101, can then read the card table base address stored in the base address parameter register 1021, the card table mask stored in the mask parameter register 1022, and the address shift bits stored in the calculation parameter register 1023, and then perform the dirty bit flag setting operation to set the dirty bit flag of the target card corresponding to the target object in the card table. Alternatively, the accelerated execution module 1024 can also read the card table base address in the base address parameter register 1021, the card table mask in the mask parameter register 1022, and the address shift bits in the calculation parameter register 1023 during its initialization process, so that the dirty bit flag setting operation can be performed subsequently triggered by the card table setting instruction executed by the processor core 101 to set the dirty bit flag of the target card corresponding to the target object in the card table.

[0040] In some embodiments, the virtual machine can configure the card table base address in the base address parameter register 1021, the card table mask in the mask parameter register 1022, and the address shift bits in the calculation parameter register 1023. For example, at startup, the virtual machine can configure the storage location of the card table in memory 103, and card table-related information such as card size to obtain the card table base address and card table mask. Then, it executes register write instructions to write the card table base address to the base address parameter register 1021, the card table mask to the mask parameter register 1022, and the address shift bits to the calculation parameter register 1023, thus completing the configuration of the relevant parameters for the dirty bit flag setting operation.

[0041] In this optional embodiment, since the base address parameter register 1021, mask parameter register 1022, and calculation parameter register 1023 are used to store the relevant parameters for the dirty bit flag setting operation, the relevant parameters can be effectively updated simply by setting the registers without modifying the instruction code, ensuring the convenience of parameter updates. Furthermore, this method can decouple the parameter update operation from the upper-level virtual machine, so that the previous virtual machine does not need to perform parameter adaptation modifications; only the adaptation instruction generation is required.

[0042] In optional embodiments of this application, please continue to refer to Figure 3 The write barrier acceleration unit 102 may further include a dirty bit parameter register 1025. The dirty bit parameter register 1025 stores a first identifier, that is, it stores data that identifies cross-generational references in the card. Correspondingly, optionally, the acceleration execution module 1024 is also used to read the first identifier stored in the dirty bit parameter register 1025. As an example, the acceleration execution module 1024 may read the first identifier stored in the dirty bit parameter register 1025 during its initialization process, so as to perform a dirty bit identifier setting operation subsequently. As another example, the first identifier is 1, and the data stored in the dirty bit parameter register 1025 is 1.

[0043] In some embodiments, the virtual machine can configure a first identifier in the dirty bit parameter register 1025. For example, the virtual machine can execute a register write instruction at startup to write the first identifier to the dirty bit parameter register 1025, thus configuring the relevant parameters for the dirty bit identifier setting operation. Since the dirty bit parameter register 1025 stores the relevant parameters for the dirty bit identifier setting operation—the first identifier—updates data with cross-generational references in the identifiable card can be easily updated by simply setting the register, ensuring convenient parameter updates.

[0044] In some embodiments, such as Figure 4 As shown, the accelerated execution module 1024 includes an index calculation submodule 10241 and an identifier writing submodule 10242. The index calculation submodule 10241, triggered by a card table setting instruction executed by the processor core 101, determines the card index of the target card corresponding to the target object in the card table based on the object address of the target object, the card table mask, and the address shift bits. It then calculates the identifier physical address of the target card based on the card index and the card table base address, and transmits this identifier physical address to the identifier writing submodule 10242. The identifier writing submodule 10242 atomically sets the dirty bit identifier of the target card to a first identifier based on the identifier physical address.

[0045] In optional embodiments of this application, such as Figure 5As shown, the data processing device also includes a cache module 105. The cache module 105 is used to cache at least a portion of the card table in memory. It is easy to understand that the cache module 105 can cache the dirty bit identifiers of at least a portion of the cards in the card table. For example, the cache module 105 caches the dirty bit identifiers of the first Y cards in the card table, where 0 < Y ≤ N, and N is the total number of dirty bit identifiers in the card table, which is also the total number of cards in memory.

[0046] The write barrier acceleration unit 102 is further configured to, when a dirty bit identifier for the target card corresponding to the target object exists in the cache module 105, read the dirty bit identifier of the target card in the cache module 105; if the dirty bit identifier is the second identifier, determine the card index of the target card in the card table based on the object address; and atomically update the dirty bit identifier of the target card in the cache module 105 and memory 103 to the first identifier based on the card index. Conversely, when a dirty bit identifier for the target card corresponding to the target object does not exist in the cache module 105, determine the card index of the target card in the card table based on the object address; read the dirty bit identifier of the target card in memory 103 based on the card index; and if the dirty bit identifier is the second identifier, atomically update the dirty bit identifier of the target card in memory 103 to the first identifier based on the card index. The second identifier is used to indicate that there is no cross-generational reference in the target card.

[0047] In this optional embodiment, the write barrier acceleration unit 102 can determine whether there is a dirty bit identifier for the target card corresponding to the target object in the card table stored in the cache module 105. If there is a dirty bit identifier for the target card in the cache module 105, the dirty bit identifier for the target card in the cache module 105 is read, and it is determined whether the dirty bit identifier is a first identifier, so as to determine whether the dirty bit identifier of the target card accurately indicates that there is a cross-generational reference in the target card, so as to determine whether the card table has marked the existence of a cross-generational reference in the target card. If the dirty bit identifier is a second identifier, indicating that the card table has not marked the existence of a cross-generational reference in the target card, the card index of the target card in the card table is determined based on the object address, and the dirty bit identifier of the target card in the cache module 105 and the memory 103 is atomically updated to the first identifier based on the card index. If the dirty bit identifier of the target card is not found in the cache module 105, the write barrier acceleration unit 102 determines the card index of the target card in the card table based on the object address, reads the dirty bit identifier of the target card in memory 103 based on the card index, and determines whether the dirty bit identifier is the first identifier to determine whether the dirty bit identifier of the target card accurately indicates that there is a cross-generation reference in the target card, so as to determine whether the card table has marked the existence of a cross-generation reference in the target card. If the dirty bit identifier is the second identifier, indicating that the card table has not marked the existence of a cross-generation reference in the target card, the dirty bit identifier of the target card in memory 103 is atomically updated to the first identifier based on the card index.

[0048] It is easy to understand that in some implementations, the write barrier acceleration unit 102 can set the dirty bit flag of the target card to the first flag, indicating that a cross-generational reference exists in the target card in the card table, without updating the dirty bit flag of the target card in memory 103 and cache module 105. Clearly, by setting the cache module 105 to cache at least part of the card table, the write barrier acceleration unit 102 can first determine whether a cross-generational reference exists in the target card in the card table based on the information in the cache module 105. This way, if a cross-generational reference exists in the target card in the card table, there is no need to access the card in memory, effectively reducing memory access. Furthermore, by pre-determining whether the dirty bit flag of the target card is the first flag, update operations on the data in memory 103 and cache module 105 are saved when a cross-generational reference exists in the target card in the card table, avoiding duplicate marking of target cards already marked with cross-generational references, and improving write barrier processing efficiency.

[0049] In some embodiments, the cache module 105 caches at least a portion of the card indexes in the card table, and the dirty bit identifiers of the cards corresponding to the card indexes. Optionally, based on this, the write barrier acceleration unit 102 can first determine the card index of the target card in the card table based on the object address, to determine whether the card index of the target card exists in the cache module 105, so as to determine whether the dirty bit identifier of the target card corresponding to the target object exists in the card table stored in the cache module 105; if the card index of the target card exists in the cache module 105, indicating that the dirty bit identifier of the target card exists in the card table stored in the cache module 105, the dirty bit identifier corresponding to the card index of the target card in the cache module 105 can be read, and then it can be determined whether the dirty bit identifier is the first identifier, so that if the dirty bit identifier is the first identifier, the dirty bit identifier corresponding to the card index of the target card in the cache module 105 is atomically updated to the first identifier, and the identifier physical address of the target card is calculated based on the card index and the card table base address, and the dirty bit identifier of the target card in the card table in memory 103 is atomically updated to the first identifier based on the identifier physical address. If the write barrier acceleration unit 102 does not have a card index for the target card in the cache module 105, indicating that the dirty bit identifier of the target card is not present in the card table stored in the cache module 105, it calculates the physical address of the identifier of the target card based on the card index and the card table base address, and reads the dirty bit identifier of the target card in the card table in memory 103 based on the physical address of the identifier. Then it determines whether the dirty bit identifier is the first identifier. If the dirty bit identifier is the first identifier, it atomically updates the dirty bit identifier of the target card in the card table in memory 103 to the first identifier based on the physical address of the identifier.

[0050] In some implementations, the cache module 105 can employ a fully associative or set-associative structure, comprising at least two entries. Each entry records a card index in the card table and the dirty bit identifier of the card corresponding to that index. For example, the cache module 105 can be a 16-entry fully associative cache structure. Each entry includes a label and data; the label is a card index in the card table, and the data is the dirty bit identifier of the card corresponding to that index.

[0051] In this optional embodiment, by setting a cache module 105 to cache at least a portion of the card table, the write barrier acceleration unit 102 can first determine whether the target card has been marked as having a cross-generational reference based on the information in the cache module 105. This way, if the target card is already marked as having a cross-generational reference in the card table, there is no need to access the card in memory, effectively reducing memory access. Furthermore, by pre-judging whether the dirty bit identifier of the target card is the first identifier, update operations on the data in memory 103 and the cache module 105 are saved if the target card is already marked as having a cross-generational reference in the card table. This avoids duplicate marking of target cards already marked as having cross-generational references, improving write barrier processing efficiency.

[0052] In an optional embodiment of this application, the write barrier acceleration unit 102 is further configured to add the dirty bit identifier of the target card to the card table in the cache module 105 when there is no dirty bit identifier of the target card corresponding to the target object in the cache module 105.

[0053] Optionally, the write barrier acceleration unit 102 can calculate the physical address of the target card's identifier based on the card index and card table base address when the dirty bit identifier of the target card corresponding to the target object does not exist in the cache module 105. It then reads the dirty bit identifier of the target card in the card table in memory 103 based on the physical address, and determines whether the dirty bit identifier is the first identifier. If the dirty bit identifier is the first identifier, it atomically updates the dirty bit identifier of the target card in the card table in memory 103 to the first identifier based on the physical address, and adds the updated dirty bit identifier of the target card to the card table in the cache module 105. If the dirty bit identifier is the first identifier, it directly adds the dirty bit identifier of the target card to the card table in the cache module 105, so that the dirty bit identifier of the target card corresponding to the target object can be directly read from the cache module 105 subsequently.

[0054] In some embodiments, the write barrier acceleration unit 102 can directly add the dirty bit identifier of the target card to the card table in the cache module 105 if the dirty bit identifier of the target card corresponding to the target object does not exist in the cache module 105.

[0055] In some embodiments, the write barrier acceleration unit 102 may add the dirty bit identifier of the target card to the card table in the cache module 105 according to the Least Recently Used (LRU) policy. The LRU policy is used to limit the write barrier acceleration unit 102 to selectively replace the dirty bit identifier of the least recently used card with the dirty bit identifier of the target card.

[0056] For example, the cache module 105 is a fully associative cache structure with 16 items. The write barrier acceleration unit 102 can overwrite the card index of the target card corresponding to the target object and the dirty bit identifier corresponding to the card index into the eviction entry if there is no dirty bit identifier of the target card corresponding to the target object in the cache module 105. The eviction entry refers to the most recently missed entry.

[0057] By updating the cache module 105 according to the LRU policy, the hit rate of the cache module 105 can be guaranteed to a certain extent, while effectively limiting the size of the cache module 105 and ensuring that the cache module 105 occupies a small circuit area.

[0058] In an optional embodiment of this application, the write barrier acceleration unit 102 is further configured to determine the card index of the target card corresponding to the target object in the card table based on the object address, read the dirty bit identifier of the target card corresponding to the target object in the card table based on the card index, and atomically update the dirty bit identifier corresponding to the target object in the card table to the first identifier based on the card index when the dirty bit identifier is the second identifier.

[0059] In this optional embodiment, the write barrier acceleration unit 102 can read the dirty bit identifier of the target card corresponding to the target object in the card table from memory based on the card index, and determine whether the dirty bit identifier is a first identifier to determine whether the target card in the card table has been identified as having a cross-generational reference. If the dirty bit identifier is the first identifier, indicating that the target card in the card table has been identified as having a cross-generational reference, the write barrier acceleration unit 102 determines that the instruction is retired; if the dirty bit identifier is the second identifier, indicating that the target card in the card table has not been identified as having a cross-generational reference, the write barrier acceleration unit 102 updates the dirty bit identifier corresponding to the target object in the card table in memory to the first identifier based on the card index.

[0060] In this optional embodiment, the write barrier acceleration unit 102 can read the dirty bit identifier of the target card page corresponding to the target object in the card table. Only when the dirty bit identifier does not indicate a cross-generational reference in the target card (i.e., it is a second identifier) ​​will the dirty bit identifier of the target card in the card table be updated. When the read dirty bit identifier indicates a cross-generational reference in the target card (i.e., it is a first identifier), there is no need to update the dirty bit identifier of the target card in the card table. This effectively reduces redundant memory write operations when the dirty bit identifier indicates a cross-generational reference in the target card, reducing the number of memory accesses. This effectively reduces bus traffic and power consumption between the write barrier acceleration unit 102 and the memory, improving write barrier processing efficiency. In particular, for object-intensive applications that frequently assign values ​​to reference type fields, reducing the number of memory accesses can more effectively improve the overall write barrier processing efficiency of the device.

[0061] In an optional embodiment of this application, the write barrier acceleration unit 102 is further configured to obtain a path selection identifier, and when the path selection identifier is a fast path identifier, determine the card index of the target card corresponding to the target object in the card table based on the object address, read the dirty bit identifier of the target card corresponding to the target object in the card table based on the card index, and when the dirty bit identifier is a second identifier, atomically update the dirty bit identifier corresponding to the target object in the card table to a first identifier based on the card index.

[0062] The path selection flag indicates whether to perform a skip update operation for the dirty bit flag. Specifically, the path selection flag determines whether to directly and atomically set the dirty bit flag of the target card corresponding to the target object in the card table to the first flag, or to determine whether the dirty bit flag is the first flag, and only atomically update the dirty bit flag corresponding to the target object in the card table to the first flag if the dirty bit flag is the second flag. Optionally, the path selection flag can be a fast path flag or a slow path flag. The fast path flag instructs the write barrier acceleration unit 102 to perform a skip update operation for the dirty bit flag; the slow path flag instructs the write barrier acceleration unit 102 not to perform a skip update operation for the dirty bit flag.

[0063] In this optional embodiment, the write barrier acceleration unit 102 can obtain a path selection identifier and determine whether the path selection identifier is a fast path identifier. If the path selection identifier is not a fast path identifier (e.g., a slow path identifier), the write barrier acceleration unit 102 atomically sets the dirty bit identifier of the target card corresponding to the target object in the card table to the first identifier directly based on the card index of the target card. If the path selection identifier is a fast path identifier, the write barrier acceleration unit 102 reads the dirty bit identifier of the target card corresponding to the target object in the card table based on the card index and determines whether the dirty bit identifier is the first identifier. If the dirty bit identifier is the first identifier, the instruction is retired; if the dirty bit identifier is the second identifier, the dirty bit identifier corresponding to the target object in the card table is updated to the first identifier based on the card index.

[0064] Optionally, the write barrier acceleration unit 102 can also, when the path selection identifier is a fast path identifier, determine whether there is a dirty bit identifier for the target card corresponding to the target object in the card table stored in the cache module 105. If there is a dirty bit identifier for the target card in the cache module 105, the dirty bit identifier for the target card in the cache module 105 is read, and it is determined whether the dirty bit identifier is a first identifier, so as to determine whether the dirty bit identifier of the target card accurately indicates that there is a cross-generational reference in the target card, in order to determine whether the card table has marked the existence of a cross-generational reference in the target card. If the dirty bit identifier is a second identifier, indicating that the card table has not marked the existence of a cross-generational reference in the target card, the card index of the target card in the card table is determined based on the object address, and the dirty bit identifier of the target card in the cache module 105 and memory 103 is atomically updated to the first identifier based on the card index. If the dirty bit identifier of the target card is not found in the cache module 105, the write barrier acceleration unit 102 determines the card index of the target card in the card table based on the object address, reads the dirty bit identifier of the target card in memory 103 based on the card index, and determines whether the dirty bit identifier is the first identifier to determine whether the dirty bit identifier of the target card accurately indicates that there is a cross-generation reference in the target card, so as to determine whether the card table has marked the existence of a cross-generation reference in the target card. If the dirty bit identifier is the second identifier, indicating that the card table has not marked the existence of a cross-generation reference in the target card, the dirty bit identifier of the target card in memory 103 is atomically updated to the first identifier based on the card index.

[0065] In some embodiments, the path selection identifier can be data pre-configured in the write barrier acceleration unit 102. Its fetching can be set based on the actual needs of the device. By flexibly configuring the path selection identifier, the specific implementation scheme of the write barrier mechanism can be flexibly changed, adapting to devices with different configurations and improving the universality of the technical solution.

[0066] In an optional configuration, the write barrier acceleration unit 102 includes a path parameter register and an acceleration execution module. The path parameter register stores a path selection identifier. The acceleration execution module reads the path selection identifier from the path parameter register. If the path selection identifier is a fast path identifier, it determines the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, it reads the dirty bit identifier of the target card corresponding to the target object in the card table. If the dirty bit identifier is a second identifier, it updates the dirty bit identifier of the target card in the card table to a first identifier based on the card index.

[0067] Optionally, the accelerated execution module can obtain the path selection identifier by reading the path selection identifier in the path parameter register under the trigger of the processor core 101. The accelerated execution module determines whether the path selection identifier is a fast path identifier. If the path selection identifier is not a fast path identifier (e.g., a slow path identifier), the accelerated execution module atomically sets the dirty bit identifier of the target card corresponding to the target object in the card table to the first identifier. If the path selection identifier is a fast path identifier, and the dirty bit identifier is the second identifier, the accelerated execution module then atomically updates the dirty bit identifier of the target object in the card table to the first identifier.

[0068] In some embodiments, the virtual machine can configure the path selection identifier in the path parameter register. For example, the virtual machine can execute a register write instruction at startup to write the path selection identifier to the path parameter register, thus configuring the relevant parameters for the dirty bit setting operation. Since the path parameter register stores the relevant parameters for the dirty bit setting operation—the path selection identifier—updating data with cross-generational references in the identifiable card can be achieved simply by setting the register, ensuring the convenience of parameter updates.

[0069] In optional embodiments of this application, please continue to refer to Figure 3 The write barrier acceleration unit 102 includes: a base address parameter register 1021, a mask parameter register 1022, a calculation parameter register 1023, a dirty bit parameter register 1025, a path parameter register 1026, and an acceleration execution module 1024.

[0070] The accelerated execution module 1024 is used to read the path selection identifier in the path parameter register 1026, the card table base address stored in the base address parameter register 1021, the card table mask stored in the mask parameter register 1022, the address shift bits stored in the calculation parameter register 1023, and the first identifier stored in the dirty bit parameter register 1025.

[0071] The accelerated execution module 1024 is further configured to, upon triggering by the processor core 101, determine whether the path selection identifier is a fast path identifier; if the path selection identifier is not a fast path identifier (e.g., a slow path identifier), determine the card index of the target card corresponding to the target object in the card table based on the object address of the target object, the card table mask, and the address shift bits, calculate the identifier physical address of the target card based on the card index and the card table base address, and atomically set the dirty bit identifier of the target card to the first identifier based on the identifier physical address; if the path selection identifier is a fast path identifier, determine the card index of the target card corresponding to the target object in the card table based on the object address of the target object, the card table mask, and the address shift bits, calculate the identifier physical address of the target card based on the card index and the card table base address, read the dirty bit identifier of the target card corresponding to the target object in the card table based on the identifier physical address, and determine whether the dirty bit identifier is the first identifier, so that if the dirty bit identifier is the first identifier, the instruction is retired; if the dirty bit identifier is the second identifier, the dirty bit identifier corresponding to the target object in the card table in memory is atomically updated to the first identifier based on the identifier physical address.

[0072] In some embodiments, please refer to Figure 4 The accelerated execution module 1024 may include an index calculation submodule 10241 and an identifier writing submodule 10242. The index calculation submodule 10241, triggered by a card table setting instruction executed by the processor core 101, determines the card index of the target card corresponding to the target object in the card table based on the object address of the target object, the card table mask, and the address shift bits; calculates the identifier physical address of the target card based on the card index and the card table base address; and transmits the identifier physical address to the identifier writing submodule 10242.

[0073] The identifier writing submodule 10242 is used to read the path selection identifier in the path parameter register and determine whether the path selection identifier is a fast path identifier. If the path selection identifier is not a fast path identifier, the dirty bit identifier of the target card corresponding to the target object in the card table in memory is atomically set to the first identifier based on the identifier physical address. If the path selection identifier is a fast path identifier, the dirty bit identifier of the target card corresponding to the target object in the card table is read and the dirty bit identifier is determined to be the first identifier. If the dirty bit identifier is the first identifier, the instruction retirement is determined. If the dirty bit identifier is the second identifier, the dirty bit identifier of the target object in the card table in memory is atomically updated to the first identifier based on the identifier physical address.

[0074] In some embodiments, the accelerated execution module 1024 may further include a comparator. The comparator is used to compare whether the dirty bit identifier of the target card being read is a first identifier. The identifier writing submodule 10242 is further used to, after reading the dirty bit identifier of the target card corresponding to the target object in the card table, transmit the dirty bit identifier to the comparator, so that the comparator compares whether the dirty bit identifier is consistent with the first identifier, and returns the comparison result to the identifier writing submodule 10242. For example, the identifier writing submodule 10242 may include an internal buffer. After reading the dirty bit identifier of the target card corresponding to the target object in the card table, the identifier writing submodule 10242 may temporarily store the dirty bit identifier in the internal buffer so as to determine whether to update the card table in memory based on the comparison result returned by the comparator.

[0075] By design, the dirty bit identifier of the target card is only updated in the card table when the dirty bit identifier of the target card does not indicate a cross-generational reference (i.e., it is the second identifier). When the read dirty bit identifier indicates a cross-generational reference (i.e., it is the first identifier), the dirty bit identifier update operation is not performed. This effectively reduces redundant memory write operations when the dirty bit identifier indicates a cross-generational reference in the target card, reducing the number of memory accesses. This, in turn, effectively reduces bus traffic and power consumption between the write barrier acceleration unit 102 and memory, improving write barrier processing efficiency. Testing shows that this can effectively reduce memory write operations by approximately 30%.

[0076] In particular, for object-intensive scenarios that frequently assign values ​​to reference type fields (such as high-concurrency write barrier mechanisms), when the dirty bit flag is the first flag, it can reduce the card table update by about 2 clock cycles, thereby reducing the overall overhead of the write barrier mechanism to about 22 clock cycles, effectively improving the overall write barrier processing efficiency of the device.

[0077] In an optional embodiment of this application, the write barrier acceleration unit 102 is further configured to perform an atomic operation to atomically set the dirty bit identifier of the target card corresponding to the target object in the card table to the first identifier.

[0078] In an optional configuration, the write barrier acceleration unit 102 can read the dirty bit identifier of the target card corresponding to the target object in the card table based on the card index of the target card, and ensure that the dirty bit identifier is written back after being set as the first identifier. For example, the write barrier acceleration unit 102 can execute a Load Reserved (LR) instruction to read the dirty bit identifier of the target card corresponding to the target object in the card table from memory based on the card index of the target card, and after setting the dirty bit identifier as the first identifier, execute a Store-Conditional (SC) instruction to write the updated dirty bit identifier back to memory. It then determines whether the SC instruction executed successfully, and if the SC instruction fails, retry writing the updated dirty bit identifier back to memory until the SC instruction executes successfully.

[0079] In another alternative scenario, the write barrier acceleration unit 102 can read the dirty bit identifier of the target card corresponding to the target object in the card table based on the card index of the target card, and ensure that the dirty bit identifier is set to the second identifier before writing it back. For example, the write barrier acceleration unit 102 can execute the LR instruction to read the dirty bit identifier of the target card corresponding to the target object in the card table from memory based on the card index of the target card, and set the dirty bit identifier to the first identifier if it is set to the second identifier; then, it executes the SC instruction to write the updated dirty bit identifier back to memory, and determines whether the SC instruction executed successfully. If the SC instruction fails, it retryes writing the updated dirty bit identifier back to memory until the SC instruction executes successfully. The LR instruction, in conjunction with the SC instruction, implements a "load-hold" read-modify-write mechanism, completing atomic operations and ensuring accurate updating of the dirty bit identifier corresponding to the target card page in the card table.

[0080] In this optional embodiment, by performing an atomic operation to set the dirty bit identifier of the target card corresponding to the target object in the card table to the first identifier, the writing failure of the dirty bit identifier corresponding to the target card page due to data contention is effectively avoided, thus ensuring the accurate updating of the dirty bit identifier corresponding to the target card page in the card table. Especially in multi-core processors, card table consistency can be guaranteed through simple atomic operations without the need for complex software locks or complex synchronization processes.

[0081] In an optional embodiment of this application, the card table setting instruction includes: a device identifier of a source register, wherein the source register is used to store the object address of the target object. Optionally, the device identifier of the source register can be the physical address of the source register, or the device number of the source register, etc., which corresponds to the physical address. For example, the source register can be a register configured by the virtual machine at its startup to store the object address of the target object in which the assignment operation of the reference type field occurs.

[0082] In some embodiments, the virtual machine can be used to write the object address of the target object to the source register in response to the assignment operation of the reference type field of the target object, and then drive the processor core 101 to execute the card table setting instruction, which includes the device identifier of the source register.

[0083] In some embodiments, such as Figure 6 As shown, the processor core 101 includes an instruction fetch unit 1011, an extended instruction decoding unit 1012, and a pre-execution unit 1013. The instruction fetch unit 1011 reads the card table setup instruction and transmits it to the extended instruction decoding unit 1012. The extended instruction decoding unit 1012 parses the card table setup instruction, obtains the device identifier of the source register, and triggers the pre-execution unit 1013 to read the object address from the source register and perform a null pointer check on the object address. The pre-execution unit 1013 is also used to, if the check result indicates that the object address is not null, trigger the write barrier acceleration unit 102 to determine the card index of the target card corresponding to the target object in the card table based on the object address, and atomically set the dirty bit identifier of the target card in the card table to the first identifier based on the card index; if the check result indicates that the object address is null, trigger null pointer exception handling.

[0084] In one optional configuration, the instruction fetch unit 1011 can be a dedicated hardware circuit that supports reading card table setup instructions under the drive of the virtual machine and transmitting the card table setup instructions to the extended instruction decoding unit 1012. For example, in response to an assignment operation of a reference type field of a target object, the virtual machine can drive the instruction fetch unit 1011 to jump and read the card table setup instructions, and after reading the card table setup instructions, transmit them to the extended instruction decoding unit 1012. In another optional configuration, the extended instruction decoding unit 1012 can be a dedicated hardware circuit that supports parsing the card table setup instructions upon receipt, obtaining the device identifier of the source registers included in the instructions, transmitting the device identifier to the pre-execution unit 1013, and triggering the pre-execution unit 1013 to read the object address from the source registers. In another optional embodiment, the pre-execution unit 1013 can be a dedicated hardware circuit that supports reading the object address from the source register and performing a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit 102 is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address, and the dirty bit identifier of the target card in the card table is atomically set to the first identifier based on the card index. Further optionally, the pre-execution unit 1013 is also configured to trigger null pointer exception handling if the check result indicates that the object address is null. In some embodiments, null pointer exception handling is used to trigger the virtual machine to throw a null pointer exception (NPE). For example, null pointer exception handling is used to be caught by the virtual machine's exception handler to throw an NPE.

[0085] In some embodiments, the card table setup instruction includes a target operation identifier, which is used to identify that the instruction type is a card table setup instruction. The processor core 101 is further configured to parse instructions executed by the virtual machine driver, and when a card table setup instruction including the target operation identifier is parsed, trigger the write barrier acceleration unit to set the dirty bit identifier of the target card corresponding to the target object in the card table to a first identifier.

[0086] The instructions executed by the processor core may include an operation identifier, which identifies the instruction type. The processor core 101 can be used to parse the instruction to obtain the operation identifier when executing instructions under the drive of a virtual machine. If the operation identifier is a target operation identifier, it indicates that a card table setting instruction has been parsed, triggering the write barrier acceleration unit to set the dirty bit identifier of the target card corresponding to the target object in the card table to the first identifier.

[0087] In some embodiments, the card table setup instructions can be custom RISC-V instructions, without modifying the underlying instruction set architecture (ISA) under the RISC-V instruction set architecture, thus fully utilizing the scalability of the RISC-V instruction set and coexisting with vector extensions.

[0088] The format of the card table setup instruction is a RISC-V custom encoding format, specifically R-type or I-type. The card table setup instruction includes the opcode, main function code (funct3), extended function code (funct7), first source operand (rs1), second source operand (rs2), and destination operand (rd). The values ​​of the opcode and main function code form the destination operation identifier; the first source operand is the device identifier of the source register; the extended function code is the custom opcode; and the second source operand and destination operand are placeholder data (e.g., 0). The opcode, main function code, and extended function code can all be used to identify the instruction type. The first source operand is the device identifier of the source register storing the object address of the target object (or address information indicating the object address). The second source operand is the device identifier of the source register storing another source data. The destination operand is the device identifier of the destination register storing the instruction result. In the card table setup instruction, the second source operand and destination operand are not used.

[0089] In other words, when the opcode value is the first value and the main function code value is the second value, the values ​​of the opcode and main function code together form the target operation identifier. That is, the card table setting instruction includes an opcode with the first value and a main function code with the second fetch instruction. The processor core 101 can be used to parse the instruction to obtain the opcode and main function code when executing the instruction under the virtual machine's drive. When the opcode value is the first value and the main function code value is the second value, it indicates that a card table setting instruction has been parsed. This triggers the write barrier acceleration unit to determine the card index of the target card corresponding to the target object in the card table based on the object address, and sets the dirty bit identifier of the target card corresponding to the target object in the card table to the first identifier based on the card index.

[0090] In some embodiments, the opcode value can be a custom opcode. The main function code value can be a value representing a write barrier instruction set under the RISC-V instruction set architecture. For example, the card table setup instruction is encoded using an encoding allocated in the opcode space reserved under the RISC-V instruction set architecture. Specifically, for example, such as... Figure 7As shown, the format of the card table setting instruction is R-type, and its encoding can be 32-bit binary data. In this encoding, bits 1 to 7 are the opcode, bits 8 to 12 are the destination operand, bits 13 to 15 are the main function code, bits 16 to 20 are the first source operand, bits 21 to 25 are the second source operand, and bits 26 to 32 are the extended function code. Specifically: In the encoding of the CST instruction, CST[6:0] = opcode = 0b0001011 (custom value), CST[11:7]rd = 0x00 (i.e., 0), CST[14:12] = funct3 = 0b011, CST[19:15] = rs1 = Id_reg, CST[24:20] = rs2 = 0x00 (i.e., 0), CST[24:20] = funct7 = 0b0010011. Id_reg represents the device identifier of the source register storing the target address of the target object.

[0091] When processor core 101 executes an instruction under the virtual machine's drive, it parses the instruction to obtain the opcode and main function code funct3. If the opcode value is 0b0001011 (i.e., the first value) and the main function code funct3 value is 0b011 (i.e., the second value), it indicates that a card table setting instruction has been parsed. This triggers the write barrier acceleration unit to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card corresponding to the target object in the card table is set to the first identifier.

[0092] In some embodiments, the extended instruction decoding unit 1012 is further configured to parse the instructions transmitted by the instruction fetch unit 1011. When a card table setting instruction including a target operation identifier is parsed, the pre-execution unit 1013 is triggered to read the object address from the source register and perform a null pointer check on the object address. The pre-execution unit 1013 is further configured to, if the check result indicates that the object address is not null, trigger the write barrier acceleration unit 102 to determine the card index of the target card corresponding to the target object in the card table based on the object address, and atomically set the dirty bit identifier of the target card in the card table to the first identifier based on the card index.

[0093] For example, when the extended instruction decoding unit 1012 receives the instruction transmitted by the instruction fetch unit 1011, it parses the instruction to obtain the opcode and the main function code funct3. If the opcode is 0b0001011 (i.e., the first value) and the main function code funct3 is 0b011 (i.e., the second value), it indicates that a card table setting instruction has been parsed. This triggers the pre-execution unit 1013 to read the object address from the source register and perform a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit 102 is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to the first identifier.

[0094] As a further example, the data processing device 10 includes: a processor core 101, a virtual machine running on the processor core 101, a write barrier acceleration unit 102 coupled to the processor core 101, and a memory subsystem 104. The memory subsystem includes memory 103 and its controller, etc.

[0095] The processor core 101 includes: an instruction fetch unit 1011, an extended instruction decoding unit 1012, and a pre-execution unit 1013. For example... Figure 8 As shown, the processor pipeline of processor core 101 includes five standard stages: fetch (F), decode (D), execute (E), memory access (M), and write back (W). That is, processor core 101 needs to go through these five standard stages sequentially to execute a single instruction.

[0096] During the instruction fetching stage, the instruction fetching unit 1011 reads the card table setting instruction (i.e., reads the binary code of the card table setting instruction) and transmits the card table setting instruction to the extended instruction decoding unit 1012.

[0097] During the decoding stage, the extended instruction decoding unit 1012 parses the received instruction, identifies the opcode with a value of 0b0001011 and the main function code funct3 with a value of 0b011, indicating that the received instruction is a card table setting instruction (i.e., it identifies opcode = 0b0001011, funct3 = 0b011, confirming the card table setting instruction), and identifies the value of the first source operand rs1 to obtain the device identifier of the source register, triggering the pre-execution unit 1013 to perform the following operations during the execution stage.

[0098] During the execution phase, the pre-execution unit 1013 reads the source register rs1 to obtain the object address in the source register and performs a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit 102 is triggered to perform the following operations during the memory access phase; if the check result indicates that the object address is null, null pointer exception handling is triggered.

[0099] During the memory access phase, the write barrier acceleration unit 102 calculates the card index of the target card corresponding to the target object in the card table based on the object address of the target object, the card table mask pre-configured in the mask parameter register, and the address shift bits pre-configured in the calculation parameter register. Based on the card index and the card table base address pre-configured in the base address parameter register, it calculates the physical address of the target card. Based on the physical address, it performs an atomic operation to send a write request for the dirty bit identifier of the target card to the memory subsystem 104. This write request is used to set the dirty bit identifier of the target card to the first identifier based on the physical address.

[0100] Since the execution result of the card table setup instruction is to write the dirty bit flag to the card table in memory, it does not generate data that needs to be written back to the destination register (which can be a general-purpose register). Therefore, in the write-back phase, after the marking instruction is completed and the instruction is determined to be retired, there is no write-back operation to the destination register.

[0101] In optional embodiments of this application, the virtual machine includes at least one of the following: a just-in-time compiler, an interpreter, The just-in-time compiler is used to generate card table setting instructions based on the system architecture description file, in response to compilation operations of reference type assignment bytecode (such as putfield) of the target object. The system architecture description file is used to indicate the matching rules between instructions and bytecode. The interpreter, in response to the interpretation and execution of reference type assignment bytecode of the target object, drives the processor core to jump and execute card table setup instructions inlined within the machine instruction template. This machine instruction template indicates the instruction sequence corresponding to the bytecode. The reference type assignment bytecode describes the assignment operation of a reference type field. The virtual machine is also used to drive the processor core 101 to jump and execute the card table setting instructions generated by the just-in-time compiler in response to the assignment operation of the reference type field of the target object, or to execute the interpretation and execution operation of the reference type assignment bytecode of the target object through the interpreter, so as to drive the processor core 101 to jump and execute the card table setting instructions inlined in the machine instruction template.

[0102] An interpreter is an execution component that translates bytecode into machine code line by line during program execution; it can be a component of a virtual machine. When the virtual machine starts, the interpreter generates a machine instruction template for the opcode of each bytecode and records the mapping relationship between the opcode and the base address of the machine instruction template. The machine instruction template is a sequence of executable machine code corresponding to the bytecode. During virtual machine execution, when interpreting and executing a bytecode, the interpreter, based on the mapping relationship between the opcode and the base address of the machine instruction template, obtains the base address of the machine instruction template corresponding to the opcode of that bytecode, and modifies the instruction register of processor core 101 to that base address, so that processor core 101 jumps to execute the machine code in the machine instruction template corresponding to that bytecode based on that base address.

[0103] Optionally, based on this, CST instructions can be pre-inlined in the machine instruction template corresponding to the opcode of the reference type assignment bytecode. When the interpreter responds to the interpretation and execution operation of the reference type assignment bytecode of the target object, it can modify the instruction register of the processor core 101 to the base address of the machine instruction template corresponding to the opcode of the reference type assignment bytecode, so that the processor core 101 jumps to execute the CST instruction in the machine instruction template corresponding to the bytecode based on the base address, thereby driving the processor core 101 to jump to execute the CST instruction inlined in the machine instruction template.

[0104] A Just-In-Time (JIT) compiler is a dynamic compilation component that translates bytecode into machine code (i.e., machine instructions) at runtime. Combining the advantages of compiled execution and interpreted execution, the JIT compiler triggers compilation and caches machine code only on the first call to the code. Subsequent calls to the same code can directly execute the cached machine code, avoiding repeated translations. The JIT compiler has a system architecture description file, which specifies the matching rules between instructions and bytecode. By modifying the matching rules between reference type assignment bytecode and CST instructions in the JIT's system architecture description file, the JIT can generate CST instructions based on the system architecture description file in response to the compilation operation of reference type assignment bytecode of the target object.

[0105] In some embodiments, it is assumed that the just-in-time compiler is a C2 compiler (a type of JIT). The C2 compiler has a system architecture description file (riscv.ad) under the RISC-V instruction set architecture. This system architecture description file is used to indicate the matching rules of registers, intermediate representation (IR) nodes, and machine instructions in the RISC-V instruction set.

[0106] The C2 compiler's bytecode compilation process includes: converting the bytecode into an IR graph, optimizing the IR graph, calling the matcher, traversing the optimized IR graph, and translating the nodes in the IR graph into machine code based on the matching rules described in the system architecture description file, ultimately obtaining all the machine code based on the bytecode conversion. The IR graph consists of nodes and directed edges. Nodes represent operations, values, or control flow; directed edges represent dependencies such as data flow, control flow, or memory dependencies between nodes.

[0107] Optionally, the matching rules of the StoreP (store pointer) node (an IR node) corresponding to the reference type assignment bytecode (putfield) in the system architecture description file can be modified in advance to change the load-cond-branch-store instruction sequence matching the StoreP node in the system architecture description file to CST instructions (a type of machine instruction), making the StoreP node match the CST instructions. When the C2 compiler responds to the compilation operation of the reference type assignment bytecode (putfield) of the target object, it can translate the StoreP node into CST instructions based on the modified system architecture description file to generate CST instructions.

[0108] Therefore, in response to the assignment operation of the reference type field of the target object, the virtual machine can directly drive the processor core 101 to jump and execute the CST instruction generated by the just-in-time compiler. For example, the C2 compiler can pre-compile the bytecode into machine code and store it in a dedicated area in memory when the virtual machine starts, so that the virtual machine can drive the processor core 101 to jump and execute the pre-generated machine code in the dedicated area by modifying the instruction register of the processor core 101 during operation.

[0109] Furthermore, when the virtual machine includes a JIT compiler and an interpreter, the interpreter can, upon detecting hot bytecode and responding to its interpretation and execution, simultaneously invoke the JIT compiler to compile the hot bytecode based on the system architecture description file, generating the corresponding machine code. This allows the virtual machine to subsequently drive the processor core 101 to jump and execute the JIT-generated machine code in response to the processing operations described by the hot bytecode. Here, hot bytecode refers to the bytecode corresponding to hot method code, which is method code whose call count exceeds a threshold. Optionally, in response to the assignment operation of a reference type field of a target object, the virtual machine can directly drive the processor core 101 to jump and execute the JIT-generated card table setting instruction if pre-generated JIT machine code exists. Otherwise, if pre-generated JIT machine code does not exist, the interpreter executes the interpretation and execution of the target object's reference type assignment bytecode to drive the processor core 101 to jump and execute the card table setting instruction inlined within the machine instruction template.

[0110] In this optional embodiment, by designing and compressing dozens of instructions implementing the write barrier mechanism in related technologies into a single instruction, the JIT compiler no longer needs to generate a complex sequence of write barrier instructions, including null pointer detection, address calculation, and memory access. Instead, it only needs to generate a single card table setting instruction, effectively simplifying the JIT compilation logic and reducing its compilation burden. Similarly, this allows the inline assembly card table setting instruction in the putfield branch of the bytecode main loop in the interpreter's machine instruction template to replace the C function calls involved in the original write barrier mechanism, simplifying the interpreter's compilation and interpretation burden.

[0111] In an optional embodiment of this application, the virtual machine is further configured to, in response to an assignment operation of a reference type field of a target object, drive the processor core 101 to execute a card table setting instruction when it is determined that the data processing device supports the write barrier acceleration unit.

[0112] In some embodiments, the virtual machine may detect whether the data processing device supports the write barrier acceleration unit when it starts up, so that if the data processing device supports the write barrier acceleration unit, the processor core 101 can be driven to execute the card table setting instruction in response to the assignment operation of the reference type field of the target object.

[0113] In an optional scenario, the virtual machine triggers an illegal instruction exception when executing a non-existent illegal instruction. The virtual machine can utilize this mechanism to drive the processor core 101 to execute a card table setup instruction during its startup. This instruction is used to detect whether the triggered illegal instruction exception will be caught, and to determine whether the data processing device supports the write barrier acceleration unit based on the capture result. If the virtual machine catches an illegal instruction exception, it can determine that the data processor does not support the write barrier acceleration unit; if it does not catch an illegal instruction exception, or if it catches the expected execution result of the card table setup instruction, it can determine that the data processor supports the write barrier acceleration unit.

[0114] Alternatively, a specific status register can be configured for the card table setup instructions. The virtual machine reads this status register during its startup to check its presence, and based on the result, determines whether the data processing device supports the write barrier acceleration unit. If the virtual machine reads data from the status register, indicating its presence, it can determine that the data processor supports the write barrier acceleration unit; if it does not read data from the status register, indicating its absence, it can determine that the data processor does not support the write barrier acceleration unit.

[0115] In this optional embodiment, the virtual machine detects whether the data processing device supports the write barrier acceleration unit to effectively determine whether it supports the custom card table setting instruction, thereby effectively avoiding the triggering of illegal instruction exceptions for the card table setting instruction when the data processing device does not support the write barrier acceleration unit, and reducing the code execution exception rate.

[0116] In this embodiment, the data processing device includes a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core. The virtual machine, in response to an assignment operation on a reference type field of a target object, drives the processor core to execute a card table setup instruction. This instruction includes address information indicating the object address of the target object. The processor core executes the card table setup instruction to perform a null pointer check on the object address. If the check indicates the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to a first identifier. This first identifier is used to identify the existence of a cross-generational reference in the target card, effectively implementing the write barrier mechanism. This technical solution can use a single card table setup instruction to replace dozens of instructions in related technologies, utilizing the write barrier acceleration unit to implement the write barrier mechanism, thereby effectively shortening the execution cycle of the write barrier and reducing device overhead. Furthermore, by performing a null pointer check on the object address of the target object before the write barrier acceleration unit executes the dirty bit flag setting operation, invalid object addresses can be effectively identified. This effectively avoids the problem of the write barrier acceleration unit failing to set the dirty bit flag due to invalid object addresses, thus ensuring the operational stability of the write barrier mechanism.

[0117] For ease of understanding, the technical solution of this application is further illustrated below with examples. For example, the data processing device includes: a processor core 101, a virtual machine running on the processor core 101, a write barrier acceleration unit 102 coupled to the processor core 101, and a memory subsystem 104. The data processing device is an OS system and supports the RISC-V instruction set architecture. The virtual machine is a JVM and includes a JIT and an interpreter.

[0118] The write barrier acceleration unit 102 can be integrated into the load-store unit (LSU) of the processor core, specifically into the memory access pipeline's storage queue within the LSU. For example... Figure 9As shown, the write barrier acceleration unit 102 includes: a Control and Status Register (CSR) group 1028, an accelerated execution module 1024, and a memory request interface 1027. The Control and Status Register group 1028 includes: a base address parameter register card_base (64 bits wide) 1021, a mask parameter register card_mask (64 bits wide) 1022, a calculation parameter register card_shift (5 bits wide) 1023, a dirty bit parameter register dirty_value (8 bits wide) 1025, and a path parameter register enable_skip (1 bit wide) 1026. It should be noted that each register in the Control and Status Register group 1028 can be allocated by the JVM in a custom CSR address space during its startup. For example, the base address of the base address parameter register card_base 1021 can be 0xBC0, the base address of the mask parameter register card_mask 1022 can be 0xBC1, and the base address of the dirty bit parameter register dirty_value 1025 can be 0xBC2, etc.

[0119] Furthermore, the accelerated execution module 1024 in the write barrier acceleration unit 102 includes: an index calculation submodule 10241 and an identifier writing submodule 10242. The index calculation submodule 10241 is used to calculate the card index idx of the target card corresponding to the target object in the card table based on the object address A of the target object, the card table mask card_mask, and the address shift number card_shift. The card index idx satisfies: idx = (A >> card_shift) & card_mask; and to calculate the identifier physical address card_addr of the target card based on the card index idx and the card table base address card_base. The identifier physical address card_addr satisfies: card_addr = card_base + idx.

[0120] The identifier write submodule 10242 is used to determine whether the path selection identifier in the path parameter register 1026 is the fast path identifier 1. If the path selection identifier is not the fast path identifier 1 (it is the slow path identifier 0), an atomic memory write operation is performed on the memory in the memory subsystem based on the identifier physical address card_addr through the memory request interface 1027 to set the dirty bit identifier of the target card in the card table to the first identifier 1. If the path selection identifier is the fast path identifier 1, the dirty bit identifier of the target card in the card table is read from the memory in the memory subsystem based on the identifier physical address card_addr through the memory request interface 1027, and it is determined whether the dirty bit identifier is the first identifier 1. If the dirty bit identifier is the second identifier 0, an atomic memory write operation is performed on the memory in the memory subsystem based on the identifier physical address card_addr through the memory request interface 1027 to set the dirty bit identifier of the target card in the card table to the first identifier 1. If the dirty bit identifier is the first identifier 1, the instruction is retired.

[0121] For example, the identifier writing submodule 10242 performs the following hardware logic without software intervention: loop: lr.b t0, (card_addr) # Execute the LR instruction to atomically load the dirty bit flag of the target card page into register t0 based on the physical address card_addr. beqz enable_skip, do_store # If the data (path selection identifier) ​​stored in the path parameter register enable_skip is 0, then jump to execute the code do_store bne t0, dirty_value, do_store # If the data stored in register t0 (the dirty bit flag of the target card page) is not equal to the data stored in the dirty bit parameter register dirty_value (the first flag), it indicates that the dirty bit flag of the target card page does not indicate the existence of a cross-generational reference in the target card, then jump to execute code do_store j skip # Unconditional jump code skip do_store: sc.b t1, dirty_value, (card_addr) # Executes the SC instruction to atomically write the data (first identifier) ​​stored in the dirty_value register to the memory space indicated by the identifier physical address card_addr, that is, to atomically write the dirty identifier of the target card page in the card table, and writes the execution status code of the SC instruction to register t1, which indicates whether the SC instruction was executed successfully.

[0122] `bnez t1, loop` # If the data stored in register t1 (execution status code) is not 0, meaning the SC instruction failed, then jump to the loop code until SC executes successfully. Testing shows that it usually succeeds after 1 to 2 loops.

[0123] skip: # Execute subsequent logic like Figure 10 As shown, the process by which the data processing device implements the write barrier mechanism includes the following steps 1001 to 1010.

[0124] Step 1001: When the JVM / OS starts, it can allocate card table memory and configure the control status register group 1028 in the write barrier acceleration unit 102 through the CSR instruction to write the card table base address to the base address parameter register 1021, write the card table mask to the mask parameter register 1022, write the address shift bits to the calculation parameter register 1023, write the first identifier to the dirty bit parameter register 1025, and write the path selection identifier to the path parameter register 1026.

[0125] Here, "card table memory" refers to the area in memory that stores the card table. The first identifier written to the dirty bit parameter register 1025 is 1, and the path selection identifier written to the path parameter register 1026 is the fast path identifier 1.

[0126] Step 1002: In response to the assignment operation of the reference type field of the target object, the JVM executes the CST instruction through the JIT / interpreter driver processor core 101.

[0127] Specifically, the JVM can respond to an assignment operation on a reference type field of the target object by driving processor core 101 to jump and execute a CST instruction generated by the JIT, or by interpreting and executing the bytecode of the target object's reference type assignment through the interpreter, thereby driving processor core 101 to jump and execute a CST instruction inlined within the machine instruction template. This also involves writing the object address A of the target object into the source register. The CST instruction includes the device identifier rs1 of the source register.

[0128] Step 1003: The processor core 101 parses the CST instruction to obtain the device identifier rs1 of the source register and reads the object address A of the target object in the source register.

[0129] Step 1004: Processor core 101 determines whether object address A is equal to 0, and performs a null pointer check on object address A. If object address A is equal to 0, proceed to step 1005; if object address A is not equal to 0, proceed to step 1006.

[0130] Step 1005: Processor core 101 triggers a null pointer exception, causing the JVM to throw an NPF.

[0131] Step 1006: The index calculation submodule 10241 of the write barrier acceleration unit 102 calculates the card index idx of the target card corresponding to the target object in the card table based on the object address A of the target object, the card table mask card_mask, and the address shift number card_shift. The card index idx satisfies: idx=(A>>card_shift)&card_mask.

[0132] Step 1007: The index calculation submodule 10241 of the write barrier acceleration unit 102 calculates the physical address card_addr of the target card based on the card index idx and the card table base address card_base. The physical address card_addr satisfies: card_addr = card_base + idx.

[0133] Step 1008: Write the identifier of the write barrier acceleration unit 102 to the submodule 10242, and determine whether the path selection identifier in the path parameter register 1026 is the fast path identifier 1. If not, proceed to step 1009; if yes, proceed to step 1010.

[0134] Step 1009: The identifier writing submodule 10242 performs an atomic memory write operation on the memory in the memory subsystem based on the identifier physical address card_addr through the memory request interface 1027, so as to set the dirty bit identifier of the target card in the card table to the first identifier 1.

[0135] It is not difficult to understand that the identifier writing submodule 10242 of the write barrier acceleration unit 102 can perform atomic operations directly and set the dirty bit identifier to the first identifier 1 without checking the current value of the dirty bit identifier of the target card in the card table when the path selection identifier is not the fast path identifier 1 (but the slow path identifier 0).

[0136] Step 1010: The identifier writing submodule 10242 reads the dirty bit identifier of the target card in the card table from memory based on the identifier physical address card_addr via the memory request interface 1027. It determines whether the dirty bit identifier is the first identifier 1. Only if the dirty bit identifier is the second identifier 0, it performs an atomic memory write operation on the memory subsystem based on the identifier physical address card_addr via the memory request interface 1027 to set the dirty bit identifier of the target card in the card table to the first identifier 1.

[0137] It is not difficult to understand that the identifier writing submodule 10242 of the write barrier acceleration unit 102 can, when the path selection identifier is fast path identifier 1, first read the current value of the dirty bit identifier of the target card in the card table, so that if the dirty bit identifier has already identified the target card as dirty, that is, if it has been identified as having a cross-generation reference, skip the write operation of the dirty bit identifier of the target card and the CST instruction is retired; only when the dirty bit identifier has not identified the target card as dirty, that is, if it has not been identified as having a cross-generation reference, perform an atomic write operation on the dirty bit identifier of the target card and update the dirty bit identifier to the first identifier 1.

[0138] In this example, by setting a path selection flag, when the path selection flag is a fast path flag, if the dirty flag indicates that the target card is dirty, there is no need to perform an update operation on the dirty flag of the target card in the card table. This can further reduce the card table update by approximately 2-3 clock cycles when the dirty flag indicates that the target card is dirty, significantly reducing the number of memory accesses. This effectively reduces bus traffic and power consumption between the write barrier acceleration unit 102 and the memory subsystem, improving write barrier processing efficiency. By running the lusearch and pmd detectors in the DaCapo suite benchmark tests, the average time for 1,000 write barrier operations implemented by the JVM using software programs in related technologies is approximately 320 clock cycles. However, the sampling method of this application, utilizing the write barrier acceleration unit as hardware, reduces the average time for 1,000 write barrier operations to 26 clock cycles, significantly improving write barrier processing efficiency by approximately 12.3 times, resulting in an overall application throughput increase of approximately 15%. Furthermore, the technical solution of this application clearly mainly involves the JVM instruction generation adaptation direction, requiring no modification to the Java application and possessing characteristics of transparency to upper layers. The parameters of the data processing device include: support for the RISC-V Rocket Chip framework, a 64-bit RV64GC processor core, and a clock speed of 200MHz; the CST instruction is in R-type format, and its encoding includes opcode: 0b0001011, main function code: 0b011, and extended function code: 0b0010011.

[0139] In this embodiment, the data processing device includes a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core. The virtual machine, in response to an assignment operation on a reference type field of a target object, drives the processor core to execute a card table setup instruction. This instruction includes address information indicating the object address of the target object. The processor core executes the card table setup instruction to perform a null pointer check on the object address. If the check indicates the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to a first identifier. This first identifier is used to identify the existence of a cross-generational reference in the target card, effectively implementing the write barrier mechanism. This technical solution can use a single card table setup instruction to replace dozens of instructions in related technologies, utilizing the write barrier acceleration unit to implement the write barrier mechanism, thereby effectively shortening the execution cycle of the write barrier and reducing device overhead. Furthermore, by performing a null pointer check on the object address of the target object before the write barrier acceleration unit executes the dirty bit flag setting operation, invalid object addresses can be effectively identified. This effectively avoids the problem of the write barrier acceleration unit failing to set the dirty bit flag due to invalid object addresses, thus ensuring the operational stability of the write barrier mechanism.

[0140] Please refer to Figure 11 This document illustrates a flowchart of a data processing method provided in an embodiment of this application. This data processing method can be applied to a data processing apparatus, which includes a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core. Optionally, the data processing method can be applied to any of the data processing methods provided in the embodiments of this application. Figure 11 As shown, the data processing methods include: Step 1101: In response to the assignment operation of the reference type field of the target object, the virtual machine drives the processor core to execute the card table setting instruction. The card table setting instruction includes address information, which indicates the object address of the target object.

[0141] Step 1102: The processor core executes the card table setting instruction to perform a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to the first identifier. The first identifier is used to identify that there is a cross-generational reference in the target card.

[0142] In this embodiment, the data processing device includes a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core. The virtual machine, in response to an assignment operation on a reference type field of a target object, drives the processor core to execute a card table setup instruction. This instruction includes address information indicating the object address of the target object. The processor core executes the card table setup instruction to perform a null pointer check on the object address. If the check indicates the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to a first identifier. This first identifier is used to identify the existence of a cross-generational reference in the target card, effectively implementing the write barrier mechanism. This technical solution can use a single card table setup instruction to replace dozens of instructions in related technologies, utilizing the write barrier acceleration unit to implement the write barrier mechanism, thereby effectively shortening the execution cycle of the write barrier and reducing device overhead. Furthermore, by performing a null pointer check on the object address of the target object before the write barrier acceleration unit executes the dirty bit flag setting operation, invalid object addresses can be effectively identified. This effectively avoids the problem of the write barrier acceleration unit failing to set the dirty bit flag due to invalid object addresses, thus ensuring the operational stability of the write barrier mechanism.

[0143] In some embodiments, the data processing apparatus further includes a caching module. The write barrier acceleration unit determines the card index of the target card corresponding to the target object in the card table based on the object address, and the process of atomically setting the dirty bit identifier of the target card in the card table to a first identifier based on the card index may include: When the write barrier acceleration unit has a dirty bit identifier for the target card corresponding to the target object in the cache module, it reads the dirty bit identifier of the target card in the cache module. If the dirty bit identifier is the second identifier, it determines the card index of the target card in the card table based on the object address, and atomically updates the dirty bit identifier of the target card in the cache module and memory to the first identifier based on the card index. If the dirty bit identifier of the target card corresponding to the target object does not exist in the cache module, it determines the card index of the target card in the card table based on the object address, reads the dirty bit identifier of the target card in memory based on the card index, and atomically updates the dirty bit identifier of the target card in memory to the first identifier based on the card index if the dirty bit identifier is the second identifier. The second identifier is used to indicate that there is no cross-generational reference in the target card.

[0144] In some embodiments, the data processing method further includes: when the write barrier acceleration unit does not have a dirty bit identifier for the target card corresponding to the target object in the cache module, adding the dirty bit identifier of the target card to the card table in the cache module.

[0145] In some embodiments, a virtual machine includes at least one of the following: a just-in-time compiler, an interpreter, The just-in-time compiler is used to generate card table setting instructions based on the system architecture description file and in response to the compilation operation of bytecode assignment to the reference type of the target object. The system architecture description file is used to indicate the matching rules between instructions and bytecode. The interpreter is used to interpret and execute the reference type assignment bytecode of the target object, driving the processor core to jump and execute the card table setting instructions inlined in the machine instruction template. The machine instruction template is used to indicate the instruction sequence corresponding to the bytecode. The reference type assignment bytecode describes the assignment operation of the reference type field. Step 1101: In response to the assignment operation of the reference type field of the target object, the process of the virtual machine driving the processor core to execute the card table setting instruction may include: In response to an assignment operation on a reference type field of the target object, the virtual machine drives the processor core to jump and execute the card table setup instructions generated by the just-in-time compiler, or it executes the interpretation and execution operation of the reference type assignment bytecode of the target object through the interpreter, thereby driving the processor core to jump and execute the card table setup instructions inlined in the machine instruction template.

[0146] In some embodiments, the card table setup instruction includes: a device identifier of a source register, the source register being used to store the object address of the target object; the processor core includes: an instruction fetch unit, an extended instruction decoding unit, and a pre-execution unit; step 1102: the processor core executes the card table setup instruction to perform a null pointer check on the object address, and if the check result indicates that the object address is not null, triggers the write barrier acceleration unit to determine the card index of the target card corresponding to the target object in the card table based on the object address, and the process of atomically setting the dirty bit identifier of the target card in the card table to the first identifier based on the card index may include: The instruction fetching unit reads the card table setting instruction and transmits the card table setting instruction to the extended instruction decoding unit; The extended instruction decoding unit parses the card table setting instruction, obtains the device identifier of the source register, and triggers the pre-execution unit to read the object address from the source register and perform a null pointer check on the object address; If the check result indicates that the object address is not null, the pre-execution unit triggers the write barrier acceleration unit to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card is atomically set to the first identifier. If the check result indicates that the object address is null, null pointer exception handling is triggered.

[0147] In some embodiments, the card table setting instruction includes a target operation identifier; the data processing method further includes: an extended instruction decoding unit parsing the instruction passed by the instruction fetching unit, and when a card table setting instruction including a target operation identifier is parsed, a pre-execution unit is triggered to read the object address from the source register.

[0148] In some embodiments, the format of the card table setup instruction is a RISC-V custom encoding format. The card table setup instruction includes an opcode, a main function code, an extended function code, a first source operand, a second source operand, and a target operand. The values ​​of the opcode and the main function code form the target operation identifier. The value of the first source operand is the device identifier of the source register. The value of the extended function code is a custom opcode. The values ​​of the second source operand and the target operand are placeholder data.

[0149] In some embodiments, step 1101, in response to the assignment operation of the reference type field of the target object, the process of driving the processor core to execute the card table setting instruction may include: when the virtual machine determines that the data processing device supports the write barrier acceleration unit, in response to the assignment operation of the reference type field of the target object, driving the processor core to execute the card table setting instruction.

[0150] In some embodiments, the write barrier acceleration unit includes: a path parameter register, a base address parameter register, a mask parameter register, a calculation parameter register, a dirty bit parameter register, and an accelerated execution module; Step 1102: The processor core executes a card table setup instruction to perform a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. The process of atomically setting the dirty bit identifier of the target card in the card table to the first identifier based on the card index may include: The accelerated execution module is used to read the path selection identifier in the path parameter register, the card table base address stored in the base address parameter register, the card table mask stored in the mask parameter register, the address shift bits stored in the calculation parameter register, and the first identifier stored in the dirty bit parameter register; and when the path selection identifier is a fast path identifier, it is used to determine the card index of the target card corresponding to the target object in the card table based on the object address, the card table mask, and the address shift bits, calculate the identifier physical address of the target card based on the card index and the card table base address, and atomically set the dirty bit identifier of the target card to the first identifier based on the identifier physical address.

[0151] In this embodiment, the data processing device includes a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core. The virtual machine, in response to an assignment operation on a reference type field of a target object, drives the processor core to execute a card table setup instruction. This instruction includes address information indicating the object address of the target object. The processor core executes the card table setup instruction to perform a null pointer check on the object address. If the check indicates that the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to a first identifier. This first identifier is used to identify the existence of a cross-generational reference in the target card, effectively implementing the write barrier mechanism. This technical solution can use a single card table setup instruction to replace dozens of instructions in related technologies, utilizing the write barrier acceleration unit to implement the write barrier mechanism, thereby effectively shortening the execution cycle of the write barrier and reducing equipment overhead. Furthermore, by performing a null pointer check on the object address of the target object before the write barrier acceleration unit executes the dirty bit flag setting operation, invalid object addresses can be effectively identified. This effectively avoids the problem of the write barrier acceleration unit failing to set the dirty bit flag due to invalid object addresses, thus ensuring the operational stability of the write barrier mechanism.

[0152] This application also provides a processor, which includes any of the data processing devices provided in this application.

[0153] This application also provides a chip that includes any of the data processing devices provided in this application.

[0154] This application also provides an electronic device, which includes any of the data processing devices, chips, or processors provided in this application.

[0155] This application also provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the data processing method provided in this application when executed by a processor.

[0156] This application also provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the data processing method provided in this application.

[0157] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A data processing apparatus, characterized in that, The data processing device includes: a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core; The virtual machine is used to drive the processor core to execute a card table setting instruction in response to an assignment operation of a reference type field of the target object. The card table setting instruction includes address information, which indicates the object address of the target object. The processor core is used to execute the card table setting instruction to perform a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to a first identifier, which is used to identify that there is a cross-generational reference in the target card.

2. The data processing apparatus according to claim 1, characterized in that, The data processing device further includes: a cache module, the cache module being used to cache at least a portion of the card table in memory; The write barrier acceleration unit is further configured to: If a dirty bit identifier for the target card corresponding to the target object exists in the cache module, read the dirty bit identifier of the target card in the cache module; if the dirty bit identifier is a second identifier, determine the card index of the target card in the card table based on the object address; and atomically update the dirty bit identifier of the target card in the cache module and the memory to the first identifier based on the card index. If a dirty bit identifier for the target card corresponding to the target object does not exist in the cache module, determine the card index of the target card in the card table based on the object address; read the dirty bit identifier of the target card in the memory based on the card index; and if the dirty bit identifier is the second identifier, atomically update the dirty bit identifier of the target card in the memory to the first identifier based on the card index. The second identifier is used to indicate that there is no cross-generational reference in the target card.

3. The data processing apparatus according to claim 2, characterized in that, The write barrier acceleration unit is further configured to add the dirty bit identifier of the target card to the card table in the cache module when the dirty bit identifier of the target card corresponding to the target object does not exist in the cache module.

4. The data processing apparatus according to claim 1, characterized in that, The virtual machine includes at least one of the following: a just-in-time compiler, an interpreter, The just-in-time compiler is used to generate the card table setting instructions based on the system architecture description file in response to the compilation operation of the reference type assignment bytecode of the target object. The system architecture description file is used to indicate the matching rules between the instructions and the bytecode. The interpreter is used to drive the processor core to jump and execute the card table setting instruction inlined in the machine instruction template in response to the interpretation and execution operation of the reference type assignment bytecode of the target object. The machine instruction template is used to indicate the instruction sequence corresponding to the bytecode. The reference type assignment bytecode describes the assignment operation of the reference type field. The virtual machine is also used to drive the processor core to jump to execute the card table setting instruction generated by the just-in-time compiler in response to the assignment operation of the reference type field of the target object, or to execute the interpretation and execution operation of the reference type assignment bytecode of the target object through the interpreter, so as to drive the processor core to jump to execute the card table setting instruction inline assembled in the machine instruction template.

5. The data processing apparatus according to claim 1, characterized in that, The card table setting instruction includes: a device identifier of a source register, the source register being used to store the object address of the target object; the processor core includes: an instruction fetch unit, an extended instruction decoding unit, and a pre-execution unit; The instruction fetching unit is used to read the card table setting instruction and transmit the card table setting instruction to the extended instruction decoding unit; The extended instruction decoding unit is used to parse the card table setting instruction, obtain the device identifier of the source register, and trigger the pre-execution unit to read the object address from the source register and perform a null pointer check on the object address; The pre-execution unit is configured to, when the check result indicates that the object address is not empty, trigger the write barrier acceleration unit to determine the card index of the target card corresponding to the target object in the card table based on the object address, atomically set the dirty bit identifier of the target card to the first identifier based on the card index, and trigger null pointer exception handling when the check result indicates that the object address is empty.

6. The data processing apparatus according to claim 5, characterized in that, The card table setting instructions include: target operation identifier; The extended instruction decoding unit is also used to parse the instructions passed by the instruction fetching unit, and when the card table setting instruction including the target operation identifier is parsed, the pre-execution unit is triggered to read the object address from the source register.

7. The data processing apparatus according to claim 6, characterized in that, The card table setup instruction is in RISC-V custom encoding format. The card table setup instruction includes an opcode, a main function code, an extended function code, a first source operand, a second source operand, and a target operand. The values ​​of the opcode and the main function code constitute the target operation identifier. The value of the first source operand is the device identifier of the source register. The value of the extended function code is a custom opcode. The values ​​of the second source operand and the target operand are placeholder data.

8. The data processing apparatus according to claim 1, characterized in that, The virtual machine is also configured to, when it is determined that the data processing device supports the write barrier acceleration unit, drive the processor core to execute a card table setting instruction in response to an assignment operation of the reference type field of the target object.

9. The data processing apparatus according to claim 1, characterized in that, The write barrier acceleration unit includes: a path parameter register, a base address parameter register, a mask parameter register, a calculation parameter register, a dirty bit parameter register, and an accelerated execution module; The accelerated execution module is configured to read the path selection identifier in the path parameter register, the card table base address stored in the base address parameter register, the card table mask stored in the mask parameter register, the address shift bits stored in the calculation parameter register, and the first identifier stored in the dirty bit parameter register; and when the path selection identifier is a fast path identifier, to determine the card index of the target card corresponding to the target object in the card table based on the object address, the card table mask, and the address shift bits, to calculate the identifier physical address of the target card based on the card index and the card table base address, and to atomically set the dirty bit identifier of the target card to the first identifier based on the identifier physical address.

10. A data processing method, characterized in that, The method is applied to a data processing device, the data processing device including a processor core, a virtual machine running on the processor core, and a write barrier acceleration unit coupled to the processor core; the method includes: In response to the assignment operation of the reference type field of the target object, the virtual machine drives the processor core to execute the card table setting instruction, which includes address information indicating the object address of the target object; The processor core executes the card table setting instruction to perform a null pointer check on the object address. If the check result indicates that the object address is not null, the write barrier acceleration unit is triggered to determine the card index of the target card corresponding to the target object in the card table based on the object address. Based on the card index, the dirty bit identifier of the target card in the card table is atomically set to a first identifier, which is used to identify that there is a cross-generational reference in the target card.

11. An electronic device, characterized in that, The electronic device includes: the data processing apparatus according to any one of claims 1 to 9.