Data processing method and related device

By allocating a unique memory pool to the processor's processing cores and avoiding access by other cores or turning off hardware prefetching, the pseudo-sharing problem between processor cores is solved, improving memory access efficiency and throughput.

WO2025145647A1PCT designated stage expired Publication Date: 2025-07-10HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/116442
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2024-09-03
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The pseudo-sharing between multiple processing cores of the processor leads to inefficient memory access and reduced throughput, which is difficult to effectively solve in the prior art.

Method used

By obtaining the purpose of the data packet and determining its corresponding memory pool, it avoids other processing to check the read and write operations of the memory pool, or turns off the hardware prefetch function and passes independent storage requests to ensure that the processing core exclusively accesses the memory pool, reducing pseudo-sharing.

Benefits of technology

Improves memory access efficiency, reduces memory access latency, increases throughput, and optimizes processor performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116442_10072025_PF_FP_ABST
    Figure CN2024116442_10072025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing method and a related device, for use in avoiding false sharing among a plurality of processing cores of a processor. The method is applied to a processor, and the processor comprises a plurality of processing cores. When a first data packet is received, a target memory pool corresponding to a target processing core can be applied for, the target memory pool being a memory pool among at least one memory pool corresponding to the target processing core, and data to be processed of the first data packet is stored in the target memory pool. Because the processing cores other than the target processing core will not perform a read / write operation on said data in the target memory pool, false sharing among the plurality of processing cores is avoided, thereby improving the memory access efficiency, reducing the memory access delay, and increasing the throughput.
Need to check novelty before this filing date? Find Prior Art

Description

A data processing method and related equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 3, 2024, with application number 202410024931.9 and application name “A data processing method and device thereof”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a data processing method and related equipment. Background Art

[0003] As the number of processing cores in a processor increases, the business process of a computing task becomes longer. Parallel acceleration can effectively improve processor performance. During the parallel acceleration process, data sharing between different processing cores is a key factor affecting the performance gain of parallel acceleration.

[0004] Currently, competition for memory between multiple processing cores of a processor will result in false sharing. For example, the processor includes two processing cores, namely processing core A and processing core B. Among them, processing core A reads the value of variable x from cache x in memory, and sets the cacheline state of cache x in the cache of processing core A to exclusive. Then, when processing core B also reads the value of variable x from cache x in memory, it triggers the cacheline state of cache x in the cache of processing core A to be updated to shared. If processing core B modifies the value of variable x in cache x in memory, based on the consistency protocol, it will trigger the cacheline state of cache x in the cache of processing core A to be invalid.

[0005] At this time, if processing core A continues to access the value of variable x, it cannot obtain the value of cache x from the cache and can only obtain the value of variable x from the memory again, which reduces memory access efficiency, increases memory access latency, and reduces throughput.

[0006] Summary of the Invention

[0007] Embodiments of the present application provide a data processing method and related devices for avoiding false sharing between multiple processing cores of a processor.

[0008] The first aspect of the present application provides a data processing method for a processor, which includes multiple processing cores. By obtaining a first data message and determining the destination processing core corresponding to the first data message, the destination processing core is one of the multiple processing cores, and each processing core in the multiple processing cores corresponds to at least one memory pool. Then, a destination memory pool corresponding to the destination processing core can be applied, and the target memory pool is one of at least one memory pool corresponding to the target processing core, and the data to be processed of the first data message is stored in the target memory pool. Since other processing cores other than the destination processing core will not perform read and write operations on the data to be processed in the destination memory pool, false sharing between multiple processing cores is avoided, thereby improving memory access efficiency, reducing memory access latency, and increasing throughput.

[0009] In some possible implementations, the queue status of the multiple processing cores is obtained, where the queue status is the number of data to be processed in the corresponding processing cores, and the destination processing core is determined based on a preset scheduling strategy and the queue status of the multiple processing cores. The destination processing core is one of the multiple processing cores, and the scheduling strategy is to give priority to the processing core with the least amount of data to be processed among the multiple processing cores, thereby determining the destination processing core corresponding to the first data packet.

[0010] In some possible implementations, the destination memory pool corresponding to the destination processing core is determined from a preset correspondence table, where the correspondence table includes multiple correspondences, and the multiple correspondences include each processing core in the multiple processing cores and at least one corresponding memory pool, thereby applying for the destination memory pool corresponding to the destination processing core.

[0011] In some possible implementations, before receiving the first data packet, the hardware prefetch function of the destination processing core can be turned off. Then, after applying for the destination memory pool corresponding to the destination processing core, an independent storage request can be transmitted to the destination processing core. The independent storage request is used to instruct the destination processing core to exclusively access the destination memory pool, avoiding competition for the destination memory pool among other processing cores and avoiding false sharing among multiple processing cores, thereby improving memory access efficiency, reducing memory access latency, and increasing throughput.

[0012] In some possible implementations, after the data to be processed of the first data message is stored in the destination memory pool, the destination processing core can perform read and write operations on the data to be processed in the destination memory pool to obtain processed data, thereby completing the processing of the data to be processed in the first data message. If the data to be processed is in the cache of the destination processing core, the destination processing core can process the data to be processed in the cache. If the data to be processed is not in the cache of the destination processing core, the destination processing core can obtain the data to be processed from the destination memory pool, store the data to be processed in the cache, and perform the read and write operations on the data to be processed in the cache based on the operation instruction. Subsequently, it is only necessary to read the data to be processed in the cache, thereby improving memory access efficiency, reducing memory access latency, and increasing throughput.

[0013] In some possible implementations, after the destination processing core performs read and write operations on the data to be processed in the destination memory pool, a second data packet can be generated, and the second data packet includes the processed data. After the second data packet is generated, the destination memory pool is released, so that the status of the destination memory pool is reset to available, avoiding unnecessary occupation of the destination memory pool.

[0014] A second aspect of the present application provides a processor, and the network device is used to execute any one of the methods in the first aspect.

[0015] A third aspect of the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the method provided by the first aspect or any possible implementation of the first aspect.

[0016] A fourth aspect of the present application provides a computer program product, which includes computer-executable instructions, which are stored in a computer-readable storage medium; at least one processor of a device can read the computer-executable instructions from the computer-readable storage medium, and at least one processor executes the computer-executable instructions so that the device implements the method provided by the above-mentioned first aspect or any possible implementation of the first aspect.

[0017] In a fifth aspect, the present application provides a communication device, which may include at least one processor, a memory, and a communication interface. The at least one processor is coupled to the memory and the communication interface. The memory is configured to store instructions, the at least one processor is configured to execute the instructions, and the communication interface is configured to communicate with other communication devices under the control of the at least one processor. When executed by the at least one processor, the instructions cause the at least one processor to perform the method of the first aspect or any possible implementation of the first aspect.

[0018] In a sixth aspect, the present application provides a chip system, which includes a processor for supporting the implementation of the functions involved in the above-mentioned first aspect or any possible implementation method of the first aspect.

[0019] In a possible design, the chip system may further include a memory for storing necessary program instructions and data. The chip system may be composed of a chip or may include a chip and other discrete devices.

[0020] Among them, the technical effects brought about by the second to sixth aspects or any possible implementation methods thereof can refer to the technical effects brought about by the first aspect or different possible implementation methods of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] FIG1-1 is a schematic diagram of the structure of a system architecture provided in an embodiment of the present application;

[0022] Figures 1-2 are schematic diagrams of the structure of a communication network system provided in an embodiment of the present application;

[0023] FIG2-1 is a schematic diagram of the structure of a network device provided in an embodiment of the present application;

[0024] FIG2-2 is a schematic diagram of the composition structure of a software entity built into a network device provided in an embodiment of the present application;

[0025] FIG3-1 is a schematic diagram of an embodiment of a data processing method provided in an embodiment of the present application;

[0026] FIG3-2 is a schematic diagram of a process for the network device to receive multiple data packets and send multiple data packets in an embodiment of the present application;

[0027] FIG4-1 is a schematic diagram of another embodiment of a data processing method provided in an embodiment of the present application;

[0028] FIG4-2 is a schematic diagram of a process for the network device to receive multiple data packets and send multiple data packets in an embodiment of the present application;

[0029] FIG5 is a schematic diagram of the structure of a network device provided in an embodiment of the present application;

[0030] FIG6 is a schematic structural diagram of a communication device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] Embodiments of the present application provide a data processing method and related devices for avoiding false sharing between multiple processing cores of a processor.

[0032] The embodiments of the present application are described below with reference to the accompanying drawings.

[0033] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0034] Please refer to Figure 1-1, which shows a schematic diagram of a system architecture involved in an embodiment of the present application. The system architecture includes a communication network system 100, a first user device and a second user device. The communication network system 100 includes multiple network devices. Messages can be transmitted between the first user device and the second user device through the network devices in the communication network system 100.

[0035] In some possible implementations, the first user device or the second user device may be a terminal device or a server. In some possible implementations, the first user device or the second user device may be a physical device or a virtual device. The first user device or the second user device may also be a container instance. The first user device and the second user device may be the same type of communication device (a terminal device or a server, a physical device or a virtual device or a container instance), or different communication devices.

[0036] Among them, the terminal device can be called a terminal, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The terminal device can be a mobile phone, a tablet computer (pad), a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc. The embodiments of the present application do not limit the specific technology and specific device form adopted by the terminal device.

[0037] Among them, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, as well as big data and artificial intelligence platforms.

[0038] Since servers need to respond to service requests and process them to provide reliable services, they should generally be able to undertake and guarantee services. They should also have strong processing capabilities, high stability, high reliability, high security, scalability, and manageability. In the embodiments of the present application, the server may be an x86 server, also known as a complex instruction set computer (CISC) architecture server, commonly known as a personal computer (PC) server. This server is based on the PC architecture and uses an Intel or other x86 instruction set-compatible processor chip and a Windows operating system.

[0039] For example, referring to Figures 1-2, a communication network system 100 may include M network devices (M is a positive integer), wherein the M network devices are configured to forward messages between a first user device and a second user device. For example, the multiple network devices include network devices 102-107 (i.e., network device 102, network device 103, network device 104, network device 105, network device 106, and network device 107).

[0040] Each of network devices 102-107 may be a switch (virtual switch or physical switch) or a router (virtual router or physical router), etc., used to forward packets in communication network system 100. Network devices 102-107 may be the same type of network devices, for example, network devices 102-107 may all be routers. Alternatively, network devices 102-107 may be different types of network devices, for example, some of network devices 102-107 may be routers and others may be switches, without limitation herein.

[0041] It should be pointed out that the system architecture shown in Figures 1-1 and 1-2 above is for example only and is not intended to limit the technical solutions of the embodiments of the present application. In the specific implementation process, the communication network system 100 may also include other devices, and the number of network devices can be configured as needed.

[0042] In some possible implementations, the communication between any two network devices 102 to 107 can be concurrent with multiple processing cores and processed in a single tunnel based on the Internet Protocol Security (IPSec), which greatly improves the communication performance.

[0043] As shown in FIG2-1 , the network device 200 is one of the network devices 102 to 107 . The network device 200 includes at least the following hardware modules: a transceiver 201 , a processor 202 , a memory 203 , and a storage 204 . The storage 204 is further used to store instructions and data.

[0044] The transceiver 201 is used to receive data messages and send data messages.

[0045] The processor 202 can be a general-purpose processor, such as but not limited to a central processing unit (CPU), or a special-purpose processor, such as but not limited to a digital signal processor (DSP), an application-specific integrated circuit (ASIC), and a field programmable gate array (FPGA). The processor 204 can also be a neural processing unit (NPU). In addition, the processor 202 can also be a combination of multiple processing cores. In particular, in the technical solution provided in the embodiment of the present application, the processor 202 can be used to execute the relevant steps in the subsequent method embodiment. In the embodiment of the present application, the processor 202 may include multiple processing cores, wherein each processing core includes multiple cache layers, such as L1, L2, and L3. The processor 202 can be used to execute instructions stored in the memory 204 to control the transceiver 201 to receive data packets or send data packets.

[0046] The memory 204 can be various types of storage media, such as non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical memory, and registers. The memory 204 is specifically used to store instructions and data. The processor 202 can perform the steps and / or operations described in the method embodiments of the present application by reading and executing the instructions stored in the memory 204.

[0047] The memory 203 can be a storage medium such as a random access memory (RAM) or a read-only memory (ROM). After the processor 202 reads data from the memory 204, it stores it in the memory 203, then obtains the data from the memory 203, stores it in the cache of the processor 202, and then processes the data in the cache.

[0048] The memory 204 may store program code, and the processor 202 may read the program code from the memory 204 to form a corresponding process to execute the corresponding program. Different programs may serve as different software entities to execute different software modules. In an embodiment of the present application, as shown in FIG2-2 , the network device 200 may run the following multiple software modules through the processor 202 and memory 203: a message receiving module, a message sending module, a scheduling module, and a memory management module.

[0049] Based on the multiple software entities running on the aforementioned network device, an embodiment of the present application provides a data processing method and related devices for avoiding false sharing between multiple processing cores of a processor.

[0050] This application is described through two embodiments, namely, embodiment 1 and embodiment 2.

[0051] In the first embodiment of the present application, after receiving the first data message, the destination processing core corresponding to the first data message can be identified in advance, and then the corresponding destination memory pool can be applied for based on the correspondence table and the destination processing core. A preset correspondence table is provided in the memory management module, and the correspondence table includes multiple correspondences, and the multiple correspondences include each processing core in the multiple processing cores and at least one corresponding memory pool. Since other processing cores other than the destination processing core will not perform read and write operations on the data to be processed in the destination memory pool, false sharing between multiple processing cores is avoided, thereby improving memory access efficiency, reducing memory access latency, and increasing throughput.

[0052] In the second embodiment of the present application, when a corresponding table is not provided in the memory management module, the hardware prefetching function of the hardware prefetcher can be disabled in advance to prevent the hardware prefetcher from pre-setting some memory pools as exclusive memory pools for the destination processing core. At the same time, by transmitting an independent storage request to the destination processing core, the independent storage request is used to instruct the destination processing core to exclusively access the destination memory pool, thereby avoiding competition among other processing cores for the destination memory pool and avoiding false sharing among multiple processing cores, thereby improving memory access efficiency, reducing memory access latency, and increasing throughput.

[0053] As shown in FIG3-1 , a data processing method provided in Example 1 of the present application mainly includes the following steps:

[0054] 301. A message receiving module receives a first data message.

[0055] The first data message may be an Ethernet-based data message or an Internet-based data message, which is not limited herein. It should be noted that the first data message includes a message header and a payload, wherein the message header may include information such as a five-tuple (destination IP address, source IP address, destination port number, source port number, and communication protocol), and the payload may be application layer information, which is not limited herein.

[0056] In some possible implementations, the first data packet may refer to multiple data packets, such as data packet 1, data packet 2, and data packet 3. The multiple data packets may be received at different times. For example, data packet 1 is received at time 1, data packet 2 is received at time 2, and data packet 3 is received at time 3.

[0057] 302. The message receiving module determines a destination processing core corresponding to the first data message.

[0058] In an embodiment of the present application, the processor may include multiple processing cores. Exemplarily, the multiple processing cores may include core1, core2, ..., coreN.

[0059] In some possible implementations, the queue status of multiple processing cores can be obtained first, where the queue status is the number of data to be processed in the corresponding processing cores. Then, the destination processing core can be determined based on a preset scheduling strategy and the queue status of the multiple processing cores. The destination processing core is one of the multiple processing cores, and the scheduling strategy is to give priority to the processing core with the least amount of data to be processed among the multiple processing cores.

[0060] For example, if the number of data to be processed by core1 is 10 and the number of data to be processed by core2 is 100, then based on the scheduling policy, core1 may be preferentially selected as the destination processing core. After the destination processing core is determined, the identifier of the destination processing core is obtained and recorded as coreid.

[0061] 303. The message receiving module applies for a destination memory pool corresponding to the destination processing core through the memory management module.

[0062] In an embodiment of the present application, the storage space in the memory can be divided into multiple memory pools, and each of the multiple memory pools is assigned an identifier, denoted as poolid. For example, the identifiers of the multiple memory pools are pool11, pool22, ..., poolNN.

[0063] In an embodiment of the present application, a memory management module may be built into the network device, and the memory management module is a software entity, and the memory management module is used to manage the storage space in the memory. A preset correspondence table may be stored in the memory management module, and the correspondence table includes multiple correspondences, each of which is used to indicate the corresponding processing core and memory pool, respectively, and is recorded as coreid and poolid. Exemplarily, the correspondence table includes N correspondences, namely, corresponding coreid=core1 and poolid=pool11, corresponding coreid=core2 and poolid=pool22, and corresponding coreid=coreN and poolid=poolNN. This is not limited here. The memory management module may also store the status of each memory pool in multiple memory pools, and the status of the memory pool may be used or available.

[0064] In some possible implementations, in the correspondence table, the processing cores and memory pools can have a one-to-many relationship. For example, the memory pools corresponding to core1 are pool11, pool12, pool13, ..., pool1M, and the memory pools corresponding to core2 are pool21, pool22, pool23, ..., pool2M. This is not limited here.

[0065] For example, if the identification of the processing core is coreid=core1, then the corresponding memory pool is poolid=pool11; if the identification of the processing core is coreid=core2, then the corresponding memory pool is poolid=pool21; if the identification of the destination processing core is coreid=core3, then the corresponding memory pool is poolid=pool31.

[0066] In some possible implementations, the message receiving module may transmit a memory request to the memory management module, where the memory request includes an identifier of the destination processing core. The memory management module may then determine, based on a correspondence table, one or more memory pools corresponding to the destination processing core, select a memory pool with an available status from the one or more memory pools as the destination memory pool, transmit the identifier of the destination memory pool to the message receiving module, and then set the status of the destination memory pool to used. This is not limited here.

[0067] 304. The message receiving module stores the to-be-processed data of the first data message in the destination memory pool.

[0068] In some possible implementations, the first data message includes data to be processed, and the data to be processed may be information in a header of the first data message, such as a destination IP address. For example, if the first data message is a data message based on Internet Protocol version 6 (IPv6), and the header of the first data message is a plurality of IPv6 addresses stacked sequentially, the data to be processed may be the plurality of IPv6 addresses. In some possible implementations, the data to be processed may be a payload in the first data message, which is not limited herein.

[0069] In an embodiment of the present application, after the destination memory pool is determined, the data to be processed of the first data packet can be stored in the destination memory pool. For example, if the information of the destination memory pool is poolid=pool11, the data to be processed of the first data packet is stored in the cache space of the destination memory pool, and the cache space is recorded as buf1; if the information of the destination memory pool is poolid=pool21, the data to be processed of the first data packet is stored in the cache space of the destination memory pool, and the cache space is recorded as buf2; if the information of the destination memory pool is poolid=pool31, the data to be processed of the first data packet is stored in the cache space of the destination memory pool, and the cache space is recorded as buf3.

[0070] 305. The message receiving module transmits a message body to the scheduling module. The message body includes an operation instruction, information of a destination memory pool, and information of a destination processing core.

[0071] It should be noted that the network device has a built-in scheduling module, which is a software entity in the network device. In an embodiment of the present application, after the message receiving module obtains the identifier of the destination processing core and the identifier of the destination memory pool, the message receiving module can pass a message body to the scheduling module, where the message body includes the operation instruction, the identifier of the destination memory pool, and the identifier of the destination processing core. The operation instruction is used to instruct the execution of the corresponding read and write operation.

[0072] Exemplarily, the destination processing core is identified as coreid=core1, and the destination memory pool is identified as poolid=pool11; or, the destination processing core is identified as coreid=core1, and the destination memory pool is identified as poolid=pool21; or, the destination processing core is identified as coreid=core3, and the destination memory pool is identified as poolid=pool31.

[0073] 306. The scheduling module transmits an operation instruction to the destination processing core, where the operation instruction is used to instruct to perform corresponding read and write operations on the to-be-processed data in the destination memory pool.

[0074] In an embodiment of the present application, the scheduling module transmits an operation instruction to the destination processing core, and the operation instruction is used to instruct corresponding read and write operations on the data to be processed in the destination memory pool. The read and write operations can be read operations or write operations, which are not limited here.

[0075] 307. The destination processing core performs the read and write operations on the data to be processed based on the operation instruction to obtain processed data.

[0076] In an embodiment of the present application, after the destination processing core receives the operation instruction and the identifier of the destination memory pool, the destination processing core can perform read and write operations on the data to be processed. If the data to be processed is in the cache of the destination processing core, the destination processing core can process the data to be processed in the cache. If the data to be processed is not in the cache of the destination processing core, the destination processing core can obtain the data to be processed from the destination memory pool, and store the data to be processed in the cache (for example, L1, L2, L3), and perform the read and write operations on the data to be processed in the cache based on the operation instruction, and subsequently only needs to read the data to be processed in the cache.

[0077] If the read / write operation is a read operation, the value of the data to be processed will not be changed, that is, the value of the data to be processed is the same as the value of the processed data; if the read / write operation is a write operation, the value of the data to be processed can be changed, that is, the value of the data to be processed is different from the value of the processed data.

[0078] Exemplarily, the data to be processed may be the multiple IPv6 addresses, and the operation instruction may be operations such as popping or pushing the multiple IPv6 addresses, which is not limited here.

[0079] For example, if the read / write operation is a write operation, the received destination memory pool information is pool11, and the data to be processed is recorded as buf1 in pool11. For example, buf1 records the variable x=x0. Then core1 can perform the write operation on buf1 in pool11, changing the variable x=x0 to the variable x=x1, that is, the processed data is the variable x=x1. If the read / write operation is a write operation, the received destination memory pool information is pool21, and the data to be processed is recorded as buf2 in pool21. For example, buf2 records If the variable y=y0, core2 can perform the write operation on buf2 in pool21, changing the variable y=y0 to the variable y=y2, that is, the processed data is variable y=y2; if the read / write operation is a write operation, the received destination memory pool information is pool31, and the data to be processed is recorded as buf3 in pool31. For example, buf3 records the variable z=z0, then core3 can perform the write operation on buf3 in pool31, changing the variable z=z0 to the variable z=z3, that is, the processed data is variable z=z3.

[0080] 308. The destination processing core transmits the processed data to the message sending module.

[0081] In an embodiment of the present application, when the destination processing core performs read and write operations on the data to be processed and obtains the processed data, the processed data can be passed to the message sending module.

[0082] For example, after core1 changes the variable x=x0 to the variable x=x1, it can pass the variable x=x1 to the message sending module; after core2 changes the variable y=y0 to the variable y=y2, it can pass the variable y=y2 to the message sending module; after core3 changes the variable z=z0 to the variable z=z3, it can pass the variable z=z3 to the message sending module.

[0083] 309. The message sending module sends a second data message, where the second data message includes the processed data.

[0084] In an embodiment of the present application, when the message sending module receives the processed data, the message sending module generates a second data message, the second data message carries the processed data, carries it in the second data message, and sends the second data message. The data in the second data message other than the processed data is consistent with the data in the first data message other than the data to be processed. Exemplarily, when a network device receives a first data message sent by a previous hop, it can modify one of the multiple IPv6s pushed therein to obtain a second data message, and send the second data message to the next hop.

[0085] 310. The message sending module releases the destination memory pool through the memory management module.

[0086] In an embodiment of the present application, after the message sending module sends the second data message, if there is no need to continue storing the processed data or the data to be processed in the destination memory pool, the message sending module can release the destination memory pool through the memory management module. Exemplarily, the message sending module can transmit a memory release instruction to the memory management module, and the memory release instruction is used to instruct the memory management module to set the status of the destination memory pool to available. Then, the memory management module deletes the data to be processed in the destination memory pool based on the memory release instruction and sets the status of the destination memory pool to available.

[0087] Exemplarily, the destination memory pool is pool11. When the memory management module receives a memory release instruction, the memory management module can delete the pending data in pool11 and set the status of the destination memory pool pool11 indicated by the memory release instruction to available; the destination memory pool is pool21. When the memory management module receives a memory release instruction, the memory management module can delete the pending data in pool11 and set the status of the destination memory pool pool21 indicated by the memory release instruction to available; the destination memory pool is pool31. When the memory management module receives a memory release instruction, the memory management module can delete the pending data in pool11 and set the status of the destination memory pool pool31 indicated by the memory release instruction to available.

[0088] In an embodiment of the present application, after receiving the first data packet, an application can be made for a destination memory pool corresponding to the destination processing core. The destination memory pool is one of at least one memory pool corresponding to the target processing core, and the data to be processed of the first data packet is stored in the destination memory pool. Since other processing cores other than the destination processing core will not perform read and write operations on the data to be processed in the destination memory pool, false sharing between multiple processing cores is avoided, thereby improving memory access efficiency, reducing memory access latency, and increasing throughput.

[0089] For example, as shown in FIG3-2 , it is a flow chart of the network device receiving multiple data packets and sending multiple data packets.

[0090] The network device can receive message 1 through the message receiving module. The message receiving module determines the processing core core 1 corresponding to message 1 and interacts with the memory management module to determine the memory pool pool 11 corresponding to core 1. The message receiving module stores the to-be-processed data of message 1 in pool 11, namely buf1. Next, the message receiving module passes the message body to the scheduling module. The message body includes an operation instruction, pool 11, and core 1. The scheduling module then passes the operation instruction to core 1 to instruct core 1 to perform the read and write operations corresponding to the buf1 row in pool 11. Core 1 performs the read and write operations on buf1 in pool 11, obtains the processed buf1, and passes the processed buf1 to the message sending module. Then, the message sending module can send message 1', which includes the processed buf1. Finally, the message sending module can release pool 11.

[0091] Similarly, the network device can receive message 2 through the message receiving module. The message receiving module determines the processing core core2 corresponding to message 2 and interacts with the memory management module to determine the memory pool pool22 corresponding to core2. The message receiving module stores buf2 of message 2 in pool22, which is buf2. Next, the message receiving module passes the message body to the scheduling module. The message body includes an operation instruction, pool22, and core2. The scheduling module then passes the operation instruction to core2 to instruct core2 to perform the corresponding read and write operations on buf2 in pool22. Core2 performs the read and write operations on buf2 in pool2, obtains the processed buf2, and then passes the processed buf2 to the message sending module. Then, the message sending module can send message 2', which includes the processed buf2. Finally, the message sending module can release pool22.

[0092] Similarly, the network device can receive message 3 through the message receiving module. The message receiving module determines the processing core core 3 corresponding to message 3 and interacts with the memory management module to determine the memory pool pool 33 corresponding to core 3. The message receiving module then stores buf3 of message 3 in pool 33, which is buf3. Next, the message receiving module passes the message body to the scheduling module. The message body includes an operation instruction, pool 33, and core 3. The scheduling module then passes the operation instruction to core 3, instructing core 3 to perform the corresponding read and write operations on buf3 in pool 33. Core 3 performs the read and write operations on buf3 in pool 3, obtains the processed buf3, and then passes the processed buf3 to the message sending module. Then, the message sending module can send message 3', which includes the processed buf3. Finally, the message sending module can release pool 33.

[0093] Then, since core1 can process data on buf1 in pool11, core2 can process data on buf3 in pool2, and core3 can process data on buf3 in pool2, competition among core1, core2, and core3 for the memory pool is avoided, and false sharing among multiple processing cores is avoided, thereby improving memory access efficiency, reducing memory access latency, and increasing throughput.

[0094] It should be noted that the processing core's access to the data to be processed in the memory needs to go through multiple levels of cache. If the data to be processed is in the cache of the processing core, the data to be processed can be read in the cache, which is called a cache hit. Otherwise, it is a cache miss. If a cache miss occurs, it is necessary to search for the data to be processed in the cache of the next layer of the processing core. If the cache of the last layer does not find the data to be processed, it needs to be copied from the memory to the cache. There is a greater delay in accessing data in the memory than in the cache of the processing core. For this reason, the industry usually uses prefetching technology for memory access, which includes hardware prefetching and software prefetching. Currently, the industry mainly uses hardware prefetching, and a hardware prefetcher is usually provided in the network equipment.

[0095] If the memory management module doesn't have a mapping table, the hardware prefetcher will pre-set some memory pools as exclusive to the destination core based on the core's read and write history. This means these memory pools are set to "used." However, the destination core may not subsequently use these memory pools, wasting memory pool resources, reducing memory access efficiency, increasing access latency, and lowering throughput.

[0096] To this end, as shown in FIG4-1, a data processing method is provided in the second embodiment of the present application. The method mainly includes the following steps:

[0097] 401. The destination processing core disables the hardware prefetch function.

[0098] It should be noted that the hardware prefetch function in the destination processing core is implemented by a hardware prefetcher. Therefore, the destination processing core can disable the hardware prefetcher before working, thereby disabling the hardware prefetch function.

[0099] 402. The message receiving module receives a first data message.

[0100] 403. The message receiving module determines a destination processing core corresponding to the first data message.

[0101] Steps 402 to 403 are the same as steps 301 to 302 and are not described in detail here.

[0102] 404. The message receiving module applies for a destination memory pool corresponding to the destination processing core through the memory management module.

[0103] In an embodiment of the present application, the storage space in the memory can be divided into multiple memory pools, and each of the multiple memory pools is assigned an identifier, denoted as poolid. For example, the identifiers of the multiple memory pools are pool11, pool22, ..., poolNN.

[0104] In an embodiment of the present application, a network device may include a built-in memory management module, which is a software entity and is used to manage storage space in the memory. The memory management module may store multiple items, each item indicating the identifier and status of a corresponding memory pool, where the status of the memory pool is either used or available.

[0105] In some possible implementations, the message receiving module may send a memory application request to the memory management module, and the memory management module may determine a memory pool from multiple memory pools as a destination memory pool and allocate the destination memory pool to the message receiving module.

[0106] 405. The scheduling module transmits an independent storage request to the destination processing core, where the independent storage request is used to instruct the destination processing core to exclusively access the destination memory pool.

[0107] In an embodiment of the present application, the scheduling module may initiate a unique storage request, i.e., the scheduling module instructs the other processing cores in the plurality of processing cores except the destination processing core to have exclusive access to the destination memory pool. In addition, the scheduling module further transmits a Read / WriteOnceCleanInvalid message to the other processing cores in the plurality of processing cores except the destination processing core, indicating that the destination memory pool is an exclusive memory pool for the destination processing core and is invalid for the other processing cores in the plurality of processing cores except the destination processing core.

[0108] Exemplarily, the scheduling module sends an independent storage request to core1, the independent storage request including the operation instruction and pool11. The scheduling module also sends a read / write clear invalidation message to core2, core3, ..., coreN, the clear invalidation message including pool11, indicating that pool11 is invalid for core2, core3, ..., coreN.

[0109] 406. The message receiving module stores the to-be-processed data of the first data message in the destination memory pool.

[0110] 407. The message receiving module sends a message body to the scheduling module. The message body includes an operation instruction, information about a destination memory pool, and information about a destination processing core.

[0111] 408. The scheduling module transmits an operation instruction to the destination processing core, where the operation instruction is used to instruct to perform corresponding read and write operations on the to-be-processed data in the destination memory pool.

[0112] 409. The destination processing core performs the read and write operations on the data to be processed in the destination memory pool based on the operation instruction to obtain processed data.

[0113] 410. The destination processing core transmits the processed data to the message sending module.

[0114] 411. The message sending module sends a second data message, where the second data message includes the processed data.

[0115] 412. The message sending module releases the destination memory pool through the memory management module.

[0116] Steps 406 to 411 are the same as steps 304 to 310 and are not described in detail here.

[0117] In the second embodiment of the present application, when a corresponding table is not provided in the memory management module, the hardware prefetch function of the hardware prefetcher can be disabled in advance to prevent the hardware prefetcher from pre-setting some memory pools as exclusive memory pools for the destination processing core. At the same time, by transmitting an independent storage request to the destination processing core, the independent storage request is used to instruct the destination processing core to exclusively access the destination memory pool, thereby avoiding competition among other processing cores for the destination memory pool and avoiding false sharing among multiple processing cores, thereby improving memory access efficiency, reducing memory access latency, and increasing throughput.

[0118] For example, as shown in FIG4-2 , it is a flow chart of the network device receiving multiple data packets and sending multiple data packets.

[0119] The network device can receive message 1 through a message receiving module. The message receiving module determines the processing core core 1 corresponding to message 1 and interacts with the memory management module to determine the memory pool pool 11 corresponding to core 1. The scheduling module can then issue a unique storage request (stashonceunique), instructing core 1 to exclusively access pool 11. The message receiving module then stores the pending data of message 1 in pool 11, designated as buf1. The message receiving module then transmits a message body to the scheduling module, which includes an operation instruction, pool 11, and core 1. The scheduling module then transmits the operation instruction to core 1, instructing core 1 to perform the corresponding read and write operations on buf1 in pool 11. Core 1 performs the read and write operations on buf1 in pool 11, obtains the processed buf1, and then transmits the processed buf1 to the message sending module. The message sending module can then transmit message 1', which includes the processed buf1. Finally, the message sending module may release pool11.

[0120] Similarly, the network device can receive message 2 through the message receiving module. The message receiving module determines the processing core core 2 corresponding to message 2 and interacts with the memory management module to determine the memory pool pool 22 corresponding to core 2. The scheduling module can then issue an independent storage request (stashonceunique), that is, the scheduling module transmits an independent storage request to core 2 and core 3, instructing core 2 to exclusively access pool 22. The message receiving module then stores the pending data of message 2 in pool 22, namely, buf2. The message receiving module then transmits a message body to the scheduling module, which includes an operation instruction, pool 22, and core 2. The scheduling module then transmits the operation instruction to core 2, instructing core 2 to perform corresponding read and write operations on buf2 in pool 22. Core 2 performs the read and write operations on buf2 in pool 22, obtains the processed buf2, and then transmits the processed buf2 to the message sending module. Then, the message sending module can send message 2', which includes the processed buf2. Finally, the message sending module can release pool22.

[0121] Similarly, the network device can receive message 3 through the message receiving module. The message receiving module determines the processing core core 3 corresponding to message 3 and interacts with the memory management module to determine the memory pool pool 33 corresponding to core 3. The scheduling module can then issue a unique storage request (stashonceunique), namely, the scheduling module transmits the unique storage request to core 2 and core 3, instructing core 3 to exclusively access pool 33. The message receiving module then stores the pending data of message 3 in pool 33, namely, buf3. The message receiving module then transmits a message body to the scheduling module, which includes an operation instruction, pool 33, and core 3. The scheduling module then transmits the operation instruction to core 3, instructing core 3 to perform the corresponding read and write operations on buf3 in pool 33. Core 3 performs the read and write operations on buf3 in pool 33, obtains the processed buf3, and then transmits the processed buf3 to the message sending module. Then, the message sending module can send message 3', which includes the processed buf3. Finally, the message sending module can release pool33.

[0122] Through the technical solution of Example 1, this application achieves that each memory pool corresponds to a unique processing core, avoiding the situation where different processing cores compete for the same memory pool. In the case of large-scale flows, data packets of different data flows can be accurately distributed to the processing core, thereby achieving performance gains. Through the technical solution of Example 2, this application achieves exclusive occupation of the destination memory pool by the destination processing core, thereby achieving performance gains.

[0123] As shown in Table 1, in the case of simulation verification, the performance of the existing technical solution, the technical solution of Example 1, and the actual solution of Example 2 are compared in various situations (whether the hardware prefetch function is enabled, whether Stashonceunique is enabled, and in a single core or 22 cores) to determine the performance gain of the technical solution of the present application.

[0124] Table 1

[0125] Among them, in the technical solution of embodiment 1, when the hardware prefetch function is enabled and Stashonceunique is not enabled, it can bring a performance improvement of 9%-24.6% compared with the existing technical solution. This is the performance benefit generated by avoiding false sharing.

[0126] In the technical solution of Example 2, the quantitative analysis of the benefits is as follows:

[0127] In case 1, when hardware prefetching is enabled, enabling Stashonceunique exacerbates contention for the memory pool among multiple processing cores. Performance degrades by 10.2% on 22 cores. This indicates that the hardware prefetcher is performing inappropriate prefetching, resulting in false sharing.

[0128] In case 2, when Stashonceunique is not enabled and the hardware prefetch function is disabled, the performance is improved by 15% with 22 cores.

[0129] In case 3, when the hardware prefetch function is disabled and Stashonceunique is enabled, the performance is further improved by 3.57% to 10% compared with case 2.

[0130] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0131] In order to better implement the above-mentioned solutions of the embodiments of the present application, relevant devices for implementing the above-mentioned solutions are also provided below.

[0132] Referring to FIG. 5 , a processor 500 provided in an embodiment of the present application may include:

[0133] The processing module 501 is configured to obtain a first data message;

[0134] The processing module 501 is further configured to determine a destination processing core corresponding to the first data message, where the destination processing core is one of the multiple processing cores, and each of the multiple processing cores corresponds to at least one memory pool.

[0135] The processing module 501 is further configured to apply for a destination memory pool corresponding to the destination processing core, where the destination memory pool is one of at least one memory pool corresponding to the destination processing core;

[0136] The storage module 502 is configured to store the to-be-processed data of the first data message in the destination memory pool.

[0137] In some possible implementations, the processing module 501 is specifically configured to:

[0138] Obtaining queue status of the plurality of processing cores, wherein the queue status is the amount of data to be processed in the corresponding processing cores;

[0139] The destination processing core is determined based on a preset scheduling policy and queue status of the multiple processing cores, the destination processing core being one of the multiple processing cores, and the scheduling policy is to give priority to selecting the processing core with the least amount of data to be processed among the multiple processing cores.

[0140] In some possible implementations, the processing module 501 is specifically configured to:

[0141] The destination memory pool corresponding to the destination processing core is determined from a preset correspondence table, where the correspondence table includes a plurality of correspondences, and the plurality of correspondences include each processing core in the plurality of processing cores and at least one corresponding memory pool.

[0142] In some possible implementations, the processing module 501 is further configured to:

[0143] Turn off the hardware prefetch function of the destination processing core;

[0144] An independent storage request is transmitted to the destination processing core, where the independent storage request is used to instruct the destination processing core to exclusively access the destination memory pool.

[0145] In some possible implementations, the processing module 501 is further configured to:

[0146] The destination processing core performs read and write operations on the data to be processed in the destination memory pool to obtain processed data.

[0147] In some possible implementations, the processing module 501 is further configured to:

[0148] generating a second data packet, wherein the second data packet includes the processed data;

[0149] After the second data message is generated, the destination memory pool is released.

[0150] It should be noted that the information interaction, execution process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and no further details will be given here.

[0151] An embodiment of the present application further provides a computer storage medium, wherein the computer storage medium stores a program, and the program executes some or all of the steps recorded in the above method embodiment.

[0152] Next, another communication device provided in an embodiment of the present application is introduced. Referring to FIG6 , the communication device 600 includes:

[0153] Receiver 601, transmitter 602, processor 603 and memory 604. In some embodiments of the present application, the receiver 601, transmitter 602, processor 603 and memory 604 may be connected via a bus or other means, wherein FIG6 takes the bus connection as an example.

[0154] The memory 604 may include a read-only memory and a random access memory, and provides instructions and data to the processor 603. A portion of the memory 604 may also include non-volatile random access memory (NVRAM). The memory 604 stores an operating system and operating instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. The operating instructions may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks.

[0155] Processor 603 controls the operation of communication device 600 and may also be referred to as a central processing unit (CPU). In specific applications, the various components of communication device 600 are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0156] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 603. Processor 603 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 603. The above processor 603 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can be located in storage media such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other well-known storage media in the art. The storage medium is located in the memory 604 , and the processor 603 reads the information in the memory 604 and completes the steps of the above method in combination with its hardware.

[0157] The receiver 601 can be used to receive input digital or character information and generate signal input related to relevant settings and function control. The transmitter 602 can include a display device such as a display screen. The transmitter 602 can be used to output digital or character information through an external interface.

[0158] In the embodiment of the present application, the processor 603 is used to execute the aforementioned data processing method.

[0159] In another possible design, when the network device 500 or the communication device 600 is a chip, it includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit may execute computer-executable instructions stored in the storage unit to enable the chip in the terminal to execute the method for sending wireless report information of any one of the above-mentioned first aspects. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit in the terminal located outside the chip, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0160] The processor mentioned in any of the above may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the above method.

[0161] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0162] Through the description of the above embodiments, it is clear to those skilled in the art that the present application can be implemented by means of software plus necessary general-purpose hardware, and of course it can also be implemented by means of dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0163] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0164] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, a computer, a server, or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium, (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).

Claims

1. A data processing method, characterized in that, For a processor, the processor includes a plurality of processing cores, including: Obtain a first data packet; Determine a destination processing core corresponding to the first data packet, the destination processing core being one of the plurality of processing cores, and each of the plurality of processing cores corresponding to at least one memory pool; Apply for a destination memory pool corresponding to the destination processing core, the target memory pool being one of the at least one memory pool corresponding to the target processing core; Store the data to be processed of the first data packet in the destination memory pool.

2. The method according to claim 1, wherein The determining the destination processing core corresponding to the first data packet includes: Obtain the queue status of the plurality of processing cores, the queue status being the number of data to be processed in the corresponding processing core; Determine the destination processing core based on a preset scheduling policy and the queue status of the plurality of processing cores, the destination processing core being one of the plurality of processing cores, and the scheduling policy being to preferentially select the processing core with the least number of data to be processed among the plurality of processing cores.

3. The method according to claim 1 or 2, characterized in that, The applying for a destination memory pool corresponding to the destination processing core includes: Determine the destination memory pool corresponding to the destination processing core from a preset correspondence table, the correspondence table including a plurality of correspondence relationships, the plurality of correspondence relationships including each of the plurality of processing cores and the corresponding at least one memory pool.

4. The method according to claim 1 or 2, characterized in that, Before receiving the first data packet, further includes: Turn off the hardware prefetch function of the destination processing core; After applying for the destination memory pool corresponding to the destination processing core, further includes: Transmit an independent storage request to the destination processing core, the independent storage request being used to instruct the destination processing core to exclusively access the destination memory pool.

5. The method according to any one of claims 1-4, characterized in that After storing the data to be processed of the first data packet in the destination memory pool, the method further includes: Perform read and write operations on the data to be processed in the destination memory pool through the destination processing core to obtain processed data.

6. The method according to claim 5, wherein After performing the read and write operations on the data to be processed in the destination memory pool through the destination processing core, the method further includes: Generate a second data packet, the second data packet including the processed data; Release the destination memory pool after generating the second data packet.

7. A processor, characterized in that, Includes: A processing module, configured to obtain a first data packet; The processing module is further configured to determine a destination processing core corresponding to the first data packet, the destination processing core being one of the plurality of processing cores, and each of the plurality of processing cores corresponding to at least one memory pool; The processing module is further configured to apply for a destination memory pool corresponding to the destination processing core, the target memory pool being one of the at least one memory pool corresponding to the target processing core; The storage module is configured to store the data to be processed of the first data packet in the destination memory pool.

8. The processor according to claim 7, wherein The processing module is specifically configured to: Obtain the queue status of the plurality of processing cores, the queue status being the number of data to be processed in the corresponding processing core; Determine the destination processing core based on a preset scheduling policy and the queue states of the multiple processing cores, where the destination processing core is one of the multiple processing cores, and the scheduling policy is to preferentially select the processing core with the smallest number of pending data among the multiple processing cores.

9. The processor according to claim 7 or 8, characterized in that The processing module is specifically configured to: Determine the destination memory pool corresponding to the destination processing core from a preset correspondence table, where the correspondence table includes multiple correspondence relationships, and the multiple correspondence relationships include each of the multiple processing cores and at least one corresponding memory pool.

10. The processor according to claim 7 or 8, characterized in that, The processing module is further configured to: Disable the hardware prefetch function of the destination processing core; Transmit an independent storage request to the destination processing core, where the independent storage request is used to instruct the destination processing core to exclusively access the destination memory pool.

11. The processor according to any one of claims 7-10, characterized in that, The processing module is further configured to: Perform read and write operations on the pending data in the destination memory pool through the destination processing core to obtain processed data.

12. The processor according to claim 11, wherein The processing module is further configured to: Generate a second data packet, where the second data packet includes the processed data; After generating the second data packet, release the destination memory pool.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program, and the program causes the computer device to execute the method according to any one of claims 1-6.

14. A computer program product, characterized in that, The computer program product includes computer-executable instructions, and the computer-executable instructions are stored in a computer-readable storage medium; at least one processor of the device reads the computer-executable instructions from the computer-readable storage medium, and the at least one processor executes the computer-executable instructions to cause the device to execute the method according to any one of claims 1-6.

15. A communication device, characterized in that, The communication device includes at least one processor, a memory, and a communication interface; The at least one processor is coupled to the memory and the communication interface; The memory is used to store instructions, the processor is used to execute the instructions, and the communication interface is used to communicate with other communication devices under the control of the at least one processor; When the instructions are executed by the at least one processor, the at least one processor is caused to execute the method according to any one of claims 1-6.

16. A chip system, characterized in that, The chip system includes a processor and a memory, the memory and the processor are interconnected by a line, the memory stores instructions, and the processor is used to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Memory access method and device of processor

    CN112231099A

  • Resource management method and corresponding device

    CN115729694A

  • Memory Pool Allocation for a Multi-Core System

    US20190347133A1

  • Memory pool management

    US20220050722A1

  • Memory management method and system, client, server and storage medium

    WO2021254330A1