A data processing system, method, device, medium and program product
By building a system with processors and switching devices, the problems of insufficient host memory and computing power were solved, and unified management of memory and computing power and improved communication efficiency were achieved.
Patent Information
- Application Number
- CN202511384938.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Current technology cannot simultaneously solve the problems of insufficient host memory and insufficient computing power.
By constructing a system with multiple processors, a first switching device, and a second switching device, communication links and resource allocation between any processor and the second device can be achieved, including memory and computing power expansion.
It enables unified management of memory space and computing power of multiple processors and second devices, solving the problem of insufficient memory and computing power, and improving communication reliability and flexibility.
Smart Images

Figure CN120872619B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a data processing system, method, device, medium and program product. BACKGROUND
[0002] At present, memory expansion can be realized for a host by using a switching device, for example, the host is connected with multiple memory devices through the switching device, and when the memory of the host is insufficient, the host can use the storage space in any memory device. However, this method only realizes memory expansion and cannot solve the problem of insufficient computing power of the host.
[0003] Therefore, how to simultaneously solve the problems of insufficient memory and insufficient computing power of the host is a problem to be solved by those skilled in the art. SUMMARY
[0004] Therefore, the present application aims to provide a data processing system, method, device, medium and program product to simultaneously solve the problems of insufficient memory and insufficient computing power of the host.
[0005] In a first aspect, the present application provides a data processing system, comprising: multiple processors, a first switching device, multiple second devices and a second switching device; the multiple processors and the multiple second devices are connected with the first switching device; the multiple processors and the multiple second devices are connected with the second switching device; the first switching device is configured to: construct a first communication link between any processor and any second device; select a target resource matching a resource requirement sent by any processor from other devices according to the resource requirement; the other devices include at least one or a combination of the multiple second devices and other processors except the processor sending the resource requirement; the target resource includes at least one or a combination of a memory resource and a computing resource; and the second switching device is configured to: construct a second communication link between any processor and any second device, communication between any processor and an external network, and communication between any second device and the external network.
[0006] In a second aspect, the present application provides a data processing method applied to a data processing system, the data processing system comprising: a plurality of processors, a first switching device, a plurality of second devices and a second switching device; the plurality of processors and the plurality of second devices are connected to the first switching device; the plurality of processors and the plurality of second devices are connected to the second switching device; the first switching device is configured to: construct a first communication link between any processor and any second device; select a target resource matching a resource requirement sent by any processor from other devices according to the resource requirement; the other devices include at least one or a combination of the plurality of second devices and other processors except the processor sending the resource requirement; the target resource includes at least one or a combination of a memory resource and a computing resource; the second switching device is configured to: construct a second communication link between any processor and any second device, communication between any processor and an external network, and communication between any second device and the external network.
[0007] Correspondingly, the data processing method comprises: any processor or any second device implements memory expansion through the first switching device when the memory is insufficient; and any processor implements computing power expansion through the first switching device when the computing power is insufficient.
[0008] In a third aspect, the present application provides an electronic device comprising: a memory configured to store a computer program; and a processor configured to execute the computer program to implement the data processing method disclosed above.
[0009] In a fourth aspect, the present application provides a non-volatile storage medium configured to store a computer program, wherein the computer program is executed by a processor to implement the data processing method disclosed above.
[0010] In a fifth aspect, the present application provides a computer program product comprising computer programs / instructions, which are executed by a processor to implement the steps of the data processing method disclosed above.
[0011] According to the above scheme, the data processing system provided by the application comprises: a plurality of processors, a first switching device, a plurality of second devices and a second switching device; the plurality of processors and the plurality of second devices are connected to the first switching device; the plurality of processors and the plurality of second devices are connected to the second switching device; the first switching device is configured to: construct a first communication link between any processor and any second device; select a target resource matching a resource requirement sent by any processor from other devices according to the resource requirement; the other devices include at least one or a combination of the plurality of second devices and other processors except the processor sending the resource requirement; the target resource includes at least one or a combination of a memory resource and a computing resource; and the second switching device is configured to: construct a second communication link between any processor and any second device, communication between any processor and an external network, and communication between any second device and the external network.
[0012] It can be seen that the application has the following advantages: the memory space of the plurality of processors and the memory space of the plurality of second devices, and the computing power of the plurality of processors and the computing power of the plurality of second devices are uniformly managed by the first switching device, so that memory allocation and computing power allocation can be realized by the first switching device regardless of which device is insufficient in memory or computing power, and the problems of insufficient memory and insufficient computing power are solved; at the same time, the first communication link is provided between any processor and any second device by the first switching device, and the second communication link is provided between any processor and any second device by the second switching device, the first communication link and the second communication link can be used in parallel to improve transmission efficiency, and the transmission link redundancy and fault tolerance between any processor and any second device are increased, thereby providing protection for communication reliability and communication efficiency; further, any processor and any second device can realize communication with the external network through the second switching device, thereby removing the limitation that only the host device can communicate with the external network, and improving communication flexibility.
[0013] Correspondingly, the data processing method, device, medium and program product provided by the application also have the above technical effects. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of the provided drawings.
[0015] Figure 1 The first data processing system disclosed by the application is shown in the figure;
[0016] Figure 2A second data processing system schematic diagram disclosed by the present application is shown in FIG. 2.
[0017] Figure 3 A CPU computing card structure schematic diagram disclosed by the present application is shown in FIG. 3.
[0018] Figure 4 An FPGA acceleration card structure schematic diagram disclosed by the present application is shown in FIG. 4.
[0019] Figure 5 A CXL switch structure schematic diagram disclosed by the present application is shown in FIG. 5.
[0020] Figure 6 A data processing method flowchart disclosed by the present application is shown in FIG. 6.
[0021] Figure 7 A server structure diagram provided by the present application is shown in FIG. 7.
[0022] Figure 8 A terminal structure diagram provided by the present application is shown in FIG. 8. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0024] It should be noted that in the description of the present application, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices comprising a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0025] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0026] At present, memory expansion can be realized for a host by using a switching device (such as a Switch), for example, the host connects multiple memory devices through the switching device, and when the host memory is insufficient, the host can use the storage space in any memory device. However, this method only realizes memory expansion and cannot solve the problem of insufficient computing power of the host. Therefore, the present application provides a data processing scheme, which can simultaneously solve the problems of insufficient host memory and insufficient computing power.
[0027] Referring to Figure 1 As shown in the drawings, the embodiments of the present application disclose a data processing system, comprising: a plurality of processors, a first switching device, a plurality of second devices and a second switching device; the plurality of processors and the plurality of second devices are connected to the first switching device; the plurality of processors and the plurality of second devices are connected to the second switching device.
[0028] In an example, the first switching device is configured to: construct a first communication link between any processor and any second device; and implement memory management and computing power allocation of any processor or any second device, specifically, according to resource requirements sent by any processor, select target resources matching the resource requirements from other devices; the other devices include at least one or a combination of the plurality of second devices and other processors except the processor sending the resource requirements; the target resources include at least one or a combination of memory resources and computing resources. For example, if processor A is short of memory, processor A sends resource requirements for how much memory to the first switching device, and applies for memory of other processors or any second device to processor A through the first switching device; for another example, if processor A is short of computing power, processor A sends resource requirements for how much computing power to the first switching device, and applies for computing power of other processors or any second device to processor A through the first switching device.
[0029] The second switching device is configured to: construct a second communication link between any processor and any second device, communication between any processor and an external network, and communication between any second device and the external network.
[0030] In an example, the memory resources of the plurality of second devices can constitute a memory pool, and the plurality of processors share the memory pool.
[0031] In this embodiment, the processor can be a processor computing card, and the second device can be an acceleration card such as an FPGA; the first switching device can be an expansion card such as a Switch, and the second switching device can be an optical switch or the like. That is, both the processor and the second device have their own memory space and their own computing cores; both the first switching device and the second switching device can realize communication between any processor and any second device, but the first communication link and the second communication link can follow different transmission technologies or transmission protocols, such as the first communication link supporting CXL, PCIe (Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard), and the second communication link being an optical path; of course, the first communication link and the second communication link can also follow the same transmission technology or transmission protocol, such as the first communication link and the second communication link both supporting the CXL protocol and the PCIe protocol. CXL (Compute Express Link) is a high-speed interface protocol that can optimize the interaction between computing, storage, and communication resources in a data center. CXL actually consists of three sub-protocols, namely CXL.io, CXL.cache, and CXL.mem. CXL.io is used for initialization, linking, device identification and enumeration, and register access, and provides a non-coherent load / store interface for devices. CXL.cache is used to access the cache and can define the interaction between the processor and the device, allowing the connected CXL device to use the request and response method to efficiently cache the processor memory with extremely low latency. CXL.mem is used to access the memory and provides access to the device's attached memory for the processor using load and store commands, where the processor acts as the master device and the CXL device acts as the slave device (such as the second device), and can support volatile and persistent memory architectures. After these protocols are dynamically multiplexed together, data transmission can be performed at a speed of 32 GT / s through the standard PCIe 5.0 physical layer.
[0032] As described above, the processor and the second device each have their own memory space and their own computing core, and the memory expansion or computing power expansion of the processor or the second device can be performed through the first exchange device. For example, if the memory of a certain processor A is insufficient, the memory of other processors or any second device can be applied to processor A through the first exchange device X for use by processor A. For another example, if the computing power of a certain processor A is insufficient, the computing power of other processors or any second device can be applied to processor A through the first exchange device X for use by processor A. In theory, the processor and the second device can have no primary and secondary distinction, and then there is: if the memory of a certain second device B is insufficient, the memory of other second devices or any processor can be applied to second device B through the first exchange device X for use by second device B. For another example, if the computing power of a certain second device B is insufficient, the computing power of other second devices or any processor can be applied to second device B through the first exchange device X for use by second device B. Correspondingly, the first exchange device and the second exchange device can also have no functional difference, that is, the first exchange device X in the example can be changed to the second exchange device Y accordingly.
[0033] Correspondingly, if the processor and the second device have primary and secondary distinction, a plurality of processors can be set as a plurality of hosts, and a plurality of second devices can be set as a plurality of memory and computing power expansion devices. Of course, a plurality of processors can also be set as a plurality of memory and computing power expansion devices, and a plurality of second devices can be set as a plurality of hosts. In this scenario, the allocation and application of memory and computing power need to be initiated by the host, and then the first exchange device or the second exchange device performs the related data exchange process. Of course, the transmission function of the first exchange device and the second exchange device can also be customized, for example, the first exchange device is used to realize the allocation and application of memory and computing power, and the transmission of original processing data between any processor and any second device, etc., and the first exchange device is used to realize the transmission of processed data (such as the processing result output by the processor or the second device after processing the original processing data) between any processor and any second device.
[0034] In this embodiment, the second exchange device can interact with the external network, and therefore in one implementation, the second exchange device is configured to: perform protocol analysis on a task (such as calculation of an algorithm, training or reasoning of a model, etc.) received from the external network, query a corresponding routing strategy according to the protocol analysis result, and transmit the task to a corresponding destination processor according to the routing strategy. Correspondingly, the second exchange device is configured to: generate a corresponding task dispatch instruction according to the routing strategy; and transmit the task to the destination processor by executing the task dispatch instruction. The routing strategy can be a set routing table, etc.
[0035] If the processor is a host class device, in an embodiment, the destination processor is configured to: extract features of the task, determine execution requirements according to the extracted features of the task, and select an execution end for the task from the plurality of processors and / or the plurality of second devices according to the execution requirements. Accordingly, the features of the task include operation logic features, complexity, parallelism, and latency sensitivity. The destination processor is configured to: determine first and second adaptabilities corresponding to the task according to the operation logic features, complexity, parallelism, and latency sensitivity, and select an execution end for the task from the plurality of processors and / or the plurality of second devices according to the first and second adaptabilities.
[0036] In an embodiment, the destination processor is configured to: calculate the first adaptability by using a first formula, wherein the first formula is Q i =a1×R+a2×S+a3×(1 / P)+a4×T; Q i is the first adaptability corresponding to the task i, R is the operation logic features, a1 is a weighting coefficient of the operation logic features in the first formula, S is the complexity, a2 is a weighting coefficient of the complexity in the first formula, P is the parallelism, a3 is a weighting coefficient of the parallelism in the first formula, and T is the latency sensitivity, a4 is a weighting coefficient of the latency sensitivity in the first formula.
[0037] In an embodiment, the destination processor is configured to: calculate the second adaptability by using a second formula, wherein the second formula is W i =b1×(1 / R)+b2×(1 / S)+b3×P+b4×(1 / T); W i is the second adaptability corresponding to the task i, R is the operation logic features, b1 is a weighting coefficient of the operation logic features in the second formula, S is the complexity, b2 is a weighting coefficient of the complexity in the second formula, P is the parallelism, b3 is a weighting coefficient of the parallelism in the second formula, and T is the latency sensitivity, b4 is a weighting coefficient of the latency sensitivity in the second formula.
[0038] In an embodiment, the destination processor is configured to: if the first adaptability is greater than the second adaptability, and the difference between the first adaptability and the second adaptability is greater than a preset first threshold, it is indicated that the task is more suitable for execution in the processor, and an execution end for the task is selected from the plurality of processors; if the first adaptability is less than the second adaptability, and the difference between the first adaptability and the second adaptability is greater than a preset second threshold, it is indicated that the task is more suitable for execution in the second device, and an execution end for the task is selected from the plurality of second devices; and if the difference between the first adaptability and the second adaptability is less than a preset third threshold, the destination processor is used as the execution end for the task under the condition that the destination processor meets the execution requirements of the task, so that the task is executed locally in the destination processor to avoid unnecessary transmission consumption.
[0039] It should be noted that, in the case that the destination processor does not meet the execution requirement of the task, the destination processor splits the task into a first sub-task and a second sub-task; the destination processor is used as an execution end of the first sub-task, and an execution end of the second sub-task is selected from the plurality of second devices or other processors except the destination processor.
[0040] In the task splitting, the task is executed locally on the destination processor as much as possible to avoid unnecessary transmission consumption. To this end, it can be considered that the execution requirement of the task includes a memory requirement and a computing power requirement; accordingly, the destination processor is used to: if the remaining memory space of the destination processor is not less than the memory requirement of the task, and the remaining computing power of the destination processor is not less than the computing power requirement of the task, it is confirmed that the destination processor meets the execution requirement of the task; otherwise, it is confirmed that the destination processor does not meet the execution requirement of the task.
[0041] In an example, the specific process of splitting the task by the destination processor includes: splitting a first sub-task from the task according to the remaining memory space of the destination processor and the remaining computing power of the destination processor; the memory required by the first sub-task is not greater than the remaining memory space of the destination processor, and the computing power required by the first sub-task is not greater than the remaining computing power of the destination processor; the remaining task amount of the task is used as a second sub-task.
[0042] It should be noted that the data transmission between any processor and any second device can include: any processor A transmits the data to be processed to the target second device through the first communication link provided by the first exchange device. The data to be processed can be part or all of the original data of the task, or the intermediate result obtained after the processor A processes part or all of the original data of the task.
[0043] Correspondingly, the data transmission between any processor and any second device can also include: the target second device processes the data to be processed to obtain a processing result; if the processing result still needs to be processed by the processor A to obtain the task response (whether processing is needed is determined by the task itself, and the processor A will send the corresponding message of the processing result to be returned to the target second device), then the processing result is transmitted to the corresponding processor through the second communication link provided by the second switching device, so that the processor A determines the response data based on the processing result, and transmits the response data (i.e. the task response) to the external network, other cabinet nodes or other servers in the same network as the processor A and the like through the second switching device. If the processing result is already the task response data (whether processing is needed is determined by the task itself, and the processor A will send the message to the target second device without processing), which does not need to be processed by the processor A, then the target second device transmits the processing result to the external network, other cabinet nodes or other servers in the same network as the processor A and the like through the second switching device without the need of forwarding by the processor A, thereby avoiding unnecessary transmission loss.
[0044] It can be seen that in the embodiment, the memory space of the plurality of processors and the memory space of the plurality of second devices, the computing power of the plurality of processors and the computing power of the plurality of second devices are uniformly managed by the first switching device. No matter which device is insufficient in memory or computing power, memory allocation and computing power allocation can be realized through the first switching device, and the problems of insufficient memory and insufficient computing power are solved. At the same time, between any processor and any second device, not only is there a first communication link provided by the first switching device, but also there is a second communication link provided by the second switching device. The first communication link and the second communication link can be used in parallel to improve transmission efficiency, and the redundancy and fault tolerance of the transmission link between any processor and any second device can be increased, thereby providing a guarantee for communication reliability and communication efficiency. Further, any processor and any second device can realize communication with the external network through the second switching device, thereby removing the limitation that only the host device can communicate with the external network, and improving the communication flexibility.
[0045] It should be noted that if the plurality of processors is a plurality of processor computing cards, and the plurality of second devices is a plurality of acceleration cards, then the plurality of processor computing cards and the plurality of acceleration cards can be plugged into the first switching device. The first switching device can be an expansion card, such as a CXL Switch expansion card. As can be seen, the plurality of processor computing cards, the plurality of acceleration cards, and the first switching device can all be in the form of a board card, and can be connected in a convenient and flexible manner through plugging, achieving plug and play. If the plurality of processors, the first switching device, and the plurality of second devices are arranged in the same cabinet node H, and the second switching device is arranged outside the cabinet node H, then the second switching device can also be connected to other cabinet nodes and / or other computer devices (such as routers, servers, etc.). The other cabinet nodes can also include a plurality of processors, a first switching device, and a plurality of second devices, that is, the other cabinet nodes are completely identical to the cabinet node H.
[0046] The second switching device can simultaneously connect other cabinet nodes and external network routers. That is, when the other cabinet nodes and the cabinet node H are in the same cluster, the second switching device can not only realize communication between the other cabinet nodes and the cabinet node H, but also realize communication between the other cabinet nodes and the external network and communication between the cabinet node H and the external network.
[0047] Please refer to Figure 2 , Figure 2 A data processing system provided includes: n CPU computing cards (processor computing cards), a CXL switch (first switching device CXL Switch), n FPGA acceleration cards, and an optical switch (second switching device); wherein the n CPU computing cards are arranged as n hosts, the n FPGA acceleration cards are shared by the hosts, and the n FPGA acceleration cards can realize computing power and memory expansion for the hosts. The optical switch realizes interaction between the n CPU computing cards, the n FPGA acceleration cards, and the external network.
[0048] Specifically, the CXL Switch has multiple CPU computing cards as hosts upstream, and multiple FPGA acceleration cards downstream. The memory of the FPGA acceleration card itself forms a memory pool, and the host can share the memory pool. The QSFP28 of the CPU computing card and the QSFP28 of the FPGA acceleration card are connected to the optical switch through optical modules and optical fibers, so that the CPU computing card and the FPGA acceleration card can communicate with each other in real time. The memory controller integrated in the CPU allocates the physical addresses of the multi-channel memory, realizes parallel access through interleaved mapping (one file corresponds to multi-channel storage), and improves the bandwidth. When the system starts, the CXL Switch enumerates all connected CXL memory devices through the CXL.io protocol, obtains information such as physical capacity and interface type (Type3 device), and then the structure manager of the CXL Switch allocates a continuous global address segment to each memory device according to the device capacity and topology. When the host CPU issues a memory access request using a local physical address, the CXL Switch converts the local physical address to a global address through a global address decoder, queries an internal routing table to determine the target memory device according to the address segment where the global address is located, converts the global address to a local device physical address, and finally the FPGA internal memory controller converts the device physical address to a specific chip row and column address (such as DDR5 Bank / Row / Column), completing the final data read and write. DDR (Double Data Rate) is a double-rate synchronous dynamic random access memory.
[0049] Please refer to Figure 3 The internal structure of any CPU computing card can include CPU, CPLD (Complex Programmable Logic Device), Flash, NVMe SSD, memory DDR5, clock, and VRM power module, etc. Among them, the CPU is mainly responsible for resource allocation and calculation; the CPLD is mainly responsible for controlling power timing and GPIO communication; the VRM is a different power module that outputs different voltage values for device use on the board; DDR5_CH1~DDR5_CH4 represents the RDIMM memory slot, used to insert DDR5 memory; QSFP28_0 and QSFP28_1 can communicate externally; USB / JTAG / UART is mainly used for debugging interface. The PCIe form of CPU computing card is essentially a motherboard, and the gold finger supports PCIe Gen6.0 and CXL3.0 protocols, and is inserted into the Slot slot of the Switch exchange board.
[0050] Please refer to Figure 4The internal structure of any one FPGA acceleration card can include an FPGA, a CPLD, a Flash, and a VRM power module, etc., wherein the FPGA is a master chip responsible for functions such as memory expansion, acceleration calculation, and communication; the CPLD is responsible for controlling power timing and GPIO communication; the VRM is a different power module outputting different voltage values for the board card device; the DDR5_CH1~DDR5_CH4 represents four-channel memory slot, used for inserting DDR5 memory; the QSFP28_0 and QSFP28_1 are optical connectors of the FPGA acceleration card, which can perform optical communication externally; the gold finger of the board card supports PCIe Gen6.0 and CXL3.0 protocols, and is inserted into the Slot slot of the Switch exchange board.
[0051] Referring to Figure 5 The internal structure of the CXL Switch can include a Switch supporting PCIe6.0 and CXL3.0, a CPLD, a VRM, and a PCIe GEN6.0 Slot. The upstream of the Switch is a CPU computing card, and the downstream is an FPGA acceleration card, etc., and the specific number of supported board cards is determined according to the Lane of the Switch; the VRM is a different power module outputting different voltage values for the board card device; the CPLD is mainly used for managing and configuring the Switch.
[0052] Referring to Figure 6 , Figure 2 The data transmission and processing flow supported by the system shown in FIG. 13 can include: when the system is started, the host is responsible for initializing the memory allocation request, and interacts with the Switch switch through the CXL protocol. The Switch bottom plate CPLD communicates with the Switch switch through the SMbus, initializes and enumerates the CXL memory device, and at the same time, the CXL Switch switch as the address allocation execution subject allocates the memory address space of the shared memory pool composed of the FPGA acceleration card. When the external network data (i.e. the task to be processed) enters the high-speed optical port of the CPU computing card through the optical switch, the optical signal is converted into an electrical signal through the optical module, and then the data is transmitted to the shared memory pool or the local memory of the CPU computing card through the IO bus (mainly depending on whether the local memory has space to temporarily store the data). Then the interrupt signal notifies the CPU that the data is ready, the CPU loads the data for preprocessing, and preliminarily judges whether the data processing task needs to be split into CPU+FPGA processing. If it does not need to be split, the CPU directly processes, and after completion, sends it out through the high-speed optical port of the CPU computing card; if it needs to be split, the CPU writes the result back to the memory of the FPGA acceleration card in the shared memory pool after processing, and the FPGA performs secondary processing on the data, and after processing, directly sends it out through the high-speed optical port of the FPGA acceleration card.
[0053] For whether the network data needs to be split, firstly judge whether the data needs flexible logical judgment, dynamic rule adaptation or complex state management in early processing, and needs fixed mode large-scale parallel computing, low delay operation in later processing, if so, it is more appropriate to use CPU processing in early stage and FPGA processing in later stage. For this, a task adaptability scoring model is needed, through model comparison, to determine whether the task needs to be split, and which of the three modes of CPU processing, FPGA processing and CPU+FPGA processing is more appropriate.
[0054] According to the characteristics of data processing tasks, four dimensions can be split to quantify the task adaptability scoring indicators, which can be specifically referred to Table 1.
[0055] Table 1 Task characteristic classification table
[0056]
[0057] Referring to the operation logic characteristics (i.e. dynamic logic demand), complexity (i.e. state management complexity), parallelism (i.e. parallel computing density) and delay sensitivity provided in Table 1, the following calculation formula can be designed.
[0058] 1. The CPU processing adaptability is calculated by the following formula (i.e. the first formula): ; wherein a1, a2, a3, a4 are weighting coefficients, which can be fixed constants according to different tasks; R represents the processing rule update frequency, the greater the value, the more frequent the logic change; S represents the state complexity, i.e. the product of the number of state variables and the number of state transitions, which can represent the complexity, the greater the value, the more complex the state management; P represents the number of parallel computing units, the smaller the value, the smaller the parallel computing scale; T represents the maximum tolerance delay, the greater the value, the more relaxed the delay requirement.
[0059] 2. The FPGA processing adaptability is calculated by the following formula (i.e. the second formula): ; wherein b1, b2, b3, b4 are weighting coefficients, which can be fixed constants according to different tasks; represents the processing rule stability, the greater the value, the more stable the logic change; represents the state simplicity, the greater the value, the simpler the state; P represents the number of parallel computing units, the greater the value, the larger the parallel computing scale; represents the delay sensitivity, the greater the value, the more stringent the delay requirement.
[0060] The calculation values output by the two formulas are formulated as follows: when , it means that the data to be processed needs to frequently adjust the logic rules, dynamically update the state complexity, the parallel computing scale is small, and the delay requirement is relaxed, which is more suitable for CPU data processing; when When the data to be processed needs logical rules that do not need to be frequently adjusted, the dynamic update state is simple, the parallel computing scale is large, and the delay requirement is strict, the data is more suitable for FPGA to perform data processing. When the data is more suitable for CPU+FPGA to jointly complete data processing, the CPU or the FPGA can also be randomly selected to complete the processing.
[0061] It should be noted that each CPU computing card enumerates all visible acceleration cards through a CXL switch, the CXLSwitch supports multi-host sharing, and resources are negotiated during enumeration. If the local memory of a CPU computing card is insufficient, it can use both the memory of the acceleration card and the memory of other CPU computing cards; the optical switch determines to which CPU computing node to forward the task according to a preset routing strategy or a dynamic control instruction; the CPU computing card can borrow other CPU computing cards to complete the task after splitting the task; the number of CPU computing cards or acceleration cards used by each task can be comprehensively considered from multiple dimensions such as resource utilization of each device, task execution state, performance index, etc.; for example, the resource utilization, queue length, execution time and other indexes of each device are monitored in real time, and the number of CPU computing cards or acceleration cards allocated to a task is comprehensively considered based on these indexes.
[0062] It can be seen that the system provided in the embodiment can not only expand the memory pool, but also expand the computing power, and the processed data can be directly forwarded through the optical port. The memory expansion system combined with the CPU and the FPGA can realize high flexibility, high throughput and low delay of data processing. That is, the FPGA acceleration card not only realizes host memory expansion, but also realizes host computing power expansion.
[0063] Next, a data processing method provided by the embodiment of the application is introduced. The data processing method described below can be mutually referred to with other embodiments described herein.
[0064] The embodiment of the application discloses a data processing method, which is applied to a data processing system, and the data processing system comprises a plurality of processors, a first switching device, a plurality of second devices and a second switching device; the plurality of processors and the plurality of second devices are connected with the first switching device; the plurality of processors and the plurality of second devices are connected with the second switching device; the first switching device is used for constructing a first communication link between any processor and any second device; according to resource demand sent by any processor, a target resource matched with the resource demand is selected from other devices; the other devices comprise at least one or a combination of the plurality of second devices and other processors except the processor sending the resource demand; the target resource comprises at least one or a combination of a memory resource and a computing resource; the second switching device is used for constructing a second communication link between any processor and any second device, communication between any processor and an external network and communication between any second device and the external network.
[0065] The data processing method provided by the embodiment comprises the following steps: any processor sends resource demand to the first switching device; the first switching device selects a target resource matched with the resource demand from other devices according to the resource demand; the other devices comprise at least one or a combination of the plurality of second devices and other processors except the processor sending the resource demand; the target resource comprises at least one or a combination of a memory resource and a computing resource.
[0066] The data processing method provided by the embodiment comprises the following steps: any processor or any second device realizes memory expansion through the first switching device when the memory is insufficient; any processor realizes computing power expansion through the first switching device when the computing power is insufficient.
[0067] The data processing method provided by the embodiment further comprises the following steps: any processor transmits processing data to a target second device through the first communication link provided by the first switching device; the target second device processes the processing data to obtain a processing result; the processing result is transmitted to a corresponding processor through the second communication link provided by the second switching device; or the processing result is transmitted to an external network through the second switching device.
[0068] In an implementation manner, the second switching device performs protocol analysis on a task received from the external network, queries a corresponding routing strategy according to a protocol analysis result, and transmits the task to a corresponding target processor according to the routing strategy.
[0069] In an implementation manner, the second switching device generates a corresponding task dispatching instruction according to the routing strategy, and transmits the task to the target processor by executing the task dispatching instruction.
[0070] In an embodiment, the destination processor extracts features of the task, determines execution requirements according to the extracted features of the task, and selects an execution end for the task from the plurality of processors and / or the plurality of second devices according to the execution requirements.
[0071] In an embodiment, the features of the task include operation logic features, complexity, parallelism, and delay sensitivity, and the destination processor determines first and second adaptation degrees corresponding to the task according to the operation logic features, the complexity, the parallelism, and the delay sensitivity, and selects an execution end for the task from the plurality of processors and / or the plurality of second devices according to the first and second adaptation degrees.
[0072] In an embodiment, the destination processor calculates the first adaptation degree by using a first formula: Q i =a1×R+a2×S+a3×(1 / P)+a4×T; Q i is the first adaptation degree corresponding to the task i, R is the operation logic features, a1 is a weighting coefficient of the operation logic features in the first formula, S is the complexity, a2 is a weighting coefficient of the complexity in the first formula, P is the parallelism, a3 is a weighting coefficient of the parallelism in the first formula, and T is the delay sensitivity, and a4 is a weighting coefficient of the delay sensitivity in the first formula.
[0073] In an embodiment, the destination processor calculates the second adaptation degree by using a second formula: W i =b1×(1 / R)+b2×(1 / S)+b3×P+b4×(1 / T); W i is the second adaptation degree corresponding to the task i, R is the operation logic features, b1 is a weighting coefficient of the operation logic features in the second formula, S is the complexity, b2 is a weighting coefficient of the complexity in the second formula, P is the parallelism, b3 is a weighting coefficient of the parallelism in the second formula, and T is the delay sensitivity, and b4 is a weighting coefficient of the delay sensitivity in the second formula.
[0074] In an embodiment, if the first adaptation degree is greater than the second adaptation degree and the difference between the first adaptation degree and the second adaptation degree is greater than a preset first threshold, the destination processor selects an execution end for the task from the plurality of processors.
[0075] In an embodiment, if the first adaptation degree is less than the second adaptation degree and the difference between the first adaptation degree and the second adaptation degree is greater than a preset second threshold, the destination processor selects an execution end for the task from the plurality of second devices.
[0076] In an embodiment, the target processor is the execution end of the task if the difference between the first and second adaptation degrees is less than a third preset threshold and the target processor meets the execution requirement of the task.
[0077] In an embodiment, the target processor splits the task into a first sub-task and a second sub-task if the target processor does not meet the execution requirement of the task; the target processor is the execution end of the first sub-task, and a processor other than the target processor is selected as the execution end of the second sub-task.
[0078] In an embodiment, the execution requirement of the task includes a memory requirement and a computing power requirement; accordingly, the target processor meets the execution requirement of the task if a remaining memory space of the target processor is not less than the memory requirement of the task and a remaining computing power of the target processor is not less than the computing power requirement of the task; otherwise, the target processor does not meet the execution requirement of the task.
[0079] In an embodiment, the target processor splits a first sub-task from the task according to the remaining memory space of the target processor and the remaining computing power of the target processor; the first sub-task requires a memory amount not greater than the remaining memory space of the target processor and a computing power not greater than the remaining computing power of the target processor; and a remaining task amount of the task is a second sub-task.
[0080] In an embodiment, the target processor transmits the data to be processed to the target second device through a first communication link provided by the first switching device.
[0081] In an embodiment, the target second device processes the data to be processed to obtain a processing result; transmits the processing result to the corresponding processor through a second communication link provided by the second switching device; or transmits the processing result to the external network through the second switching device.
[0082] In an embodiment, the corresponding processor determines response data based on the processing result and transmits the response data to the external network through the second switching device.
[0083] In an embodiment, the corresponding processor determines response data based on the processing result and transmits the response data to the external network through the second switching device.
[0084] It can be seen that the embodiment provides a data processing method, and no matter which device is insufficient in memory or computing power, memory allocation and computing power allocation can be realized through the first switching device; the first communication link is provided between any processor and any second device through the first switching device, and the second communication link is also provided through the second switching device, the first communication link and the second communication link can be used in parallel to improve transmission efficiency, and the redundancy and fault tolerance of the transmission link between any processor and any second device can be increased, thereby providing guarantee for communication reliability and communication efficiency; any processor and any second device can realize communication with the external network through the second switching device, thereby removing the limitation that only the host device can communicate with the external network, and improving communication flexibility.
[0085] An electronic device provided in the embodiment of the application is described below, and the electronic device described below can be referred to with other embodiments described herein. The electronic device provided in the embodiment can be the processor, the second device, the first switching device, the second switching device, and the like described in other embodiments.
[0086] The embodiment of the application discloses an electronic device, including: a memory for saving a computer program; a processor for executing the computer program to realize the method disclosed in any of the above embodiments.
[0087] In the embodiment, when the processor executes the computer program saved in the memory, the following steps can be specifically realized: when the memory is insufficient, memory expansion is realized through the first switching device; and when the computing power of any processor is insufficient, computing power expansion is realized through the first switching device.
[0088] In the embodiment, when the processor executes the computer program saved in the memory, the following steps can be specifically realized: the first communication link provided by the first switching device is used to transmit the data to be processed to the target second device.
[0089] In the embodiment, when the processor executes the computer program saved in the memory, the following steps can be specifically realized: the data to be processed is processed to obtain a processing result; the second communication link provided by the second switching device is used to transmit the processing result to the corresponding processor; or the second switching device is used to transmit the processing result to the external network.
[0090] In the embodiment, when the processor executes the computer program saved in the memory, the following steps can be specifically realized: the task received from the external network is protocol-analyzed, the corresponding routing strategy is queried according to the protocol analysis result, and the task is transmitted to the corresponding target processor according to the routing strategy.
[0091] In the embodiment, the processor executes the computer program stored in the memory, and the following steps can be implemented: generating a corresponding task dispatch instruction according to the routing strategy; and transmitting the task to the target processor by executing the task dispatch instruction.
[0092] In the embodiment, the processor executes the computer program stored in the memory, and the following steps can be implemented: extracting features of the task, determining an execution requirement according to the extracted task features; and selecting an execution end for the task from the plurality of processors and / or the plurality of second devices according to the execution requirement.
[0093] In the embodiment, the processor executes the computer program stored in the memory, and the following steps can be implemented: determining the first adaptation degree and the second adaptation degree corresponding to the task according to the operation logic features, the complexity, the parallelism and the delay sensitivity; and selecting an execution end for the task from the plurality of processors and / or the plurality of second devices according to the first adaptation degree and the second adaptation degree.
[0094] In the embodiment, the processor executes the computer program stored in the memory, and the following steps can be implemented: if the first adaptation degree is greater than the second adaptation degree, and the difference between the first adaptation degree and the second adaptation degree is greater than a preset first threshold, selecting an execution end for the task from the plurality of processors.
[0095] In the embodiment, the processor executes the computer program stored in the memory, and the following steps can be implemented: if the first adaptation degree is less than the second adaptation degree, and the difference between the first adaptation degree and the second adaptation degree is greater than a preset second threshold, selecting an execution end for the task from the plurality of second devices.
[0096] In the embodiment, the processor executes the computer program stored in the memory, and the following steps can be implemented: if the difference between the first adaptation degree and the second adaptation degree is less than a preset third threshold, selecting the target processor as the execution end of the task if the target processor meets the execution requirement of the task.
[0097] In the embodiment, the processor executes the computer program stored in the memory, and the following steps can be implemented: if the target processor does not meet the execution requirement of the task, splitting the task into a first subtask and a second subtask; selecting the target processor as the execution end of the first subtask, and selecting an execution end for the second subtask from the plurality of second devices or other processors except the target processor.
[0098] In the embodiment, the processor, when executing the computer program stored in the memory, can specifically implement the following steps: if the remaining memory space of the target processor is not less than the memory requirement of the task and the remaining computing power of the target processor is not less than the computing power requirement of the task, it is determined that the target processor meets the execution requirement of the task; otherwise, it is determined that the target processor does not meet the execution requirement of the task.
[0099] In the embodiment, the processor, when executing the computer program stored in the memory, can specifically implement the following steps: according to the remaining memory space of the target processor and the remaining computing power of the target processor, a first subtask is split from the task; the memory required by the first subtask is not greater than the remaining memory space of the target processor, and the computing power required by the first subtask is not greater than the remaining computing power of the target processor; and the remaining task amount of the task is taken as a second subtask.
[0100] In the embodiment, the processor, when executing the computer program stored in the memory, can specifically implement the following steps: the data to be processed is transmitted to the target second device through the first communication link provided by the first switching device.
[0101] In the embodiment, the processor, when executing the computer program stored in the memory, can specifically implement the following steps: the data to be processed is transmitted to the target second device through the first communication link provided by the first switching device.
[0102] In the embodiment, the processor, when executing the computer program stored in the memory, can specifically implement the following steps: the data to be processed is transmitted to the target second device through the first communication link provided by the first switching device.
[0103] Further, the embodiment of the present application also provides an electronic device. Wherein, the above-mentioned electronic device can be a server as shown in Figure 7 , or a terminal as shown in Figure 8 . Figure 7 , and Figure 8 are structural diagrams of electronic devices according to an exemplary embodiment, and the contents in the diagrams cannot be considered as any limitation to the use range of the present application.
[0104] Figure 7 A structural diagram of a server provided by the embodiment of the present application. The server can specifically include at least one processor, at least one memory, a power supply, a communication interface, an input / output interface and a communication bus. Wherein, the memory is used to store a computer program, the computer program is loaded and executed by the processor to realize the related steps in the data processing disclosed in any of the preceding embodiments.
[0105] In this embodiment, the power supply is used to provide working voltage for each hardware device on the server; the communication interface can create a data transmission channel between the server and external devices, and the communication protocol followed by the communication interface is any communication protocol applicable to the technical solution of the present application, which is not specifically limited here; the input and output interface is used to obtain external input data or output data to the outside world, and the specific interface type can be selected according to the specific application needs, which is not specifically limited here.
[0106] In addition, the memory as a carrier of resource storage can be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the resources stored thereon include an operating system, a computer program and data, etc., and the storage mode can be temporary storage or permanent storage.
[0107] The operating system is used to manage and control each hardware device and computer program on the server to realize the operation and processing of the processor on the data in the memory, which can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program capable of completing the data processing method disclosed in any of the preceding embodiments, the computer program can further include a computer program capable of completing other specific work. In addition to the data including the update information of the application program, the data can also include the developer information of the application program.
[0108] Figure 8 A structural schematic diagram of a terminal provided by the embodiment of the present application, which specifically can include but is not limited to a smart phone, a tablet computer, a notebook computer or a desktop computer, etc.
[0109] Generally, the terminal in this embodiment includes a processor and a memory.
[0110] The processor can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content required to be displayed on the display screen. In some embodiments, the processor can also include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.
[0111] The memory can include one or more computer non-volatile storage media, which can be non-transitory. The memory can also include a high-speed random access memory and a non-volatile memory such as one or more disk storage devices, flash memory devices. In this embodiment, the memory is at least used to store the following computer programs, wherein the computer programs are loaded and executed by the processor, and can realize the related steps in the data processing method executed by the terminal side disclosed in any of the preceding embodiments. In addition, the resources stored in the memory can also include an operating system and data, and the storage mode can be temporary storage or permanent storage. The operating system can include Windows, Unix, Linux, and the like. The data can include but is not limited to application update information.
[0112] In some embodiments, the terminal can also include a display screen, an input / output interface, a communication interface, a sensor, a power supply, and a communication bus.
[0113] Those skilled in the art can understand that the structure shown in the above embodiments is not a limitation on the terminal, and can include more or fewer components than the illustrated components. Figure 8 The structure shown in the above embodiments is not a limitation on the terminal, and can include more or fewer components than the illustrated components.
[0114] A non-volatile storage medium provided by an embodiment of the application is described below. The non-volatile storage medium described below can be mutually referred to with other embodiments described herein.
[0115] A non-volatile storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the data processing method disclosed in the foregoing embodiments. The non-volatile storage medium is a computer-readable non-volatile storage medium, and serves as a carrier for storing resources. The non-volatile storage medium can be a read-only memory, a random access memory, a magnetic disk, an optical disk, or the like. The resources stored on the non-volatile storage medium include an operating system, a computer program, data, and the like. The storage mode can be temporary storage or permanent storage.
[0116] A computer program product is described below. The computer program product described below can be combined with other embodiments described herein.
[0117] A computer program product includes computer programs / instructions, which, when executed by a processor, implement the steps of the data processing method disclosed above.
[0118] Another computer program product is provided in the embodiments of the present application. The computer program product includes a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium is used to store a computer program. The computer program, when executed by a processor, implements the steps of any of the above embodiments.
[0119] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts of each embodiment can be referred to each other.
[0120] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination thereof. The software module can be stored in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of non-volatile storage medium known in the art.
[0121] The principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method and core idea of the present application. For those skilled in the art, the specific implementation manners and application scope of the present application can be changed according to the idea of the present application. The above description of the embodiments should not be understood as a limitation of the present application.
Claims
1. A data processing system, characterized by The method comprises the following steps: a plurality of processors, a first switching device, a plurality of second devices and a second switching device; the plurality of processors and the plurality of second devices are connected to the first switching device; the plurality of processors and the plurality of second devices are connected to the second switching device; the first switching device is configured to: build a first communication link between any processor and any second device; and select a target resource matching a resource requirement sent by any processor from other devices according to the resource requirement; the other devices include at least one or a combination of the plurality of second devices and other processors except the processor sending the resource requirement; and the target resource includes a memory resource and a computing resource; the second switching device is configured to: build a second communication link between any processor and any second device, communication between any processor and an external network, and communication between any second device and the external network; wherein the second switching device is configured to: perform protocol analysis on a task received from the external network, query a corresponding routing strategy according to the protocol analysis result, and transmit the task to a corresponding destination processor according to the routing strategy; wherein the destination processor is configured to: perform feature extraction on the task; the task features include operation logic features, complexity, parallelism and delay sensitivity; determine a first adaptation degree and a second adaptation degree corresponding to the task according to the operation logic features, the complexity, the parallelism and the delay sensitivity; if the first adaptation degree is greater than the second adaptation degree, and the difference between the first adaptation degree and the second adaptation degree is greater than a preset first threshold, select an execution end for the task in the plurality of processors; if the first adaptation degree is less than the second adaptation degree, and the difference between the first adaptation degree and the second adaptation degree is greater than a preset second threshold, select an execution end for the task in the plurality of second devices; and if the difference between the first adaptation degree and the second adaptation degree is less than a preset third threshold, select the destination processor as the execution end of the task if the destination processor meets the execution requirement of the task; the execution requirement of the task includes a memory requirement and a computing power requirement; The processor is configured to calculate the first fitness degree by using a first formula, wherein the first formula is Q i =a1×R+a2×S+a3×(1 / P)+a4×T i is the first fitness degree corresponding to the task i, R is the operation logic feature, a1 is a weight coefficient of the operation logic feature in the first formula, S is the complexity, a2 is a weight coefficient of the complexity in the first formula, P is the parallelism, a3 is a weight coefficient of the parallelism in the first formula, T is the delay sensitivity, and a4 is a weight coefficient of the delay sensitivity in the first formula. Wherein, the processor is configured to calculate the second fitness by using a second formula, and the second formula is: W i = b1 x (1 / R) + b2 x (1 / S) + b3 x P + b4 x (1 / T); W i is a second fitness corresponding to task i, R is the operation logic feature, b1 is a weighting coefficient corresponding to the operation logic feature in the second formula; S is the complexity, b2 is a weighting coefficient corresponding to the complexity in the second formula; P is the parallelism, b3 is a weighting coefficient corresponding to the parallelism in the second formula; T is the delay sensitivity, and b4 is a weighting coefficient corresponding to the delay sensitivity in the second formula.
2. The data processing system of claim 1, wherein, the second switching device is configured to: generate a corresponding task dispatch instruction according to the routing strategy; and transmit the task to the destination processor by executing the task dispatch instruction.
3. The data processing system of claim 1, wherein, the destination processor is configured to: split the task into a first subtask and a second subtask if the destination processor does not meet the execution requirement of the task; select the destination processor as the execution end of the first subtask, and select an execution end for the second subtask in the plurality of second devices or other processors except the destination processor.
4. The data processing system of claim 1, wherein, the destination processor is configured to: confirm that the destination processor meets the execution requirement of the task if the remaining memory space of the destination processor is not less than the memory requirement of the task, and the remaining computing power of the destination processor is not less than the computing power requirement of the task; otherwise, confirm that the destination processor does not meet the execution requirement of the task.
5. The data processing system of claim 3, wherein, The target processor is configured to split the first subtask from the task according to a remaining memory space of the target processor and a remaining computing power of the target processor; a required memory amount of the first subtask is not greater than the remaining memory space of the target processor, and a required computing power of the first subtask is not greater than the remaining computing power of the target processor; and a remaining task amount of the task is taken as the second subtask.
6. The data processing system according to any one of claims 1 to 5, characterized in that, The arbitrary processor is configured to transmit the data to be processed to a target second device through a first communication link provided by the first switching device; Correspondingly, the target second device is configured to process the data to be processed to obtain a processing result, and transmit the processing result to the corresponding processor through a second communication link provided by the second switching device, or transmit the processing result to an external network through the second switching device.
7. The data processing system of claim 6, wherein, Correspondingly, the processor is configured to determine response data based on the processing result, and transmit the response data to the external network through the second switching device.
8. The data processing system according to any one of claims 1 to 5, characterized in that, The plurality of processors are a plurality of processor computing cards, the plurality of second devices are a plurality of accelerator cards, and the plurality of processor computing cards and the plurality of accelerator cards are plugged into the first switching device.
9. The data processing system of claim 8, wherein, The first switching device is an expansion card.
10. The data processing system according to any one of claims 1 to 5, characterized in that, The plurality of processors, the first switching device, and the plurality of second devices are arranged in a cabinet node, the second switching device is arranged outside the cabinet node, and the second switching device is further connected to other cabinet nodes.
11. A data processing method, characterized by, The application is applied to a data processing system, which comprises a plurality of processors, a first switching device, a plurality of second devices, and a second switching device; the plurality of processors and the plurality of second devices are connected to the first switching device; the plurality of processors and the plurality of second devices are connected to the second switching device; the first switching device is configured to construct a first communication link between an arbitrary processor and an arbitrary second device; the second switching device is configured to construct a second communication link between an arbitrary processor and an arbitrary second device, communication between an arbitrary processor and an external network, and communication between an arbitrary second device and the external network; Correspondingly, the data processing method comprises: An arbitrary processor sends resource requirements to the first switching device; The first switching device selects target resources matching the resource requirements from other devices according to the resource requirements; the other devices include at least one or a combination of the plurality of second devices and other processors except the processor sending the resource requirements; the target resources include memory resources and computing resources; The second switching device is configured to perform protocol analysis on a task received from an external network, query a corresponding routing strategy according to a protocol analysis result, and transmit the task to a corresponding target processor according to the routing strategy. The destination processor is configured to: perform feature extraction on the task; the task features include operation logic features, complexity, parallelism, and delay sensitivity; determine the first and second adaptation degrees corresponding to the task according to the operation logic features, the complexity, the parallelism, and the delay sensitivity; if the first adaptation degree is greater than the second adaptation degree, and the difference between the first adaptation degree and the second adaptation degree is greater than a preset first threshold, select an execution end for the task in the plurality of processors; if the first adaptation degree is less than the second adaptation degree, and the difference between the first adaptation degree and the second adaptation degree is greater than a preset second threshold, select an execution end for the task in the plurality of second devices; and if the difference between the first adaptation degree and the second adaptation degree is less than a preset third threshold, select the destination processor as the execution end of the task if the destination processor meets the execution requirement of the task; the execution requirement of the task includes memory requirement and computing power requirement. The processor is configured to calculate the first fitness degree by using a first formula, wherein the first formula is Q i =a1×R+a2×S+a3×(1 / P)+a4×T; Q i is the first fitness degree corresponding to the task i, R is the operation logic feature, a1 is a weight coefficient of the operation logic feature in the first formula; S is the complexity, a2 is a weight coefficient of the complexity in the first formula; P is the parallelism, a3 is a weight coefficient of the parallelism in the first formula; T is the delay sensitivity, and a4 is a weight coefficient of the delay sensitivity in the first formula. The processor is configured to calculate the second fitness by using a second formula, wherein the second formula is W i = b1 x (1 / R) + b2 x (1 / S) + b3 x P + b4 x (1 / T); W i is a second fitness corresponding to the task i, R is the operation logic feature, b1 is a weight coefficient corresponding to the operation logic feature in the second formula; S is the complexity, b2 is a weight coefficient corresponding to the complexity in the second formula; P is the parallelism, b3 is a weight coefficient corresponding to the parallelism in the second formula; T is the delay sensitivity, and b4 is a weight coefficient corresponding to the delay sensitivity in the second formula.
12. An electronic device, comprising: The computer program is configured to implement the method of claim 11 when executed by a processor. The computer program is configured to implement the method of claim 11 when executed by a processor. The computer program is configured to implement the method of claim 11 when executed by a processor.
13. A non-volatile storage medium, comprising: 14. A computer program product comprising computer programs / instructions, characterized in that,
Citation Information
Patent Citations
Switch, memory sharing method and system, computing equipment and storage medium
CN117118930A
Data processing system and method and medium
CN117806833A