Single-channel high-resolution output circuit and data processing method thereof
By designing a single-channel high-resolution output circuit, optimizing the layout of NUMA nodes and accelerator resources, an efficient image high-resolution reconstruction output under limited hardware resources is achieved, solving the problem of excessive resource occupation.
Patent Information
- Application Number
- CN202510223513.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Under the limited hardware resource limitations, how to reasonably allocate the resource occupation requirements of different tasks and realize high-resolution reconstruction output of images based on available single-channel resources.
A single-channel high-resolution output circuit is designed, including a processor slot, an acceleration processing unit, an input unit, a split unit, a single-channel determination unit, a mapping unit and an image reconstruction unit. Through these modules, the layout of NUMA nodes and accelerator resources is optimized, the most suitable target image processing channel is selected, and the acceleration processing unit is mapped to some NUMA nodes to achieve high-resolution reconstruction of the image.
It realizes efficient image high-resolution reconstruction output under limited hardware resources, optimizes resource utilization, and avoids the problem of excessive resource occupation in multi-channel processing mode.
Smart Images

Figure CN119722464B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and hardware layout optimization, and in particular relates to a single-channel high-resolution output circuit and a data processing method thereof, a computer-readable storage medium for implementing the method, a computer program product and an electronic device. Background Art
[0002] In the digital age, pictures / videos have become an indispensable part of our daily life and work. However, many times, the pictures / videos we have are not satisfactory, especially when we need to zoom in to see the details, the image is often blurred due to insufficient resolution, which is particularly obvious when watching on large screens and at close range.
[0003] To this end, high (ultra-high) resolution image / video output (reconstruction) technology has emerged. For example, 8K screen resolution is a very high display resolution, usually 7680×4320 pixels. This means that the screen has 7680 horizontal pixels and 4320 vertical pixels, which is often referred to as "8K Ultra High Definition" (8K UHD). Compared with the more common high-definition (HD) and ultra-high-definition (UHD) resolutions, 8K provides an extremely high level of detail, making the picture clearer and sharper, especially on large screens and when viewed at close range. 8K screens provide extremely high image clarity and are suitable for large TVs, monitors, and professional-grade display equipment (such as filmmaking, medical imaging, etc.).
[0004] The way to achieve 8K screen resolution is usually based on screen splicing technology, such as using multi-channel image processing resources to display multiple input images with lower resolution (such as 2K / 4K) on multiple sub-screens, and then using image interpolation, smoothing, enhancement and other technologies to perform high (ultra-high) resolution image reconstruction, thereby achieving 8K resolution output on high-size large screens. For example, the Chinese invention patent application with publication number CN119052533A proposes a multi-channel video super-resolution server processing method and processing system, which can realize efficient real-time multi-channel video processing and super-resolution processing to improve the user's video experience, reduce waiting time, and reduce the concurrent pressure of the server.
[0005] However, ultra-high-resolution reconstruction and output means long-term and high-proportion occupation of high-specification hardware resources. If the utilization of multiple hardware channel resources is not properly allocated, it will also affect other processing tasks that require high-specification hardware resources for processing. Therefore, it is particularly important to study the reasonable allocation of resource occupancy requirements of different tasks under the constraints of limited hardware resources (for example, only single-channel resources are currently available) and achieve high-resolution reconstruction and output of images based on available single-channel resources. Summary of the invention
[0006] In view of the above technical problems, the present invention proposes a single-channel high-resolution output circuit and a data processing method thereof, a computer-readable storage medium for implementing the method, a computer program product and an electronic device.
[0007] In a first aspect of the present invention, a single-channel high-resolution output circuit is provided, wherein the output circuit comprises at least one processor slot and a plurality of accelerated processing units; the output circuit further comprises:
[0008] An input unit, the input unit being used to input image data of a first resolution;
[0009] A splitting unit, wherein the splitting unit splits the multiple processor cores on the at least one processor socket into a plurality of different NUMA nodes, each NUMA node including one or more processor cores;
[0010] A single channel determination unit, the single channel determination unit is used to select a NUMA node from the multiple different NUMA nodes as a target image processing channel at a current moment;
[0011] A mapping unit, wherein the mapping unit maps the plurality of accelerated processing units to some NUMA nodes among the plurality of different NUMA nodes;
[0012] an image reconstruction unit, wherein the image reconstruction unit receives the first image data of the first resolution input by the input unit at the current moment and the second image data of the second resolution input at a moment before the current moment by using the target image processing channel at the current moment, reconstructs an image of a third resolution and outputs the image;
[0013] The third resolution is greater than the first resolution and greater than the second resolution.
[0014] The first resolution is 2K, and the third resolution is 8K.
[0015] In a simple-design output circuit, the output circuit includes only one processor socket, and the one processor socket includes a plurality of processor cores;
[0016] The single channel determination unit is used to select a NUMA node from the multiple different NUMA nodes as a target image processing channel at the current moment, specifically including:
[0017] For each NUMA node, obtain the total number N of processor cores contained in the NUMA node and the total number M of currently idle processor cores;
[0018] Will The largest NUMA node is used as the target image processing channel at the current moment.
[0019] In an output circuit of a complex design, the output circuit includes P processor sockets, each of the processor sockets includes Q processor cores; ;
[0020] The single channel determination unit is used to select a NUMA node from the multiple different NUMA nodes as a target image processing channel at the current moment, specifically including:
[0021] Calculate the number of NUMA nodes included in each processor socket;
[0022] The processor socket with the largest number of NUMA nodes is used as the target processor socket;
[0023] For each NUMA node of the target processor socket, obtain the total number N of processor cores contained in the NUMA node and the total number M of currently idle processor cores;
[0024] Will The largest NUMA node is used as the target image processing channel at the current moment.
[0025] The mapping unit maps the multiple accelerated processing units to some NUMA nodes among the multiple different NUMA nodes, specifically including:
[0026] Obtain the number Acc of accelerated processing units available at the current moment and the total number NAcc of accelerated processing units;
[0027] For each NUMA node, when it meets the following conditions, it is regarded as the partial NUMA node: Where N is the total number of processor cores contained in the NUMA node; M is the total number of processor cores currently idle in the NUMA node;
[0028] Get the node number Numpart of the part of NUMA nodes;
[0029] Based on the node number Numpart of the partial NUMA nodes, multiple acceleration processing units are divided into at least one acceleration resource queue; and by associating the at least one acceleration resource queue to the corresponding one of the partial NUMA nodes, multiple processor cores divided into one of the partial NUMA nodes use the at least one acceleration resource queue based on the association.
[0030] The multiple acceleration processing units include multiple FPGA acceleration resources;
[0031] The image reconstruction unit receives the first image data of the first resolution input by the input unit at the current moment and the second image data of the second resolution input at a moment before the current moment by using the target image processing channel at the current moment to reconstruct and output an image of the third resolution, specifically including:
[0032] The image reconstruction unit utilizes the multiple processor cores of the NUMA node corresponding to the target image processing channel at the current moment and uses the associated FPGA acceleration resources to perform the reconstruction.
[0033] In response to a change in the node number Numpart of the partial NUMA nodes, the association between the at least one acceleration resource queue and one of the corresponding partial NUMA nodes is adjusted.
[0034] In a second aspect of the present invention, a data processing method is further provided. The data processing method is implemented based on the output circuit described in the first aspect, and the method comprises the following steps:
[0035] S810: Selecting a NUMA node from a plurality of different NUMA nodes as a target image processing channel at a current moment;
[0036] S820: Mapping the plurality of accelerated processing units to some NUMA nodes among the plurality of different NUMA nodes;
[0037] S830: Reconstruct an image with a third resolution by using the target image processing channel at the current moment to receive the first image data with a first resolution input at the current moment and the second image data with a second resolution input at a moment before the current moment and output the image;
[0038] Wherein, the third resolution is greater than the first resolution, and greater than the second resolution;
[0039] Wherein, the multiple acceleration processing units include multiple FPGA acceleration resources;
[0040] The step S830 specifically includes:
[0041] The reconstruction is performed using the associated FPGA acceleration resources using multiple processor cores of the NUMA node corresponding to the target image processing channel at the current moment.
[0042] The step S820 specifically includes:
[0043] Obtain the number Acc of accelerated processing units available at the current moment and the total number NAcc of accelerated processing units;
[0044] For each NUMA node, when it meets the following conditions, it is regarded as the partial NUMA node: Where N is the total number of processor cores contained in the NUMA node; M is the total number of processor cores currently idle in the NUMA node;
[0045] Get the node number Numpart of the part of NUMA nodes;
[0046] Based on the node number Numpart of the partial NUMA nodes, multiple acceleration processing units are divided into at least one acceleration resource queue; and by associating the at least one acceleration resource queue to the corresponding one of the partial NUMA nodes, multiple processor cores divided into one of the partial NUMA nodes use the at least one acceleration resource queue based on the association.
[0047] The aforementioned data processing method can also be automatically implemented through various forms of electronic devices through computer program instructions; the computer program instructions can be stored in different forms of storage media and loaded into computer electronic devices for execution.
[0048] Therefore, in the third aspect of the present invention, a computer-readable storage medium is also provided for storing computer instructions, which, when executed on an electronic device, enables the electronic device to execute the aforementioned data processing method.
[0049] In a third aspect of the present invention, a computer device is further proposed, comprising a processor and a memory, wherein the memory is used to store instructions, and the processor is used to call the instructions in the memory so that the computer device executes the aforementioned data processing method.
[0050] In a fourth aspect of the present invention, a computer program product is also proposed, wherein the product comprises a computer program, and when the computer program is executed, the aforementioned data processing method is implemented.
[0051] The technical solution of the present invention can utilize the optimized layout matching of NUMA idle node resources and accelerator resources, thereby realizing high-resolution reconstruction output of images based on available single-channel resources.
[0052] Further advantages of the present invention will be further reflected in detail in the specific embodiments section in conjunction with the drawings of the specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0054] Figure 1 This is a schematic diagram of the functional modules of a single-channel high-resolution output circuit according to an embodiment of the present invention;
[0055] Figure 2 yes Figure 1 Schematic diagram of idle NUMA nodes in the embodiment;
[0056] Figure 3 yes Figure 1 Schematic diagram of the association between the embodiment NUMA node and accelerator resources;
[0057] Figure 4 is based on Figure 1 A schematic diagram of the main flow of the data processing method implemented in the embodiment. DETAILED DESCRIPTION
[0058] Before introducing various specific embodiments of the present invention, the meanings of relevant technical features that may be involved in the present invention are first introduced.
[0059] High-resolution image / video (in the context of the present invention, the image / video reconstruction principles are basically the same, so the subsequent introduction regards "image" and "video" as one concept, and uniformly adopts the concept of "image") reconstruction: also known as high-definition image magnification, high-definition reconstruction, resolution enhancement, etc., its core principle is to use advanced algorithms and computer technology to interpolate and reconstruct images. The interpolation algorithm can infer the value of unknown pixels based on the known pixel information in the image, thereby generating more pixels. The reconstruction algorithm can reconstruct a higher-resolution image based on these pixels. This process involves software implementations such as deep learning and neural network models, as well as hardware processing units such as central processing units (CPUs), image processing units (GPUs), general-purpose image processing units (GPGPUs), and accelerated processing units (ALUs).
[0060] Accelerated processor: referred to as accelerator, acceleration resource, accelerated processing resource, etc., is a concept relative to the ordinary central processing unit (CPU). The central processing unit (CPU) has general general capabilities, but there are obvious limitations in the capabilities of the CPU for specific task processing (such as high-resolution image reconstruction). At this time, the accelerator needs to be configured to perform certain tasks faster than the tasks originally executed by software on such accelerators and / or cores on the central processing unit for unloading CPU workload.
[0061] Common types of accelerated processors include field programmable gate arrays (FPGAs), graphics processing units (GPUs), general purpose GPUs (GP-GPUs), application specific integrated circuits (ASICs), and similar devices.
[0062] The acceleration processor used in the embodiments of the present invention is mainly an FPGA, which is used to assist the central processor in completing the corresponding image super-resolution output reconstruction process.
[0063] Processor sockets, processor cores and NUMA architecture: Currently, processors generally use non-uniform memory access (NUMA) architecture. In this architecture, different memory devices and processor cores belong to different NUMA nodes. Each NUMA node contains one or more processor cores, and each NUMA node has its own integrated memory controller (IMC).
[0064] With the improvement of semiconductor process technology and the widespread application of chiplet technology, more and more NUMA nodes need to be split on a single processor socket. That is, multiple processor cores on a single processor socket are split into multiple different NUMA nodes.
[0065] As an exemplary introduction, each NUMA node includes multiple CPU cores, a level 3 (L3) cache, a PCIe port, and a universal information interface link (not shown), etc. Each NUMA node also includes two unified memory controllers, where each UMC is used as a memory channel. Therefore, in each NUMA node, all CPUs can access the memory through a two-channel interleaved mode. Since the memory access within each NUMA node is isolated from other nodes, the communication between two NUMA nodes is limited by the memory bandwidth that the remote NUMA node can provide.
[0066] NUMA node and accelerator resource access restrictions and delays: Since multiple NUMA nodes in the NUMA architecture can be connected and exchange information, each CPU can access the memory of the entire system, and the speed of accessing local memory will be much higher than the speed of accessing remote memory (memory of other nodes in the system). However, when the accelerator accesses memory across NUMA nodes, there is a problem of relatively large access delay due to the interconnection restrictions between processor sockets or chips (dies). In order to avoid this problem, the common practice in current computing devices is to configure the corresponding NUMA node for the accelerator and restrict the processor core that uses the accelerator and the accelerator to be in the same NUMA node, so that the accelerator can only be used by the processor core in the NUMA node to which it belongs, and cannot be used by the processor core in other NUMA nodes.
[0067] Image processing channel: Based on the above division, a NUMA node is usually regarded as an image processing channel, which contains at least one CPU (processor) core and at least one accelerator resource (FPGA) associated with the core.
[0068] Multiple processing channels: Use multiple NUMA nodes to perform image processing (resolution enhancement) in parallel, occupying more channels and hardware resources, affecting the execution of other parallel tasks.
[0069] Single processing channel: Each time a NUMA node is selected to receive the image data to be processed, and it is fused with the previously processed image data to perform resolution reconstruction (high-resolution output). In single-channel processing mode, each time the idle status of the NUMA node, the idle status of the accelerator resources, and the matching correlation between the two need to be taken into account to ensure that resources are fully utilized while taking into account multiple tasks.
[0070] It can be seen that the multi-processor channel mode is relatively simple to implement, but it requires more hardware resources and is relatively expensive. The single-processor model does not have special requirements for the hardware itself (at least one available NUMA node is sufficient). The focus is on optimizing resource scheduling at the software level based on resource detection and hardware layout optimization.
[0071] In the scenario of actual hardware cost (budget) constraints, the present invention focuses on studying the technical solution of reasonably allocating resource requirements of different tasks under limited hardware resource constraints (for example, only single-channel resources are currently available) and realizing high-resolution reconstruction output of images based on available single-channel resources.
[0072] See first Figure 1 , Figure 1 The present invention is a schematic diagram of the functional modules of a single-channel high-resolution output circuit according to an embodiment of the present invention.
[0073] exist Figure 1 , it is shown that the output circuit includes at least one processor socket and a plurality of accelerated processing units;
[0074] The output circuit further includes:
[0075] An input unit, the input unit being used to input image data of a first resolution;
[0076] A splitting unit, wherein the splitting unit splits the multiple processor cores on the at least one processor socket into a plurality of different NUMA nodes, each NUMA node including one or more processor cores;
[0077] A single channel determination unit, the single channel determination unit is used to select a NUMA node from the multiple different NUMA nodes as a target image processing channel at a current moment;
[0078] A mapping unit, wherein the mapping unit maps the plurality of accelerated processing units to some NUMA nodes among the plurality of different NUMA nodes;
[0079] an image reconstruction unit, wherein the image reconstruction unit receives the first image data of the first resolution input by the input unit at the current moment and the second image data of the second resolution input at a moment before the current moment by using the target image processing channel at the current moment, reconstructs an image of a third resolution and outputs the image;
[0080] The third resolution is greater than the first resolution and greater than the second resolution.
[0081] In a specific embodiment, the first resolution is 2K, and the third resolution is 8K.
[0082] Preferably, the second resolution is 2K or 4K.
[0083] It can be seen that in the embodiment of the present invention, the first resolution and the second resolution are lower resolutions (relative to the third resolution), and the third resolution is a higher resolution.
[0084] The image data of the first resolution and the second resolution are successively inputted through the input unit;
[0085] For example, at the first moment, the input is the image data of the first resolution, and at the second moment, the input is the image data of the second resolution; or at the first moment, the input is the image data of the second resolution, and at the second moment, the input is the image data of the first resolution. The first moment and the second moment are two consecutive input moments.
[0086] For the convenience of description, in the subsequent embodiments, the first moment is referred to as the current moment, and the second moment is the moment before the current moment.
[0087] Since the first resolution is 2K, 2K resolution itself is a relatively high resolution in the art (2K is called a low resolution only compared to 8K). Therefore, the image data of the first resolution input by the input unit itself has a large amount of data to be processed, and a single processor core is generally not enough to process it. Therefore, this embodiment first needs to determine the processing channel of the input image data of the first resolution (and the second resolution).
[0088] A processing channel includes at least one graphics processing thread. In the image processing (super-resolution reconstruction) scenario of the present application, a processing channel generally includes a graphics processing main thread (executed by a CPU core) and at least one acceleration processing co-thread (executed by a corresponding acceleration processor).
[0089] As mentioned above, in current hardware architecture, processors usually adopt a non-uniform memory access (NUMA) architecture.
[0090] In one topology under the NUMA architecture, there are multiple CPU sockets on the motherboard, corresponding to multiple physical CPUs, and one CPU socket corresponds to one logical NUMA node; each physical CPU has multiple logical CPU cores. In another topology, multiple processor cores on a single processor socket are split into multiple different NUMA nodes.
[0091] It should be noted that the "split" here should be understood as "logical split", such as using virtualization technology to divide each physical CPU into multiple logical CPU cores, etc. Therefore, the "processor core" mentioned in the embodiments of the present invention refers to the "logical CPU core".
[0092] Therefore, in the case where one CPU socket includes one physical CPU, one physical CPU can also be split into multiple logical CPU cores (processor cores), and then the multiple processor core logical CPU cores (processor cores) are split into multiple different NUMA nodes.
[0093] The case where multiple CPU sockets include multiple physical CPUs is handled similarly.
[0094] To this end, the splitting unit included in the output circuit splits the multiple processor cores on the at least one processor socket into multiple different NUMA nodes, each NUMA node including one or more processor cores.
[0095] When hardware resources are sufficient, multiple processing channels can be called simultaneously to receive, process and perform subsequent reconstruction and fusion steps on input image data at different times. However, using multiple NUMA nodes for parallel image processing (resolution enhancement) will occupy more channels and hardware resources, affecting the execution of other parallel tasks (especially general processing tasks and small tasks), and may even lead to overall task blocking, affecting the output of subsequent reconstructed image effects. For example, the premise of outputting super-resolution images is that other parallel tasks must be completed.
[0096] Therefore, in the scenario of actual hardware cost (budget) constraints, the present invention focuses on studying the technical solution of reasonably allocating resource requirements of different tasks under limited hardware resource constraints (for example, only single-channel resources are currently available) and realizing high-resolution reconstruction output of images based on available single-channel resources.
[0097] In a simple output circuit architecture, the output circuit includes only one processor socket, and the one processor socket includes Q processor cores;
[0098] In a complex output circuit architecture, the output circuit includes P processor sockets, each of the processor sockets includes Q processor cores; ;
[0099] The single channel determination unit is used to select a NUMA node from the multiple different NUMA nodes as a target image processing channel at a current moment.
[0100] In the above simple output circuit architecture, the single channel determination unit is used to select a NUMA node from the multiple different NUMA nodes as the target image processing channel at the current moment, specifically including:
[0101] For each NUMA node, obtain the total number N of processor cores contained in the NUMA node and the total number M of currently idle processor cores;
[0102] Will The largest NUMA node is used as the target image processing channel at the current moment.
[0103] Figure 2 Show Figure 1 Schematic diagram of idle NUMA nodes in the embodiment. Figure 2 In the processor socket shown, one or more existing physical CPUs have been "split" into 14 logical CPU cores and divided into four groups. Each group includes 4 CPU cores, forming a NUMA node. Figure 2 Shown as NUMA0, NUMA1, NUMA2, NUMA3.
[0104] At the current moment, all four processor cores of NUMA0 are busy (running 4 threads); only one processor core of NUMA1 is busy (the remaining three are idle); all four processors of NUMA2 are idle; and 2 processors of NUMA3 are idle (the remaining 2 are busy).
[0105] Therefore, in the case of a single processor socket, the single channel determination unit selects the NUMA2 node as the target image processing channel at the current moment.
[0106] Correspondingly, in the above-mentioned complex output circuit architecture, the single channel determination unit is used to select a NUMA node from the multiple different NUMA nodes as the target image processing channel at the current moment, specifically including:
[0107] Calculate the number of NUMA nodes included in each processor socket;
[0108] The processor socket with the largest number of NUMA nodes is used as the target processor socket;
[0109] For each NUMA node of the target processor socket, obtain the total number N of processor cores contained in the NUMA node and the total number M of currently idle processor cores;
[0110] Will The largest NUMA node is used as the target image processing channel at the current moment.
[0111] In this case, it is necessary to first determine the target processor socket, and then determine the target NUMA node on the target processor socket by referring to the same method.
[0112] Determining the target NUMA node is to determine the target image processing channel at the current moment. Since the target NUMA node contains at least one idle processor core, the idle processor core can be used as the main thread of subsequent super-resolution reconstruction to perform general tasks such as receiving image input data, image data preprocessing, image segmentation, smoothing, noise reduction, task segmentation (in order to enter the accelerator unit), etc.
[0113] However, it is not enough for the target image processing channel (NUMA node) to have only general main thread resources (CPU cores), and corresponding graphics processing resource accelerators are also required.
[0114] Preferably, the accelerator resource is FPGA.
[0115] However, when the accelerator accesses memory across NUMA nodes, there is a problem of relatively large access latency due to the interconnection limitations between processor sockets or chips (dies). To avoid this problem, the common practice in current computing devices is to configure the corresponding NUMA node for the accelerator and restrict the processor core that uses the accelerator to be in the same NUMA node as the accelerator, so that the accelerator can only be used by the processor core in the NUMA node to which it belongs, and cannot be used by the processor core in other NUMA nodes.
[0116] exist Figure 2 In the diagram, the multiple processor cores of the processor socket are divided into four NUMA nodes, NUMA0, NUMA1, NUMA2, and NUMA3, and the corresponding NUMA node configured for the accelerator is NUMA1. Then, only the processor cores in the NUMA1 node can use the accelerator, and the processor cores in other NUMA nodes cannot use the accelerator.
[0117] In order to solve the problem of limited accelerator usage, there is currently a solution that uses accelerator software to adjust the NUMA configuration and then allocate the use of accelerators so that processor cores in multiple NUMA nodes can use accelerators. This solution requires the configuration of complex software logic. In addition, the software needs to dynamically check the current NUMA configuration of the computing device and then dynamically allocate the use of accelerators based on the NUMA configuration. This further increases the complexity of software configuration.
[0118] To solve the above problem, the output circuit is configured with a mapping unit, and the mapping unit maps the multiple accelerated processing units to some NUMA nodes among the multiple different NUMA nodes.
[0119] Specifically, the mapping unit obtains the number Acc of accelerated processing units available at the current moment and the total number NAcc of accelerated processing units;
[0120] For each NUMA node, when it meets the following conditions, it is regarded as the partial NUMA node: Where N is the total number of processor cores contained in the NUMA node; M is the total number of processor cores currently idle in the NUMA node;
[0121] Get the node number Numpart of the part of NUMA nodes;
[0122] Based on the node number Numpart of the partial NUMA nodes, multiple acceleration processing units are divided into at least one acceleration resource queue; and by associating the at least one acceleration resource queue to the corresponding one of the partial NUMA nodes, multiple processor cores divided into one of the partial NUMA nodes use the at least one acceleration resource queue based on the association.
[0123] Figure 3 yes Figure 1 A schematic diagram of the association between NUMA nodes and accelerator resources in the embodiment.
[0124] Figure 3 is Figure 2 Continue the description based on Figure 2 In the figure, four NUMA nodes (NUMA0-NUMA3) are shown, corresponding to The values are 0, 0.75, 1.0, and 0.5;
[0125] exist Figure 3 In the figure, the number Acc of available accelerated processing units at the current moment is 3 (A2-A4 in the figure are idle, and A1 is busy), and the total number NAcc of accelerated processing units is 4.
[0126] Therefore, only NUMA1 nodes and NUMA2 nodes meet the above conditions ( ), therefore, the NUMA1 node and the NUMA2 node, as the candidate partial nodes, can be included in the accelerator allocation process at the current moment.
[0127] At this time, the currently available acceleration processing units are divided into two groups of acceleration resource queues, the first group of acceleration resource queues is A2 and A3, and the second group of acceleration resource queues is A4;
[0128] The first group of acceleration resource queues is associated with the NUMA1 node, and the second group of acceleration resource queues is associated with the NUMA2 node;
[0129] At this time, multiple (idle) processor cores of the NUMA1 node can use the first group of acceleration resource queues A2 and A3; and multiple (idle) processor cores of the NUMA2 node can use the second group of acceleration resource queues A4.
[0130] It can be seen that the NUMA3 node does not participate in the allocation of accelerator resources at this time because the idleness of the processor core does not match the idleness of the accelerator resources. Because NUMA3 is relatively busy, it is unlikely to be selected as the target image processing channel (target NUMA node) when the single channel determination unit determines the target image processing channel (target NUMA node), so there is no need to consider allocating associated accelerator resources to it.
[0131] Of course, the above accelerator mapping association process is dynamic, and in response to a change in the node number Numpart of the partial NUMA nodes, the association between the at least one acceleration resource queue and one of the corresponding partial NUMA nodes is adjusted.
[0132] That is to say, if the node number Numpart of the candidate nodes changes at the next moment, the above association process needs to be executed again.
[0133] Preferably, the multiple acceleration processing units include multiple FPGA acceleration resources;
[0134] The image reconstruction unit receives the first image data of the first resolution input by the input unit at the current moment and the second image data of the second resolution input at a moment before the current moment by using the target image processing channel at the current moment to reconstruct and output an image of the third resolution, specifically including:
[0135] The image reconstruction unit utilizes the multiple processor cores of the NUMA node corresponding to the target image processing channel at the current moment and uses the associated FPGA acceleration resources to perform the reconstruction.
[0136] In a specific example, each acceleration resource queue includes multiple FPGAs. At this time, each target image processing channel includes at least one main thread and multiple co-threads, the main thread is executed by the central processing unit, and the co-threads are executed by the FPGA, so as to complete the above-mentioned image super-resolution reconstruction output.
[0137] It can be understood that the use of a host computer (including a CPU processor core) and FPGA accelerator resources to achieve image super-resolution reconstruction output itself belongs to the prior art, and the embodiments of the present invention will not elaborate on this.
[0138] Based on the above output circuit, Figure 4 Shown based on Figure 1 A schematic diagram of the main flow of the data processing method implemented in the embodiment.
[0139] The data processing method is implemented based on the output circuit described in the first aspect, and the method comprises the following steps:
[0140] S810: Selecting a NUMA node from a plurality of different NUMA nodes as a target image processing channel at a current moment;
[0141] S820: Mapping the plurality of accelerated processing units to some NUMA nodes among the plurality of different NUMA nodes;
[0142] S830: Reconstruct an image with a third resolution by using the target image processing channel at the current moment to receive the first image data with a first resolution input at the current moment and the second image data with a second resolution input at a moment before the current moment and output the image;
[0143] Wherein, the third resolution is greater than the first resolution, and greater than the second resolution;
[0144] Wherein, the multiple acceleration processing units include multiple FPGA acceleration resources;
[0145] The step S830 specifically includes:
[0146] The reconstruction is performed using the associated FPGA acceleration resources using multiple processor cores of the NUMA node corresponding to the target image processing channel at the current moment.
[0147] The step S820 specifically includes:
[0148] Obtain the number Acc of accelerated processing units available at the current moment and the total number NAcc of accelerated processing units;
[0149] For each NUMA node, when it meets the following conditions, it is regarded as the partial NUMA node: Where N is the total number of processor cores contained in the NUMA node; M is the total number of processor cores currently idle in the NUMA node;
[0150] Get the node number Numpart of the part of NUMA nodes;
[0151] Based on the node number Numpart of the partial NUMA nodes, multiple acceleration processing units are divided into at least one acceleration resource queue; and by associating the at least one acceleration resource queue to the corresponding one of the partial NUMA nodes, multiple processor cores divided into one of the partial NUMA nodes use the at least one acceleration resource queue based on the association.
[0152] For other technologies, principles, algorithms or models not elaborated in detail in this application, please refer to the prior art.
[0153] In the above-mentioned embodiment section, the present invention provides multiple embodiments, each of which can constitute an independent technical solution and may contribute to the prior art and solve corresponding technical problems. However, it should be pointed out that different embodiments can be combined with each other without violating logic; at the same time, each embodiment can solve at least one technical problem, but it is not required that each individual embodiment solves multiple or all technical problems.
[0154] It can be seen that since the super-resolution reconstruction output circuit and data processing method of the present application follow the existing NUMA structure and are executed based on single-channel resources, there is no need for hardware resource support of specific specifications, that is, the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0155] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0156] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.
Claims
1. A single-channel high-resolution output circuit, comprising at least one processor slot and a plurality of accelerated processing units; characterized in that: The output circuit also includes: An input unit, used for inputting image data of a first resolution; A splitting unit, splitting a plurality of processor cores on at least one processor socket into a plurality of different NUMA nodes, each NUMA node including one or more processor cores; A single channel determination unit, used to select a NUMA node from a plurality of different NUMA nodes as a target image processing channel at a current moment; A mapping unit maps the multiple acceleration processing units to some NUMA nodes among the multiple different NUMA nodes, specifically including: Obtain the number Acc of acceleration processing units available at the current moment and the total number NAcc of acceleration processing units; For each NUMA node, if it meets the following conditions, it is considered a partial NUMA node: N is the total number of processor cores contained in the NUMA node; M is the total number of processor cores currently idle in the NUMA node; Get the node number Numpart of the part of NUMA nodes; Based on the node number Numpart of the partial NUMA nodes, the plurality of accelerated processing units are divided into at least one accelerated resource queue; and by associating at least one accelerated resource queue with a corresponding one of the partial NUMA nodes, the plurality of processor cores divided into one of the partial NUMA nodes use the at least one accelerated resource queue based on the association; An image reconstruction unit receives, by using a target image processing channel at a current moment, first image data of a first resolution input by an input unit at a current moment and second image data of a second resolution input at a moment before the current moment, reconstructs an image of a third resolution and outputs the image; The third resolution is greater than the first resolution and greater than the second resolution.
2. A single-channel high-resolution output circuit as claimed in claim 1, characterized in that: The first resolution is 2K, and the third resolution is 8K.
3. A single-channel high-resolution output circuit as claimed in claim 1, characterized in that: The output circuit includes only one processor socket, and the one processor socket includes a plurality of processor cores; The single channel determination unit is used to select a NUMA node from the multiple different NUMA nodes as a target image processing channel at the current moment, specifically including: For each NUMA node, obtain the total number N of processor cores contained in the NUMA node and the total number M of currently idle processor cores; Will The largest NUMA node is used as the target image processing channel at the current moment.
4. A single-channel high-resolution output circuit as claimed in claim 1, characterized in that: The output circuit includes P processor slots, each of which includes Q processor cores; ; The single channel determination unit is used to select a NUMA node from the multiple different NUMA nodes as a target image processing channel at the current moment, specifically including: Calculate the number of NUMA nodes included in each processor socket; The processor socket with the largest number of NUMA nodes is used as the target processor socket; For each NUMA node of the target processor socket, obtain the total number N of processor cores contained in the NUMA node and the total number M of currently idle processor cores; Will The largest NUMA node is used as the target image processing channel at the current moment.
5. A single-channel high-resolution output circuit as claimed in claim 1, characterized in that: The multiple acceleration processing units include multiple FPGA acceleration resources; The image reconstruction unit receives the first image data of the first resolution input by the input unit at the current moment and the second image data of the second resolution input at a moment before the current moment by using the target image processing channel at the current moment to reconstruct and output an image of the third resolution, specifically including: The image reconstruction unit utilizes the multiple processor cores of the NUMA node corresponding to the target image processing channel at the current moment and uses the associated FPGA acceleration resources to perform the reconstruction.
6. A single-channel high-resolution output circuit as claimed in claim 1, characterized in that: In response to a change in the node number Numpart of the partial NUMA nodes, the association between the at least one acceleration resource queue and one of the corresponding partial NUMA nodes is adjusted.
7. A data processing method based on the single-channel high-resolution output circuit according to any one of claims 1 to 6, characterized in that: The method comprises the following steps: S810: Selecting a NUMA node from a plurality of different NUMA nodes as a target image processing channel at a current moment; S820: Mapping the plurality of accelerated processing units to some NUMA nodes among the plurality of different NUMA nodes; S830: Reconstruct an image with a third resolution by using the target image processing channel at the current moment to receive the first image data with a first resolution input at the current moment and the second image data with a second resolution input at a moment before the current moment and output the image; Wherein, the third resolution is greater than the first resolution, and greater than the second resolution; Wherein, the multiple acceleration processing units include multiple FPGA acceleration resources; The step S830 specifically includes: The reconstruction is performed using the associated FPGA acceleration resources using multiple processor cores of the NUMA node corresponding to the target image processing channel at the current moment.
8. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the data processing method according to claim 7.
Citation Information
Patent Citations
Multi-channel video super-resolution server processing method and processing system
CN119052533A
Processor binding techniques
CN117632468A