GPU (Graphics Processing Unit) and NPU (Network Processing Unit) hybrid operation method applied to security platform, medium and system
By adopting a hybrid computing method between GPU and NPU on the security platform, the problems of insufficient utilization of single-core CPU resources and high system upgrade costs in traditional designs are solved, and efficient secure computing and highly adaptable data processing capabilities are achieved.
Patent Information
- Application Number
- CN202411977109.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional security platform design relies on single-core CPU for secure computing, communication and control, resulting in the underutilization of heterogeneous core resources, high system upgrade costs, and difficult to meet the needs of big data processing and edge sensing.
The hybrid operation method of GPU and NPU is adopted, and the input data is hosted to the GPU and NPU for processing according to the characteristics of the input data. The tasks are allocated using OPENCL technology, the results are mapped to the CPU memory through DMA, and the data consistency is ensured through the interrupt mechanism.
It improves the computing efficiency of secure computing, makes full use of multi-core architecture and heterogeneous core performance, reduces the change cost of system upgrades, and adapts to the needs of big data processing and edge sensing.
Smart Images

Figure CN120067002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of secure computing, and in particular, to a GPU and NPU hybrid computing method, medium and system applied to a secure platform. Background Art
[0002] Currently, in general secure platforms such as industrial control, automotive electronics, and rail transit signals, the traditional design is to use a single-core CPU to participate in secure computing, comparison, and output, where the secure computing function is divided into different tasks for processing according to the data path.
[0003] Generally, according to the task division of the secure platform, the entire task is cyclically run basically according to the following process: data (communication) input, data processing, data comparison, data (communication) output. Among them, data (communication) input and data (communication) output are strongly related to hardware and must rely on hardware interfaces (such as network chips, serial port chips, digital IO interfaces) and related interrupt notification mechanisms. And data processing and data comparison are the parts independent of hardware, but they consume CPU resources and occupy task processing time. When multiple tasks run in parallel and switch between tasks, it will inevitably affect the real-time performance of the system.
[0004] With the rapid development of chip integration technology and software technology, this traditional implementation method gradually reveals its backwardness. The traditional design has the following disadvantages:
[0005] 1. Only the single-core CPU can be used for functional safety computing, communication, and control, and other heterogeneous cores are idle, and the performance of heterogeneous multi-cores is not fully reflected;
[0006] 2. When the traditional system is upgraded, the system software and hardware need to be redesigned, developed, and debugged, making the change cost of developing a new system much greater than that of this platform;
[0007] 3. With the rise and development of artificial intelligence and edge data intelligence technologies, the traditional architecture is increasingly unable to meet the requirements of big data processing from edge to cloud and cloud to edge, as well as the processing of edge sensing data.
[0008] As disclosed in Chinese Patent CN109308216B, a real-time task scheduling method for a single-core system for inaccurate calculations is provided. This method establishes a task model according to the task characteristics in a single-core system; constructs an offline scheduling module and an online scheduling module according to the task scheduling time points in the task model; constructs a task calculation accuracy unit for the online scheduling module according to the execution accuracy of the task; constructs a task calculation average degree of exceeding the deadline unit for the online scheduling module according to the average degree of the task exceeding the deadline; constructs a task exceeding the deadline frequency calculation unit for the online scheduling module according to the frequency of the task exceeding the deadline. This method can only make more full use of the computing resources of the system by separately ensuring the calculation accuracy, the degree of exceeding the deadline, and the upper limit of the probability of exceeding the deadline, on the premise of ensuring a certain QoS (Quality of Service), and cannot improve the system computing efficiency by using the multi-core architecture and multi-core performance. Summary of the Invention
[0009] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a GPU and NPU hybrid operation method, medium and system applied to a security platform, which improves the operation efficiency through a multi-core architecture and multi-core performance.
[0010] The purpose of the present invention can be achieved through the following technical solutions:
[0011] A GPU and NPU hybrid operation method applied to a security platform includes the following steps:
[0012] According to the characteristics of the input data, the input data is respectively entrusted to the GPU and the NPU for processing, wherein,
[0013] If the input data is the data received through the communication interface, using the OPENCL technology, the input data is divided into different core modules according to different interface types, and the core modules are entrusted to the GPU core for calculation to obtain a data result,
[0014] If the input data is the data detected by the anti-collision unit, the data is entrusted to the NPU for discrimination to obtain a discrimination result, and an obstacle recognition model is set on the NPU;
[0015] Adopt the DMA method to map the data result and the discrimination result to the CPU memory unit, and output the calculation result and the discrimination result in the CPU memory unit to different output interfaces for data output.
[0016] Further, the specific steps of obtaining the input data include:
[0017] Receive external data, unpack the received external data, convert the external data from its transmission format to a processable format, perform data verification to verify data integrity and correctness, and store it as input data after completing the data verification.
[0018] Further, the interface types include physical interfaces and logical interfaces.
[0019] Further, the physical interfaces include multiple types among Ethernet, serial ports, RS485 interfaces, and CAN interfaces.
[0020] Further, the logical interfaces include multiple types among inter - process communication, FIFO, and circular buffers.
[0021] Further, the GPU core is an independent computing unit, and the independent computing unit includes multiple stream processors.
[0022] Further, the anti - collision unit includes lidar, IMU, odometer, and radio frequency devices.
[0023] Further, the data detected by the anti - collision unit includes laser SLAM data and radio frequency data.
[0024] Further, when mapping the data result and discrimination result to the CPU memory unit in DMA mode, ensure the data consistency between the CPU memory unit and the CPU cache through the MESI protocol.
[0025] Further, after mapping the data result and discrimination result to the CPU memory unit in DMA mode, perform data consistency check on the data in the CPU memory unit through the interrupt mechanism.
[0026] Further, the output interface includes a communication interface and a digital IO interface.
[0027] Further, the specific steps for data output of the digital IO interface include:
[0028] Determine the data represented by each bit field, combine the bit fields into a complete byte, and output the combined byte to the digital IO interface; control the level of the digital IO interface according to the voltage and current requirements between different devices.
[0029] Further, the specific steps for data output of the communication interface include:
[0030] Determine the format and structure of the data packet according to the communication protocol, create the data packet header, add the actual data to be transmitted as the payload to the data packet, add the data packet tail, and serialize each part of the data packet into a byte stream; calculate the check value for the payload of the data packet using a check algorithm and append the check value to the tail of the data packet or send it as a separate field; send the packetized data through the communication interface.
[0031] Further, the control information in the data packet header includes multiple types among the source address, destination address, and data length.
[0032] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement a GPU and NPU hybrid operation method applied to a security platform.
[0033] A GPU and NPU hybrid operation system applied to a security platform, comprising:
[0034] A task division module, configured to respectively entrust the input data to the GPU and NPU for processing according to the characteristics of the input data, where
[0035] If the input data is the data received through the communication interface, using the OPENCL technology, divide the input data into different core modules according to different interface types, entrust the core modules to the GPU core for calculation to obtain a data result,
[0036] If the input data is the data detected by the anti-collision unit, entrust the data to the NPU for discrimination to obtain a discrimination result, and an obstacle recognition model is set on the NPU;
[0037] A data mapping module, configured to map the data result and the discrimination result to the CPU memory unit in the DMA mode;
[0038] A data output module, configured to output the calculation result and the discrimination result in the CPU memory unit to different output interfaces for data output.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] 1. The present invention divides tasks according to the characteristics of input data, and separately entrusts the input data to the GPU and NPU for processing; adopting the OPENCL technology, the input data received through the communication interface is divided into different core modules according to different interface types, and the core modules are entrusted to the GPU core for calculation to obtain data results; the input data detected by the anti-collision unit is entrusted to the NPU for discrimination, and the data results and discrimination results are mapped to the CPU memory unit by using the DMA method to obtain the discrimination results, which are output to different output interfaces for data output, improving the operation efficiency of secure operation.
[0041] 2. The present invention is not limited to a certain model of chip integrated with a multi-core heterogeneous CPU, GPU, and NPU. According to the multi-core heterogeneous architecture, corresponding solutions can be flexibly adopted to ensure that the chip performance is fully exerted. When the chip is upgraded, the present invention can be slightly changed to adapt to the new type of chip, improving the adaptability and reliability of secure operation. Brief Description of the Drawings
[0042] Figure 1 It is a schematic flowchart proposed by the present invention. Detailed Embodiments
[0043] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.
[0044] Abbreviations involved:
[0045] Graphics Processing Unit: GPU
[0046] Neural network Processing Unit: NPU
[0047] Central Processing Unit: CPU
[0048] Open Computing Language: OpenCL
[0049] Direct Memory Access: DMA
[0050] First Input First Output: FIFO
[0051] Interrupt Handler: Interrupt Service Routine, ISR
[0052] Embodiment 1
[0053] This embodiment provides a GPU and NPU hybrid operation method applied to a security platform, as Figure 1 shown, including the following steps:
[0054] S1. According to the characteristics of the input data, the input data is respectively entrusted to the GPU and NPU for processing.
[0055] The specific steps for obtaining the input data include: receiving external data, unpacking the received external data, converting the external data from its transmission format to a processable format, performing data verification, verifying data integrity and correctness, and storing the data after completing the data verification.
[0056] If the input data is the data received through the communication interface, using the OPENCL technology, the input data is divided into different core modules according to different interface types, and the core modules are entrusted to the GPU core for calculation to obtain the data result.
[0057] The interface types of the communication interface include physical interfaces and logical interfaces. The physical interfaces include multiple types such as Ethernet, serial port, RS485 interface, and CAN interface. The logical interfaces include multiple types such as inter-process communication, FIFO, and circular buffer.
[0058] The input data from the physical interface is input into the physical interface core module for data processing and operation, and the input data from the logical interface is input into the logical interface core module and entrusted to the GPU for data processing and operation.
[0059] The GPU core is an independent computing unit, and the independent computing unit includes multiple stream processors.
[0060] If the input data is the data detected by the anti-collision unit, the data is entrusted to the NPU for discrimination to obtain the discrimination result.
[0061] The anti-collision unit includes lidar, IMU, odometer, and radio frequency device. The data detected by the anti-collision unit includes laser SLAM data and radio frequency data.
[0062] An obstacle recognition model is set on the NPU, and the data detected by the anti-collision unit is input into the obstacle recognition model for obstacle discrimination to obtain the discrimination result.
[0063] S2. Use the DMA method to map the data result and the discrimination result to the CPU memory unit.
[0064] After the DMA mapping is completed, data consistency checking is performed on the data in the CPU memory unit through the interrupt mechanism.
[0065] The specific steps for performing data consistency checking on the data in the CPU memory unit through the interrupt mechanism are as follows: When the DMA controller completes the data transfer task, it triggers an interrupt signal to the CPU; after receiving the interrupt signal, the CPU pauses the currently executing task and saves the context of the current task (such as the program counter and register status); the CPU jumps to the ISR to process the data consistency checking after the DMA transfer is completed; in the ISR, the CPU can check the status of the DMA transfer to confirm whether the data is correctly transferred to the target address. By comparing the expected data with the actually transferred data, or by checking the checksum to verify the integrity and correctness of the data. If data inconsistency is found, the ISR can execute error handling logic, such as retrying the transfer, logging the error, sending an error notification, etc. After completing the data consistency checking, the ISR restores the previously saved task context and returns to the interrupted task to continue execution.
[0066] S3. Output the calculation result and the discrimination result in the CPU memory unit to different output interfaces for data output.
[0067] When the DMA transfers data, it bypasses the CPU and directly accesses the CPU memory unit, which may cause the data in the CPU cache to be inconsistent with the data in the CPU memory unit. In this embodiment, the MESI (Modified Exclusive Shared Invalid) protocol is used to ensure the data consistency between the CPU cache and the CPU memory unit.
[0068] The data output interface includes a communication interface and a digital IO interface.
[0069] The specific steps for data output of the digital IO interface are as follows:
[0070] Determine the data represented by each bit field, combine the bit fields into a complete byte, and output the combined byte to the digital IO interface; control the level of the digital IO interface according to the voltage and current requirements between different devices.
[0071] The specific steps for data output of the communication interface include: determining the format and structure of the data packet according to the communication protocol, creating the data packet header, and the control information in the data packet header includes multiple types among the source address, destination address, and data length; adding the actual data to be transmitted as the payload to the data packet, adding the data packet tail, and serializing each part of the data packet into a byte stream; selecting an appropriate verification algorithm according to the protocol requirements, such as CRC, Checksum, MD5, and SHA, etc., to calculate the verification value for the payload of the data packet, and appending the verification value to the tail of the data packet or sending it as a separate field; sending the data after packet assembly through the communication interface.
[0072] Embodiment 2
[0073] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the GPU and NPU hybrid operation method applied to the security platform.
[0074] The rest is the same as Embodiment 1.
[0075] Embodiment 3
[0076] This embodiment provides a GPU and NPU hybrid operation system applied to the security platform, including:
[0077] A task division module, configured to respectively entrust the input data to the GPU and NPU for processing according to the characteristics of the input data, where,
[0078] If the input data is the data received through the communication interface, using the OPENCL technology, dividing the input data into different core modules according to different interface types, and entrusting the core modules to the GPU core for calculation to obtain the data result,
[0079] If the input data is the data detected by the anti-collision unit, entrusting the data to the NPU for discrimination to obtain the discrimination result, and an obstacle recognition model is set on the NPU;
[0080] A data mapping module, configured to map the data result and the discrimination result to the CPU memory unit by using the DMA method;
[0081] A data output module, configured to output the calculation result and the discrimination result in the CPU memory unit to different output interfaces for data output.
[0082] The rest is the same as Embodiment 1.
[0083] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in this technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art shall fall within the protection scope determined by the claims.
Claims
1. A GPU and NPU hybrid computing method applied to a security platform, characterized in that: The following steps are involved: According to the characteristics of the input data, the input data is respectively hosted on the GPU and the NPU for processing, wherein: If the input data is data received through a communication interface, OPENCL technology is used to divide the input data into different core modules according to different interface types, and the core modules are hosted in the GPU core for calculation to obtain data results. If the input data is data obtained through detection by the anti-collision unit, the data is hosted on the NPU for identification to obtain an identification result, and the NPU is provided with an obstacle identification model; The data results and the discrimination results are mapped to the CPU memory unit by using the DMA method, and the calculation results and the discrimination results in the CPU memory unit are output to different output interfaces for data output.
2. The GPU and NPU hybrid computing method applied to a security platform according to claim 1, characterized in that: The specific steps of obtaining the input data include: Receive external data, unpack the received external data, convert the external data from its transmission format into a processable format, perform data verification, verify the data integrity and correctness, and store it as input data after completing data verification.
3. The GPU and NPU hybrid computing method applied to a security platform according to claim 1, characterized in that: The interface types include physical interfaces and logical interfaces.
4. The GPU and NPU hybrid computing method applied to a security platform according to claim 3, characterized in that: The physical interface includes multiple ones of Ethernet, serial port, RS485 interface and CAN interface.
5. The GPU and NPU hybrid computing method applied to a security platform according to claim 3, characterized in that: The logical interface includes multiple ones of inter-process communication, FIFO and ring buffer.
6. The GPU and NPU hybrid computing method applied to a security platform according to claim 1, characterized in that: The GPU core is an independent computing unit, and the independent computing unit includes multiple stream processors.
7. The GPU and NPU hybrid computing method applied to a security platform according to claim 1, characterized in that: The anti-collision unit includes a laser radar, an IMU, an odometer and a radio frequency device.
8. The GPU and NPU hybrid computing method applied to a security platform according to claim 1, characterized in that: The data detected by the anti-collision unit includes laser SLAM data and radio frequency data.
9. The GPU and NPU hybrid computing method applied to a security platform according to claim 1, characterized in that: When the data result and the discrimination result are mapped to the CPU memory unit by adopting the DMA mode, the data consistency between the CPU memory unit and the CPU cache is ensured by the MESI protocol.
10. The GPU and NPU hybrid computing method applied to a security platform according to claim 1, characterized in that: After the data result and the discrimination result are mapped into the CPU memory unit by adopting the DMA mode, the data consistency check is performed on the data in the CPU memory unit by using the interrupt mechanism.
11. The GPU and NPU hybrid computing method applied to a security platform according to claim 1, characterized in that: The output interface includes a communication interface and a digital IO interface.
12. The GPU and NPU hybrid computing method applied to a security platform according to claim 11, characterized in that: The specific steps of data output of the digital IO interface include: Determine the data represented by each bit field, combine the bit fields into a complete byte, and output the combined byte to the digital IO interface; control the level of the digital IO interface according to the voltage and current requirements between different devices.
13. The GPU and NPU hybrid computing method applied to a security platform according to claim 11, characterized in that: The specific steps of data output of the communication interface include: Determine the format and structure of the data packet according to the communication protocol, create a data packet header, add the actual data to be transmitted as a payload to the data packet, add a data packet tail, and serialize the various parts of the data packet into a byte stream; use a checksum algorithm to calculate a checksum value for the payload of the data packet, and append the checksum value to the tail of the data packet or send it as a separate field; and send the packetized data through the communication interface.
14. The GPU and NPU hybrid computing method applied to a security platform according to claim 13, characterized in that: The control information in the data packet header includes multiple types of source address, destination address and data length.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it can implement the GPU and NPU hybrid computing method applied to the security platform as described in any one of claims 1 to 14.
16. A GPU and NPU hybrid computing system applied to a security platform, characterized in that: include: The task division module is used to delegate the input data to the GPU and NPU for processing according to the characteristics of the input data, wherein: If the input data is data received through a communication interface, OPENCL technology is used to divide the input data into different core modules according to different interface types, and the core modules are hosted in the GPU core for calculation to obtain data results. If the input data is data obtained through detection by the anti-collision unit, the data is hosted on the NPU for identification to obtain an identification result, and the NPU is provided with an obstacle identification model; A data mapping module, used for mapping the data result and the discrimination result to the CPU memory unit by adopting DMA mode; The data output module is used to output the calculation results and the judgment results in the CPU memory unit to different output interfaces for data output.
Citation Information
Patent Citations
A Real-Time Task Scheduling Method for Single-Core Systems with Inaccurate Computation
CN109308216B