Heterogeneous multi-core sealing microsystem for satellite-borne on-orbit intelligent calculation and fusion method
By designing a heterogeneous multi-core integrated microsystem for satellites, integrating scalar, vector and tensor computing units, and equipped with real-time scheduling units and global shared memory, the existing satellite-on-mounted intelligent processor architecture has solved the problems of low energy efficiency and large size during intelligent computing, and achieved high energy efficiency, low latency and radiation-resistant intelligent computing capabilities.
Patent Information
- Application Number
- CN202510574831.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
AI Technical Summary
The existing satellite-borne intelligent processor architecture has problems such as low energy efficiency, large size, large mass, large power consumption, and the inability to support signal processing and data processing calculations in satellite-borne scenarios when performing intelligent computing.
A heterogeneous multi-core integrated microsystem for satellite-mounted on-orbit intelligent computing is designed, integrating multi-core scalar, vector and tensor intelligent computing units, equipped with real-time and efficient computing scheduling units, and real-time memory sharing is realized through global shared memory, reducing data exchange costs and delays. At the same time, thermal backup and protective reinforcement treatment are set up for the anti-irradiation, anti-cosmic rays and high-energy particles requirements of the on-site environment.
It realizes parallel heterogeneous computing under the premise of high energy efficiency ratio, meets the complex needs of satellite-based intelligent computing, reduces the system's volume, quality and energy consumption, and improves the system's reliability and safety.
Smart Images

Figure CN120086026A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electrical data processing, and particularly relates to a heterogeneous multi-core co-packaged microsystem and a fusion method for on-orbit intelligent computing in spaceborne applications. Background Art
[0002] Spaceborne intelligent processors are key components in satellite systems. They are used for real-time data processing, key information extraction, supporting complex algorithm operations, autonomous decision-making and control, mission planning and scheduling, optimizing data transmission, supporting high-speed communication, fault diagnosis and reconstruction, enhancing anti-interference capabilities, etc. during satellite operation.
[0003] Currently, typical spaceborne intelligent processor architectures include the following:
[0004] 1. Multi-core CPU architecture: It has advantages such as ecological friendliness, flexible programming, and rich human-computer interfaces, and can also be used to support intelligent computing. However, due to the complex architecture of the CPU core and the small proportion of the logical quantity of the computing unit in the processor, when implementing intelligent computing, the CPU has a relatively low energy efficiency ratio.
[0005] 2. CPU + GPU processing architecture: The main form of this architecture is to add a GPU card to a PC server, and the communication and control between the main control CPU and the acceleration unit GPU are realized through high-speed buses such as PCIe. The main control CPU and the GPU each have an independent DDR memory. This architecture can make full use of the flexibility of the CPU and the high computing power of the GPU, and is currently the mainstream solution for cloud intelligent computing. In spaceborne scenarios, it has problems such as large volume, large mass, and high power consumption.
[0006] 3. CPU + NPU processing architecture: The main form of this architecture is to add an NPU card to a PC server, and the communication and control between the main control CPU and the acceleration unit NPU are realized through high-speed buses such as PCIe. The main control CPU and the NPU each have an independent DDR memory. This architecture can make full use of the flexibility of the CPU and the high computing power of the NPU, and has been applied to a certain extent in pre-cloud intelligent computing. Compared with the PC + GPU solution, the NPU has a better "performance / power consumption" ratio when implementing deep learning computing. However, in spaceborne scenarios, it still needs to face a large number of data processing tasks in multi-spectral and hyperspectral signals.
[0007] 4. CPU + GPU + NPU processing architecture: In this architecture, the CPU is the main control, and at the same time, the GPU and NPU coprocessors are integrated to achieve data processing and neural network computing. It is an energy-efficient spaceborne intelligent computing architecture. However, if a spaceborne intelligent computing system is built with a CPU motherboard + GPU PCIe card + NPU PCIe card, it will have a large volume, mass, and energy consumption.
[0008] Therefore, the above architectures have the following defects:
[0009] 1. The deep learning computing efficiency of the multi-CPU architecture is extremely low;
[0010] 2. The computing efficiency of the CPU+GPU architecture is not high enough, and it has the disadvantages of large volume, large mass, and high power consumption;
[0011] 3. The flexibility and application ecosystem of the CPU+NPU architecture are poor, and it cannot support signal processing and data processing calculations in on-orbit scenarios.
[0012] 4. The CPU+GPU+NPU architecture has a large volume, mass, and energy consumption, and it is difficult to perform radiation-hardening and reinforcement processing. Summary of the Invention
[0013] To solve the above technical problems, the present invention provides a heterogeneous multi-core co-packaged microsystem and fusion method for on-orbit intelligent computing in space. Compared with traditional single-core processors, multi-core homogeneous processors, and "host-accelerator" architectures with a CPU as the core, the present invention integrates multi / single-core scalar, vector, and tensor intelligent computing units on a single chip to meet the data processing, intelligent reasoning, and decision-making control requirements in on-orbit intelligent application scenarios. It is equipped with a real-time and efficient computing scheduling unit, which can perform parallel heterogeneous computing under the premise of high energy efficiency ratio. According to the requirements of radiation protection, cosmic ray protection, and high-energy particle protection in on-orbit scenarios, thermal backup and protection and reinforcement processing are set up.
[0014] To achieve the above object, the present invention adopts the following technical solutions:
[0015] A heterogeneous multi-core co-packaged microsystem for on-orbit intelligent computing in space, including a communication interface unit, a multi-core scalar computing unit, a multi-core vector computing unit, a multi-core tensor computing unit, a real-time scheduling unit, and a global shared memory; the communication interface unit is connected to the multi-core scalar computing unit to form an external communication interface of the heterogeneous multi-core co-packaged microsystem; after receiving computing instructions and data, the multi-core scalar computing unit creates a computing task table in the global shared memory, and distributes the computing tasks to the multi-core scalar computing unit, the multi-core vector computing unit, and the multi-core tensor computing unit for computing through the real-time scheduling unit, and collects the computing results; the real-time scheduling unit obtains a computing task list from the computing task table, queries the idle status of the multi-core scalar computing unit, the multi-core vector computing unit, and the multi-core tensor computing unit, and distributes the computing tasks to the multi-core scalar computing unit, the multi-core vector computing unit, and the multi-core tensor computing unit for computing according to the priority control mechanism; the global shared memory is used to realize memory sharing.
[0016] The present invention also provides a fusion method for a heterogeneous multi-core co-packaged microsystem for on-orbit intelligent computing in space, which fuses multi-core scalar computing, multi-core vector computing units, and multi-core tensor computing units, including:
[0017] Use YOLOv5 to perform object detection on the input video frames to obtain the initial position and category information of the objects; initialize the tracking module of the TLD algorithm according to the detection results of YOLOv5;
[0018] Adopt the tracking module of the TLD algorithm to start real-time tracking of the objects according to the initialized object positions;
[0019] During the tracking process, if the tracking module of the TLD algorithm determines that the object is lost or the tracking fails, then re-call YOLOv5 for object detection; YOLOv5 re-detects the positions and categories of the objects and feeds the detection results back to the tracking module of the TLD algorithm to re-initialize the tracking module; the detection module of the TLD algorithm also detects the video frames and makes a comprehensive judgment with the results of the tracking module to determine the final positions of the objects;
[0020] The learning module of the TLD algorithm dynamically updates the feature points and model parameters of the objects according to the results of the tracking module and the detection module;
[0021] Process the video frames frame by frame and repeat the above process until the video ends.
[0022] Furthermore, the heterogeneous computing in the field of on-orbit computing in space is not independent of each other and requires mutual fusion and data exchange. Therefore, the present invention is equipped with a global shared memory to achieve memory sharing, reduce the data exchange cost and latency. At the same time, the heterogeneous multi-core co-packaged microsystem architecture needs to be equipped with a high-performance real-time task scheduling module to manage the computing tasks and timely allocate the computing tasks to their respective multi-core computing resources for computing.
[0023] Furthermore, in view of the particularity of the spaceborne environment, it is necessary to protect against irradiation, cosmic rays, high-energy particles, etc. to ensure the safety of the system. In the heterogeneous multi-core co-packaged microsystem architecture, hot backup configurations are implemented for core computing units such as CPUs and NPUs, and physical means are used for reinforcement processing.
[0024] Beneficial effects:
[0025] 1. The present invention specifically allocates scalar, vector, and tensor computing power resources in a single co-packaged microsystem to meet the complex computing task requirements of on-orbit intelligent computing in space;
[0026] 2. The present invention is equipped with a high-performance real-time task scheduling unit to achieve high-efficiency and low-latency computing task scheduling;
[0027] 3. The present invention is equipped with a global shared memory, reducing the data bandwidth requirements between computing units;
[0028] 4. Achieve miniaturized packaging and radiation-hardening of the microsystem to meet the requirements of the on-orbit application environment for satellites;
[0029] 5. The present invention adopts a high-performance, small-size, and light-weight solution, which is suitable for on-orbit application scenarios for satellites.
[0030] In summary, the present invention provides a specific application solution for the multi-modal fusion algorithm, which can be widely applied to civilian and industrial product models under this architecture, and can generally be applied to other chip architectures with multi-modal fusion solutions in the market. After verification by the actual product solution, the fusion algorithm is superior to the algorithms running on a single computing unit in terms of detection accuracy (mAP) and tracking success rate (Precision). It has advantages such as high-precision detection, robust tracking, real-time performance, and strong adaptability, which are not possessed by current single algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is the overall architecture diagram of a heterogeneous multi-core co-packaged microsystem for on-orbit intelligent computing in satellite applications according to an embodiment of the present invention;
[0032] Figure 2 is the task scheduling flowchart of the real-time scheduling unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0034] In view of the data processing, intelligent reasoning, and decision-making control requirements in the on-orbit intelligent computing scenario for satellites, and considering factors such as the respective capabilities, underlying computing operator frameworks, and energy efficiency ratios of processor units with different architectures, the present invention provides a heterogeneous multi-core co-packaged microsystem and a fusion method for on-orbit intelligent computing in satellite applications, which adopt a domain-specific architecture (Domain Specification Architecture, DSA) for on-orbit computing in satellite applications. The DSA architecture is designed to be optimized for the specific requirements of on-orbit intelligent computing in satellite applications, not only meeting the requirements of on-orbit intelligent computing in satellite applications but also providing support for the efficient deployment of multi-modal fusion intelligent algorithms.
[0035] According to the requirements of on-orbit intelligent computing for spaceborne applications, this invention optimally configures scalar, vector, and tensor computing units to ensure high performance, low energy consumption, and low latency. Through a real-time scheduling unit, tasks are dynamically allocated to each computing unit to improve the overall efficiency of the system. A global shared memory is equipped to reduce the data exchange cost and latency. At the same time, considering the particularity of the spaceborne environment, thermal backup and physical reinforcement are carried out to ensure the reliability and security of the system.
[0036] As Figure 1 shown, an heterogeneous multi-core co-packaged microsystem for on-orbit intelligent computing for spaceborne applications provided by the DSA architecture in an embodiment of this invention includes:
[0037] Multi-core scalar computing units (i.e., the scalar computing units in Figure 1 , Figure 2 ) are used to execute scalar computing tasks such as attitude estimation, region extraction, spectral band selection, auxiliary data arrangement, etc. The computing processes of these tasks are complex but with small data volumes, and are suitable for using a CPU processor.
[0038] Multi-core vector computing units (i.e., the vector computing units in Figure 1 , Figure 2 ) are used to process vector algorithms such as relative radiometric correction, multi-spectral registration, geometric positioning, atmospheric correction, etc. These algorithms need to process a large amount of data and are suitable for using a GPU that supports SIMD (Single Instruction Multiple Data) and SIMT (Single Instruction Multiple Threads) instructions.
[0039] Multi-core tensor computing units (i.e., the tensor computing units in Figure 1 , Figure 2 ) are used to execute tensor algorithms such as cloud detection, spatio-spectral joint target detection, terrain and landform recognition, etc. These algorithms are based on deep neural networks and are suitable for using an NPU (Neural Network Processing Unit).
[0040] Global shared memory (DDR) realizes memory sharing and reduces the data exchange cost and latency.
[0041] A real-time scheduling unit is responsible for task scheduling and allocates computing tasks to each computing unit.
[0042] A communication interface unit is used for data communication with the scalar computing unit.
[0043] Preferably, considering the particularity of the spaceborne environment, thermal backup configuration is carried out for core computing units such as the CPU and NPU, and physical means are used for reinforcement processing.
[0044] The multi-core scalar computing units, multi-core vector computing units, and multi-core tensor computing units each include local memory for storing their respective data.
[0045] Specifically, the communication interface unit is connected to the multi-core scalar computing unit to form an external communication interface for the heterogeneous multi-core integrated microsystem, including obtaining sensor data, instruction data from other devices and returning calculation and intelligent recognition results, etc.
[0046] Specifically, after receiving calculation instructions and data, the multi-core scalar computing unit creates a calculation task table in the global shared memory, and distributes the calculation tasks to each computing unit for calculation through the real-time scheduling unit, and collects the calculation results.
[0047] Specifically, the multi-core vector computing unit and the multi-core tensor computing unit are respectively connected to the global shared memory, and are used to execute the corresponding vector algorithms and tensor algorithms to process the image data in on-orbit satellite computing.
[0048] Specifically, the real-time scheduling unit obtains the calculation task list from the calculation task table, and distributes the calculation tasks to each computing unit for calculation according to the priority control mechanism by querying the idle status of each computing unit (NPU / CPU / GPU).
[0049] Specifically, the global shared memory is used to achieve memory sharing, reduce the data exchange cost and delay, and provide access and storage of data for each computing unit.
[0050] Specifically, the local memory provides local storage space for each computing unit, and is used to store temporary data and intermediate results.
[0051] Specifically, for the scalar algorithms in the multi-core scalar computing unit, its processor unit uses a CPU; it processes tasks with complex calculation processes but small data volumes, such as attitude estimation, region extraction, spectral band selection, auxiliary data arrangement, etc. Its underlying calculation operator framework is for general mathematical calculations and trigonometric function calculations. Due to the small amount of task data, the flexibility and versatility of the CPU make it have a higher energy efficiency ratio in these tasks.
[0052] Specifically, for the vector algorithms of the multi-core vector computing unit, its processor unit uses a GPU to process the real-time data of large image frames, multi-channels and large data streams generated by multi / hyperspectral sensors, such as relative radiometric correction, multi-spectral registration, geometric positioning, atmospheric correction, etc. Its underlying calculation operator framework supports vector calculations of SIMD and SIMT instructions. The GPU has a high energy efficiency ratio when processing large-scale parallel data and is suitable for processing vector algorithms.
[0053] Specifically, for the tensor algorithms of the multi-core tensor computing unit, its processor unit uses an NPU to execute algorithms based on deep neural networks, such as cloud detection, spatio-spectral joint target detection. The NPU has a higher performance / power ratio in deep learning calculations and is suitable for processing tensor algorithms.
[0054] Such asFigure 2 As shown in Figure 2 , in addition to performing scalar calculation tasks, the multi-core scalar calculation unit is also responsible for the system's external communication interface, including obtaining sensor data, instruction data from other devices, and returning calculation and intelligent recognition results, etc. After receiving the calculation instructions and data, the multi-core scalar calculation unit creates a calculation task table in the global shared memory. The real-time scheduling unit obtains the calculation task list from the calculation task table, queries the idle status of each heterogeneous calculation unit, and according to the priority control mechanism, allocates the calculation tasks to each calculation unit for calculation, and collects the calculation results. According to the topological structure of the calculation graph, after a node calculation is completed, it will trigger the next node to start calculating until the entire calculation graph is completed and the calculation results are output.
[0055] Specifically, the core of the priority control mechanism is to assign a priority value to each task, event or request. The higher the priority value, the more urgent or important the task is and should be processed first. The priority can be static (determined when the task is created and does not change over time) or dynamic (dynamically adjusted according to the execution situation of the task or the system state).
[0056] In the present invention, the task waiting queue and the task execution queue are important components of the real-time scheduling unit, which are used to efficiently manage and schedule calculation tasks. The following are the specific contents and working principles of the task waiting queue and the task execution queue:
[0057] 1. The task waiting queue is a queue that stores tasks that have not been assigned to any calculation unit. These tasks are ready to execute, but there are no available calculation resources currently. The main role of the task waiting queue is to ensure that tasks can be managed and scheduled orderly while waiting for execution, and it includes:
[0058] Task object: Each task object contains detailed information about the task, such as task ID, task type (scalar, vector, tensor), input data pointer, output data pointer, priority, etc.
[0059] Task dependency: Some tasks may depend on the output results of other tasks. These dependencies will be recorded in the task object to ensure that the execution order of tasks is logical.
[0060] Task priority: Each task has a priority, and tasks with higher priority will be taken out from the waiting queue first and assigned to the execution queue.
[0061] The working principle of the task waiting queue is as follows:
[0062] Task enqueue: When a new task is generated, if there are no available calculation resources currently, the task will be put into the task waiting queue.
[0063] Task Sorting: The tasks in the task waiting queue are sorted according to their priorities. Tasks with higher priorities are placed at the front of the queue and are scheduled first.
[0064] Task Scheduling: The real-time scheduling unit periodically checks the task waiting queue. When computing resources are idle, it takes out the task with the highest priority from the queue and moves it to the task execution queue.
[0065] 2. The task execution queue is a queue that stores tasks that have been assigned to computing units but have not yet completed execution. These tasks are ready to be executed on the corresponding computing units but may be waiting for execution due to the busyness of the computing units. The main function of the task execution queue is to ensure that tasks can be orderly assigned to computing units and efficiently executed on them. It includes:
[0066] Task Object: Similar to the task objects in the task waiting queue, each task object contains detailed information about the task, such as task ID (name), task type, input data pointer, output data pointer, etc.
[0067] Computing Unit Allocation Information: The task object records the ID of the computing unit to which the task is assigned and the status of that computing unit (idle, busy).
[0068] Task Status: The task object records the current status of the task, such as "assigned", "executing", "completed", etc.
[0069] The working principle of the task execution queue is as follows:
[0070] Task Allocation: When the real-time scheduling unit takes out a task from the task waiting queue, it assigns the task to a suitable computing unit according to the task type and the current status of the computing unit, and puts the task object into the task execution queue.
[0071] Task Execution: The computing unit takes out the task object from the task execution queue and starts executing the task. During the task execution process, the computing unit updates the status of the task object.
[0072] Task Completion: After the task execution is completed, the computing unit updates the status of the task object to "completed" and removes the task object from the task execution queue. If the output result of the task is the input of other tasks, it will trigger the scheduling of dependent tasks.
[0073] The specific process of the real-time scheduling unit is as follows:
[0074] Task Enqueue: The task object is put into the task waiting queue.
[0075] Task Sorting: The task waiting queue is sorted according to the priorities of the tasks.
[0076] Task scheduling: The real-time scheduling unit checks the task waiting queue and finds the task with the highest priority.
[0077] Task allocation: The real-time scheduling unit allocates tasks to appropriate computing units according to the task type and the current state of the computing units, and puts the task objects into the task execution queue.
[0078] Task execution: The computing unit takes out the task object from the task execution queue and starts to execute the task.
[0079] Task completion: After the task execution is completed, the computing unit updates the status of the task object to "completed" and removes the task object from the task execution queue. If the output result of a task is the input of other tasks, it will trigger the scheduling of dependent tasks.
[0080] Through this mechanism, the DSA architecture can efficiently manage and schedule tasks, ensuring the full utilization of computing resources and the orderly execution of tasks.
[0081] Preferably, the present invention also provides an integrated multi-core heterogeneous system compilation toolchain for performing integrated processing on the user's deep learning model implementation, etc., reducing the workload of developers during the implementation of projects, and focusing their energy on the algorithm itself rather than getting stuck in the programming and debugging work of details such as the hardware structure and instruction system.
[0082] Preferably, the present invention also provides an integrated intelligent programming framework, providing underlying support, integrated control, and fusion examples to help R & D personnel quickly understand the algorithm principle, quickly deploy the algorithm model, and quickly verify the algorithm effect.
[0083] In response to the requirements of the on-orbit intelligent field of satellites, the present invention realizes a high-performance, low-power, and low-latency intelligent computing solution by specifically optimizing the configuration of multi-core scalar computing, multi-core vector computing units, multi-core tensor computing units, real-time scheduling units, and global shared memory, and realizes protection against irradiation, cosmic rays, and high-energy particles in response to the particularity of the satellite-borne environment to meet the requirements of satellite-borne intelligent computing.
[0084] The present invention adopts a multi-modal fusion algorithm that combines deep learning object detection and TLD (Tracking-Learning-Detection). The core of the algorithm is to combine deep learning object detection (such as YOLOv5) with the TLD algorithm to achieve stable recognition and tracking of the object recognition algorithm under strong interference conditions. Taking YOLOv5 as an example, its principle mainly includes the following aspects:
[0085] In the NPU, deep learning object detection (such as YOLOv5) includes:
[0086] Automatically learn image features using a Convolutional Neural Network (CNN) to achieve the localization and classification of targets. YOLOv5 adopts a One-stage (single-stage) detection method and can quickly detect the position and category of targets through four parts: the input layer, BackBone (main network), Neck (neck), and DetHead (detection head). It features fast speed, high accuracy, and strong adaptability, and is especially suitable for complex background and multi-target scenarios.
[0087] Adopt the TLD algorithm in the CPU and GPU. TLD is a single-target long-term tracking algorithm that combines a tracking module, a detection module, and a learning module. The tracking module is responsible for real-time tracking of the target's position; the detection module is used to re-detect the target when the target is lost or occluded; the learning module updates the target's feature points and model parameters through an online learning mechanism to improve the robustness of tracking. TLD can adapt to the appearance changes of the target (such as lighting changes, deformation, etc.) and re-capture the target after a short occlusion.
[0088] Fuse the high-precision detection ability of YOLOv5 and the robust tracking ability of TLD to form complementary advantages. YOLOv5 provides the high-precision initial target position, and TLD is responsible for real-time tracking; when the target is occluded, deformed, or changes in scale, YOLOv5 can correct the tracking error of TLD to avoid tracking drift.
[0089] The method of this invention for fusing multi-core scalar computing, multi-core vector computing units, and multi-core tensor computing units includes:
[0090] (1) Initialization stage: Use YOLOv5 to perform target detection on the input video frames to obtain the initial position and category information of the targets. Initialize the tracking module of the TLD algorithm according to the detection results of YOLOv5 and set the initial parameters of the tracking module.
[0091] (2) Tracking stage: Adopt the tracking module of the TLD algorithm to perform real-time tracking of the target according to the initialized target position. The tracking module continuously tracks the target by predicting the position of the target in the next frame.
[0092] (3) Detection and correction stage: During the tracking process, if the tracking module of the TLD algorithm determines that the target is lost or the tracking fails (such as the target is occluded, deformed, etc.), then re-call YOLOv5 for target detection. YOLOv5 re-detects the position and category of the target and feeds the detection results back to the tracking module of the TLD algorithm to re-initialize the tracking module. The detection module of the TLD algorithm also detects the video frames and makes a comprehensive judgment with the results of the tracking module to determine the final position of the target.
[0093] (4) Learning and updating stage: The learning module of the TLD algorithm dynamically updates the feature points and model parameters of the target according to the results of the tracking module and the detection module. Through the online learning mechanism, the TLD algorithm can adapt to the appearance changes of the target and improve the robustness and accuracy of tracking.
[0094] (5) Loop processing: Process the video frames frame by frame, repeating the above loop process of tracking, detection, correction, and learning until the video ends. In each loop, according to the tracking and detection results of the current frame, dynamically adjust the tracking strategy of the next frame to ensure stable tracking of the target.
[0095] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A heterogeneous multi-core packaged microsystem for on-orbit intelligent computing, characterized in that: It includes a communication interface unit, a multi-core scalar computing unit, a multi-core vector computing unit, a multi-core tensor computing unit, a real-time scheduling unit, and a global shared memory; the communication interface unit is connected to the multi-core scalar computing unit to form an external communication interface of a heterogeneous multi-core sealed microsystem; after receiving computing instructions and data, the multi-core scalar computing unit creates a computing task table in the global shared memory, and distributes computing tasks to the multi-core scalar computing unit, the multi-core vector computing unit, and the multi-core tensor computing unit through the real-time scheduling unit for computing, and collects computing results; The real-time scheduling unit obtains the computing task list from the computing task table, queries the idle status of the multi-core scalar computing unit, the multi-core vector computing unit, and the multi-core tensor computing unit, and allocates the computing tasks to the multi-core scalar computing unit, the multi-core vector computing unit, and the multi-core tensor computing unit for calculation according to the priority control mechanism; the global shared memory is used to realize memory sharing.
2. A heterogeneous multi-core packaged microsystem for on-orbit intelligent computing according to claim 1, characterized in that: It also includes local memory, which provides local storage space for multi-core scalar computing units, multi-core vector computing units, and multi-core tensor computing units for storing temporary data and intermediate results.
3. The heterogeneous multi-core packaged microsystem for on-orbit intelligent computing according to claim 1, characterized in that: A scalar algorithm is used in the multi-core scalar computing unit, and its processor unit is a CPU.
4. The heterogeneous multi-core packaged microsystem for on-orbit intelligent computing according to claim 1, characterized in that: The multi-core vector computing unit adopts vector algorithms, and its processor unit adopts GPU.
5. The heterogeneous multi-core packaged microsystem for on-orbit intelligent computing according to claim 1, characterized in that: The multi-core tensor computing unit adopts tensor-type algorithms, and its processor unit adopts NPU to execute algorithms based on deep neural networks.
6. The heterogeneous multi-core packaged microsystem for on-orbit intelligent computing according to claim 1, characterized in that: The real-time scheduling unit includes a task waiting queue and a task execution queue. The task waiting queue stores the queue of tasks that have not yet been assigned to the multi-core scalar computing unit, the multi-core vector computing unit, and the multi-core tensor computing unit; the task execution queue stores the queue of tasks that have been assigned to one of the multi-core scalar computing unit, the multi-core vector computing unit, and the multi-core tensor computing unit but have not yet completed execution.
7. The heterogeneous multi-core packaged microsystem for on-orbit intelligent computing according to claim 6, characterized in that: The real-time scheduling unit realizes task queuing, task sorting, task scheduling, task allocation, task execution, and task completion. Through the real-time scheduling unit, the heterogeneous multi-core packaged microsystem manages and schedules tasks to ensure full utilization of computing resources and orderly execution of tasks.
8. The heterogeneous multi-core packaged microsystem for on-orbit intelligent computing according to claim 1, characterized in that: A multi-mode fusion algorithm that combines deep learning target detection and TLD algorithm is adopted; the TLD algorithm is a tracking-learning-detection algorithm.
9. The heterogeneous multi-core packaged microsystem for on-orbit intelligent computing according to claim 8, characterized in that: Deep learning target detection uses YOLOv5, which is applied to multi-core tensor computing units and uses convolutional neural networks to automatically learn image features to achieve target positioning and classification. YOLOv5 uses a single-stage detection method to quickly detect the location and category of the target through the input layer, backbone network, neck and detection head; The TLD algorithm is applied to multi-core scalar computing units and multi-core vector computing units. The TLD algorithm includes a tracking module, a detection module, and a learning module. The tracking module is responsible for tracking the position of the target in real time; the detection module is used to re-detect the target when it is lost or occluded; the learning module updates the feature points and model parameters of the target through an online learning mechanism to improve the robustness of tracking; The high-precision detection capability of YOLOv5 and the robust tracking capability of the TLD algorithm are combined. YOLOv5 provides high-precision initial target location, and the TLD algorithm is responsible for real-time tracking. When the target is occluded, deformed, or scaled, YOLOv5 corrects the tracking error of the TLD algorithm to avoid tracking drift.
10. A fusion method for heterogeneous multi-core packaged microsystems for on-orbit intelligent computing, characterized in that: The multi-core scalar computing, multi-core vector computing unit and multi-core tensor computing unit are integrated, including: Use YOLOv5 to detect the target on the input video frame and obtain the initial position and category information of the target; initialize the tracking module of the TLD algorithm based on the detection results of YOLOv5; The tracking module using TLD algorithm tracks the target in real time according to the initial target position; During the tracking process, if the tracking module of the TLD algorithm determines that the target is lost or the tracking fails, YOLOv5 is called again for target detection; YOLOv5 re-detects the position and category of the target, and feeds the detection results back to the tracking module of the TLD algorithm to reinitialize the tracking module; the detection module of the TLD algorithm also detects the video frame and makes a comprehensive judgment with the results of the tracking module to determine the final position of the target; The learning module of the TLD algorithm dynamically updates the target’s feature points and model parameters based on the results of the tracking module and the detection module; The video frames are processed frame by frame, and the above process is repeated until the video ends.
Citation Information
Patent Citations
Airborne embedded intelligent micro-processing system
CN112579512A
Method and system for dynamically scheduling multi-core fusion computing processor
CN114371933A
Multi-core heterogeneous processor chip, multi-core heterogeneous processor equipment, readable medium and article identification method
CN117573606A
Satellite-borne heterogeneous AI computer
CN118606028A
Space-based intelligent load calculation system
CN119357122A
Cited By
Satellite task planning computing power dynamic distribution system based on heterogeneous multi-core processor
CN120803752A