Processor Core Notification Circuit for Low Latency Register Updates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In processors with multiple cores, the existing methods for notifying register values across cores lead to increased latency and decreased arithmetic operation performance due to sequential core selection and arbitration processes, which are inefficient and difficult to scale without increasing wiring complexity.
Innovation Solution
A processor architecture that includes destination notification dedicated lines, a destination core selection circuit, and a data matching circuit to efficiently select and update register values across cores, allowing for simultaneous notification and reduction of latency through parallel data transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If sequential arbitration and core selection is used for register value notification, then device complexity is reduced, but latency increases and productivity decreases
Solution Approach 1:
The notification process is segmented into two independent parallel paths: (1) destination core selection circuit selects a first core from multiple destination cores, and (2) data matching circuit simultaneously identifies second cores with matching transmission data. This segmentation eliminates sequential arbitration, allowing multiple cores to be notified in parallel, thereby reducing latency without proportionally increasing device complexity.
Solution Approach 2:
The destination core selection circuit and data matching circuit perform their selections and comparisons in advance simultaneously, before the actual data transmission. By pre-identifying all destination cores (both first core selected by destination selection and second cores with matching data) beforehand, the system avoids time-consuming sequential arbitration during the notification phase, thus reducing overall latency.
2Speed
If dedicated notification lines are provided to each core, then speed of notification is improved, but device complexity and wiring complexity increase
Solution Approach 1:
The patent merges the destination selection function and data matching function into a unified notification mechanism. The destination core selection circuit and data matching circuit work together to identify multiple destination cores simultaneously, and the notification is transmitted through shared notification infrastructure rather than requiring separate dedicated lines for each possible destination, thus reducing wiring complexity while maintaining high notification speed.
Solution Approach 2:
The notification circuitry is designed with multi-functionality to handle both the selection of the first core via destination core selection and the identification of second cores via data matching. This universal notification mechanism can serve multiple destination cores simultaneously through a shared notification bus or line, eliminating the need for separate dedicated notification lines for each core pair, thereby reducing wiring complexity while preserving high notification speed.
3Productivity
If parallel data transmission to multiple cores is implemented, then productivity is improved, but device complexity increases
Solution Approach 1:
The parallel notification process is segmented into two independent concurrent operations: destination core selection (identifying the first core) and data matching (identifying second cores with matching transmission data). This segmentation allows the system to notify multiple cores in parallel without requiring a completely complex new architecture, as each segment uses dedicated but relatively simple circuitry that works independently and simultaneously, thus improving productivity with controlled device complexity.
Solution Approach 2:
Both the destination core selection and data matching operations perform their selections and comparisons in advance simultaneously, before the actual data transmission phase. By pre-identifying all destination cores through these preliminary parallel operations, the system enables rapid parallel data transmission to multiple cores, thereby improving arithmetic operation performance without requiring complex real-time arbitration during the critical data transmission path.
Data Source
AI summary
A processor includes a plurality of cores to which individual destination notification dedicated lines are coupled. A destination core selection circuit receives a plurality of packets having some of the cores as destinations, respectively, for arbitration, and selects one first core from among the cores serving as the destinations of the plurality of packets. A data matching circuit compares first transmission data included in a packet for the first core with second transmission data included in a packet for the core other than the first core serving as a destination to participate in the arbitration, extracts one or more second cores of which the second transmission data matches the first transmission data, designates the first core and the second core as the destinations by using the destination notification dedicated lines, and transmits the first transmission data via a data notification line.


