Multi-threaded wireless communication processor with refined thread processes

The multi-threaded superscalar processor with a thread mapping register and dynamic interrupt handling addresses unequal bandwidth allocation and interrupt inefficiencies, achieving reduced processor requirements and improved efficiency in multi-threaded environments.

CN112486572BActive Publication Date: 2025-07-15SILICON LABORATORIES INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202010940096.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-02
Filing Date
2020-09-09
Publication Date
2025-07-15
Estimated Expiration
2040-09-09

AI Technical Summary

Technical Problem

Existing multithreaded processors are uneven when allocating CPU bandwidth, cannot be dynamically adjusted, and cannot handle interrupts in time, resulting in inefficient processing.

Method used

The combination of thread-mapping registers and thread-by-thread program counters is adopted to dynamically allocate CPU cycles and provide separate interrupt processing for each thread to ensure the timeliness of interrupt processing.

Benefits of technology

It realizes dynamic allocation of CPU bandwidth and timely processing of interrupts, improves processor processing efficiency and reduces power consumption requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112486572B_ABST
    Figure CN112486572B_ABST
Patent Text Reader

Abstract

This application relates to a multi-threaded wireless communication processor with refined thread processes. The communication processor is operable to adapt thread allocation on an instruction-by-instruction basis to communication processes handled by the multi-threaded processor. A thread mapping register controls the allocation of each processor cycle to a specific thread and reprograms the thread mapping register according to the network process loads of multiple communication processors such as WLAN, Bluetooth, Zigbee, or LTE with increasing or decreasing load requirements. A thread management process can dynamically allocate processor cycles to each corresponding process during the active time of each associated communication process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-threaded processor. More specifically, the present invention relates to a multi-threaded processor having features of fine-grained and dynamic thread allocation, such that a variable percentage of central processing unit (CPU) processing power can be dynamically allocated to each thread. Background Art

[0002] When a system runs multiple processes, a multi-threaded processor is used, and each process runs with its own separate thread. Examples of prior art multi-threaded processors and their use are described in U.S. Patents Nos. 7,761,688, 7,657,683, and 8,396,063. In a typical application using a two-threaded processor for wireless communication, the processor alternates execution cycles between executing instructions of a high-priority program on a first thread and executing instructions of a low-priority program on a second thread, and the alternation results in allocating 50% of the CPU processing power to each thread. In addition, the allocation of CPU bandwidth to each thread is protected, because during a thread stop, such as when the first thread accesses an external peripheral device and has to wait for data to return, the second thread can continue execution without being affected by the stop of the first thread.

[0003] There are problems where a multi-threaded processor needs to allocate bandwidth unevenly or needs to change the allocation dynamically. It is desirable to provide a dynamic allocation of thread utilization for each task such that during each interval consisting of a set of processor execution cycles, each of the threads during that interval receives a fixed percentage of the CPU cycles. During subsequent intervals, other threads can be added or removed, or the percentage allocation of CPU cycles for each thread can be changed. It is also desirable to provide an unequal allocation of CPU capabilities among multiple threads and to perform the allocation dynamically.

[0004] Another problem in a multi-threaded processor is the timely handling of interrupts. During interrupt handling, new interrupts are disabled so that the handling of a particular previous interrupt can be completed. Subsequentially received interrupts cannot be recognized until the previous interrupt handling is completed and the interrupt is revealed. It is desirable to provide interrupt handling that can timely recognize new interrupts that arrive during the pendency of the handling of a previous interrupted task.

[0005] Object of the Invention

[0006] A first object of the present invention is a multi-threaded superscalar processor having a series of cascaded stages, each cascaded stage providing the result of an operation to a subsequent stage, the first of the cascaded stages receiving instructions from a program memory address referenced by a thread identifier and an associated program counter, the thread identifier being provided by a thread mapping register containing a sequence of thread identifiers, each thread identifier indicating which of the program counter and register file is to be used by a particular processor stage, selecting a particular instruction using the thread identifier and per-thread program counter provided to a sequence of pipelined stages including an instruction fetch stage, an instruction decode stage, a decode / execution stage, an execution stage, a load / store stage, and a write-back stage, the decode / execution stage being coupled to a register file selected by the thread identifier.

[0007] A second object of the present invention is a multi-threaded superscalar processor operable to handle multiple interrupt processes, each interrupt process being associated with a particular thread.

[0008] A third object of the present invention is a multi-threaded superscalar processor having a thread mapping register that is reprogrammable to dynamically identify a sequence of threads to be executed, each thread being associated with a program counter register and a register file, the program counter register and the register file being coupled to at least one subsequent stage: a prefetch stage, an instruction fetch stage, an instruction decode stage, a decode / execution stage, an execution stage, a load-store stage, and an optional write-back stage.

[0009] A fourth object of the present invention is the dynamic allocation of thread bandwidth from a first protocol process to a second protocol process, each protocol process handling data packets arriving via a separate interface and being handled by different threads in a multi-threaded processor, the multi-threaded processor having a refined control over the allocation cycle for each thread.

[0010] A fifth object of the present invention is a communication interface having concurrent handling of unrelated communication protocols such as Bluetooth and WLAN, the Bluetooth interface being active during regular time intervals separated by gaps during which the Bluetooth protocol is inactive and gaps during which the Bluetooth protocol is inactive for WLAN communication, the communication protocol running on the multi-threaded processor providing a dynamic assignment of a greater number of thread cycles to the Bluetooth protocol during active Bluetooth intervals and a dynamic assignment of a greater number of thread cycles to the WLAN protocol during active WLAN intervals. SUMMARY OF THE INVENTION

[0011] In an example of the present invention, a superscalar processor has (in sequence) a prefetch stage, a fetch stage, a decode stage, a decode / execution stage, an execution stage, a load / store stage, and an optional write-back stage. The prefetch stage receives instructions provided by a per-thread program counter under the guidance of a thread mapping register, which provides a canonical sequence of thread identifiers that index into the per-thread program counters to select the identified threads, and the selected program counters direct the prefetch stage to receive instructions from an instruction memory. The decode / execution stage is coupled to a register file that selects the register file associated with the thread being executed by the decode / execution stage at that time, thereby addressing a thread-specific register bank.

[0012] The thread mapping register identifies the specific thread being executed, where the thread mapping register can refer to any number of different threads, subject to the number of per-thread program counters and per-thread register files. For example, the thread mapping register can contain 10 entries, and the number of per-thread program counters and per-thread register files can be 4. In this case, the granularity of each of the 4 threads can be specified as 10%, such that thread_0 can receive 1 cycle, thread_1 can receive 4 cycles, thread_2 can receive 3 cycles, and thread_3 can receive 2 cycles. The thread register (non-limiting) can specify any one of [0, 1, 1, 1, 1, 2, 2, 2, 3, 3] for canonical execution. The thread register can be updated to change the number of threads or thread allocation, such as, by writing the new value [0, 0, 0, 0, 1, 2, 2, 2, 3, 3] to the thread register, thread 0 can be expanded and thread 1 can be reduced.

[0013] In another example of the present invention, per-thread interrupt masking is provided on a superscalar multi-threaded processor such that each thread has its own separate interrupt register. In this example of the present invention, each thread has its own separate interrupt handling such that an interrupt to thread_0 is masked by thread_0, and other threads (such as thread_1, thread_2,..., thread_n) continue execution, each having the ability to separately handle interrupts directed to each corresponding thread. In this example architecture, each thread can be capable of handling different protocol types, for example, a packet buffer coupled to a processor interface of a multi-protocol baseband processor with a general packet buffer interface can be used to handle each of wireless protocol WLAN, Bluetooth, and Zigbee packet handling. In this example, the multi-threaded processor can handle acknowledgment and retransmission requests, each of which must be completed in a timely manner using interrupt handling, each protocol type occurs on a separate interrupt dedicated to a separate thread, and the thread register is rewritten as needed to allocate larger thread cycles on an adaptive basis. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 Block diagram showing a multi-threaded superscalar processor with a per-thread program counter and a per-thread register file.

[0015] Figure 1A Block diagram showing the organization for a per-thread program counter.

[0016] Figure 1B Block diagram showing an example of a thread mapping register.

[0017] Figure 2A Thread mapping register showing an example of a thread mapping register for thread sequential mapping and a given thread assignment.

[0018] Figure 2B Showing for Figure 2A Thread mapping register for non-sequential mapping of a thread.

[0019] Figure 3 Showing for Figure 1 Per-thread interrupt controller and handling for a multi-threaded processor.

[0020] Figure 4 Block diagram showing a Bluetooth and WLAN processor using a separate CPU.

[0021] Figure 5 Block diagram showing a Bluetooth and WLAN processor using a multi-threaded processor.

[0022] Figure 5A Showing an example assignment of program code and related tasks for a multi-threaded processor.

[0023] Figure 5B Showing an example assignment of RAM for a packet buffer. Detailed Description

[0024] Figure 1An example of the present invention showing a superscalar processor 100 having a sequence of levels: a prefetch level 102, a fetch level 104, a decode level 106, a decode / execute level 108, an execute level 110, a load / store level 112, and an optional write-back level 114. Instructions delivered to the prefetch level 102 are sequentially executed by each subsequent level on separate clock cycles, carrying forward any context and intermediate results required for the next level. In one example of the present invention, a thread mapping register 103 provides a canonical sequence of thread identifiers (thread_id) to be delivered to a per-thread program counter 105, which provides the associated current program counter 105 address to the prefetch level 102. The prefetch level 102 retrieves the associated instructions from an instruction memory 116 and delivers them to the fetch level 104 in a subsequent clock cycle. The decode / execute level 108 is coupled to a per-thread register file 118, which responds to read requests from the decode / execute level 108 or write-back operations from level 114, each of which is dedicated to a thread, so that data read or written to the register file 118 corresponds to the thread_id for which the data is being requested or provided.

[0025] Figure 1A Showing a plurality of per-thread program counters 105: PC_T0 for thread_0, PC_T1 for thread_1, ..., PC_Tn for thread n, with one program counter operable for use with each thread.

[0026] Figure 1B Showing the thread mapping register 103, which includes a sequence of thread identifiers T0 130 to Tn132 for canonical execution. The number of threads (each thread being a separate process executed in a CPU cycle at a particular level) is m, limited by the number of register files 118 and program counters 105, while the thread mapping register 103 can support m threads to evenly distribute CPU bandwidth among the threads, or for coarser-grained thread control, can provide n time slots, where n > m. For example, a thread mapping with 16 entries can support 4 threads, each with a granularity of 1 / 16 of the available CPU processing power, and support any value between 0 / 16 and 16 / 16 of the available CPU processing power (depending on the allocation of CPU processing power to the remaining threads).

[0027] Figure 2A Showing an example 16-entry thread mapping register 103 on a canonical cycle length 204, with the thread mapping register canonically repeating at the end of each 16 entries. Figure 2AThis example is shown for 4 threads and a sequential mapping, which can be applicable to applications in the case of no thread stops, for example. In this case, due to the delay in receiving results from external resources, the threads cannot execute sequential cycles. For an n = 16 thread mapping register location, the thread mapping register provides a 1 / 16 resolution of the processor application for each task, and the processor can be used with one thread at each thread mapping register location. However, this provides a fixed time allocation for each thread. In a preferred utilization, the number m of thread identifiers is less than the number n of thread mapping register locations, which provides that the allocation of a particular thread to a task can have a granularity of p / n, where n is usually fixed, while p is programmable to the number of cycles allocated to a particular thread and can vary from 0 to n to allocate more or less computing resources to each thread. In another example of the present invention, the length n of the thread mapping register can itself be programmable to provide greater granularity in task cycle management or to support a larger number of threads.

[0028] Figure 2A An example thread mapping register for a four-thread processor in a 16-location thread mapping register 202 is shown, where threads 0, 1, 2, and 3 (T0, T1, T2, T3 respectively) and 12.5%, 25%, 50%, and 12.5% of the processor capabilities are allocated to each corresponding thread. There is a problem where a particular thread has to wait for an external resource response (referred to as thread stop). In Figure 2A the example, the decode / execute stage 108 may need to read an external shared memory or a media access controller (MAC) not shown, and the delay in reading the external resource may require 4 clock cycles. In Figure 2A In the case where the threads that show thread allocation and access the external resource are T0 and T3, or in the case of a delay when reading or writing to a device, T0 will be in thread stop in operation 208, and T3 will be in thread stop 214 in cycle 210. In Figure 2A In the case of the arrangement of the thread identifiers shown in

[0029] Figure 2B An alternative mapping is shown, which uses the same time allocation as Figure 2A but rearranges the thread sequence 220 for the same thread stop situation as Figure 2A shown. The rearrangement of T0 to positions 0 and 7 and T3 to positions 1 and 8 is reflected in the arrangement of Figure 2B . The T0 thread is only stopped when the thread stop is longer than 6 clock cycles 224, while the thread stop 212 is 4 clock cycles, so in Figure 2BThe arrangement performs two occurrences of T0 instead of one of those in Figure 2A Similarly, unless the thread stop has a duration of 226, there is no occurrence in Figure 2B of the T3 stop that causes the second T3 cycle delay of Figure 2A .

[0030] Figure 3 FIG. shows another aspect of the present invention, an example for wireless signal processing, wherein the process thread 308 can be executed as different threads on a multi-threaded processor of Figure 1 , and the multi-threaded processor has an interface 310 that is part of the multi-threaded CPU 100, and each interface is associated with a specific MAC. Wireless signals are received and transmitted on the antenna 301, converted to baseband when received, or modulated to RF when transmitted at 302, and provided to the multi-protocol baseband processor 304. When a data packet arrives at a specific interface of the multi-protocol MAC, an interrupt for a specific thread can be sent to the interrupt controller 306, where each interrupt can be masked by an associated process 308 operating in the multi-protocol processor. Each process is capable of controlling the associated interrupt mask (shown as IM0, IM1, IM2, IM3), which is provided to the interrupt controller 306 to mask the interrupts such that the associated process does not process new interrupts until the previous interrupt for that process has been completed.

[0031] The current multi-task handling of interrupts has specific advantages over the prior art. In the prior art, the interrupt service routine on thread 0 can be handling packet acknowledgments for multiple packet interfaces. In this task, after receiving a packet, the receive buffer is checked to detect any lost packets in the sequence, and the process acknowledges the received packet or makes a retransmission request to the sender for any lost packets. There are critical timing windows associated with packet acknowledgment and retransmission, so it is important to make an acknowledgment or retransmission request in a timely manner after receiving a packet. We can consider the following scenario: where a retransmission request must be made within 30 us after receiving a packet, and the first retransmission task 0 takes 5 us to complete, the second retransmission task 1 takes 10 us to process and complete, and the third retransmission task 3 takes 5 us to process and complete, and a single process handles three tasks on a single thread. In this example where the three tasks are handled by a common thread and a common interrupt mask is used as in the prior art, after receiving a packet, the process on thread 0 handles task 0 masking the interrupt to prevent other packet acknowledgments from slowing down the current acknowledgment that takes 5 us. If a second interrupt associated with task 1 of thread 0 arrives during the handling of task 0, task 1 is not handled until at least 5 us after it arrives because thread 0 is still busy with task 0. The following situation may further occur due to burst packets on different interfaces. While task 1 (which takes 10 us) is waiting for task 0 (which takes 5 us) to complete, the third task 3 that takes 5 us to complete may arrive. When task 0 is completed, the interrupt mask is removed, task 1 generates an interrupt and is detected, the interrupt mask is asserted again, and the handling of task 1 is completed. Thereafter, the interrupt mask is cleared, and task 2 is detected by asserting its interrupt. Thereafter, the interrupt mask is asserted again, task 2 starts at least 15 us after the request arrives, and after passing the required retransmission request window, the request is completed in 20 us. After task 2 is completed, the interrupt mask is cleared, however the remote station does not receive the retransmission request from task 2 in a timely manner, and the retransmission protocol fails. The prior art solution to the waiting time delay problem of task 2 after handling earlier tasks 1 and 2 is a faster processor. Additionally, when a multi-core processor is reading the MAC interface, thread locking can occur, which can be avoided by the rearrangement of thread identifiers as previously shown in Figure 2B. In this case, a small number of thread cycles may be required for acknowledgment and retransmission tasks, but spreading these three tasks across separate threads, each thread allocated a small amount of time, will overcome the interface read / write waiting time as well as the waiting time delay of the interrupt mask by associating each thread with a separate interrupt and interrupt mask.

[0032] In the prior art, where each of the tasks is executed on a single thread and each task requires 50 MIPS, due to the latency and delay of sequential interrupt handling, 300 MIPS of processing power is required to successfully handle three tasks. Whereas with Figure 3 the novel method, only approximately 150 MIPS is required, thus saving the MIPS requirement to half of the original, thereby reducing the power consumption requirement.

[0033] In Figure 1 another example of a multi-protocol processor, each of the wireless protocols can be handled by a separate thread. For example, the processes for handling WLAN, Bluetooth, and Zigbee can each operate on separate processes on their own threads, and the retransmission processes for each can be handled by separate processes for each protocol, each protocol operating on its own thread.

[0034] In another example of the present invention, the thread mapping register can be interactively changed according to the process requirements detected by the process managed by a separate thread. Since the context from each stage is forwarded to Figure 1 the subsequent stage, the change to the thread mapping register can be completed at any time, depending on the synchronization clock requirements of the prefetch stage 102 for receiving the deterministic thread_ID and the associated per-thread program counter 105.

[0035] Figure 4 Shows an example wireless local area network (WLAN) and Bluetooth (BT) combined transceiver, which has an interface 480 for exchanging data with a communication system. Due to the dedicated WLAN and BT processing operations required for each protocol and the required response timeliness for each, each interface type requires a CPU. The requirement of the CPU for low-latency processing of each interface results in the WLAN and BT processing performed by the system architecture shown by Figure 4 the figure.

[0036] Figure 4 Shows a WLAN processor including an analog front end and a MAC 401 coupled to a WLAN CPU 424, and a BT process including an analog front end and a MAC 450 coupled to a BT CPU 482. Each of the WLAN CPU 424 and the BT CPU 482 is capable of responding in a timely manner to interrupts and emergencies that require immediate processing by the software programs associated with each respective WLAN processor 401 and BT processor 450.

[0037] In the WLAN processor 401, the antenna 402 is coupled to the transmit / receive switch 404 for coupling the received signal to the low noise amplifier 406 and the transmit signal from the power amplifier 414. The input signal is mixed 408 to the baseband using the clock source 418 and low pass filtering 410, and the analog baseband signal is digitized and processed by the combined ADC and baseband processor 412, which demodulates the received symbols into a data stream that is formed into layer 2 packets by the media access controller (MAC) 422 across a serial data interface (SDI) such as to the CPU 424. The CPU 424 has an associated random access memory (RAM) 428 for storing received and to-be-transmitted packets, program code executed by the CPU 424, and other non-persistent information of the system during power-down. The read-only memory (ROM) or flash memory 426 is used to store program instructions that are typically downloaded from the flash / ROM to the RAM during the power-up sequence. The MAC 422 receives data transmitted on the interface 423 (such as a serial data interface (SDI)) and provides the received packets along with sequence numbers to the CPU 424 such that the CPU 424 can detect and manage retransmission of any lost data, as well as set any WLAN authentication protocols and perform any required per-packet operations such as encapsulation and decapsulation, channel management, packet aggregation, and connection management and authentication.

[0038] Figure 4 An example Bluetooth processor 450 including an analog front end and a BT MAC is shown, which similarly operates with an antenna 452, a transmit / receive switch 454, a low noise amplifier 456, a mixer 458, a bandpass filter 460, and an analog / digital converter and baseband processor 462 as the ADC / baseband processor 412 does for WLAN 802.11 packets, and the analog / digital converter and baseband processor 462 can convert the baseband Bluetooth frequency hopping pattern into a data stream. The Bluetooth transmit chain includes a baseband processor and DAC 470, a mixer 466 that modulates the baseband frequency hopping stream to the RF carrier frequency using a modulation clock source 468, and a power amplifier 464 that couples the modulated Bluetooth frequency hopping stream to the transmit / receive switch 454. The BT CPU disposes of various connection management including pairing.

[0039] The WLAN MAC 422 is coupled to the WLAN CPU 424 via a digital interface 423 (such as a serial peripheral interface (SPI)), and the BT MAC 480 is coupled to the BT CPU 482 via a digital interface 481. Figure 4 The architecture thus provides separate CPU processing power for each of the WLAN and Bluetooth processes operating independently, including low latency for processing connection or packet requests from each interface.

[0040] Figure 5 shows Figure 4 an alternative architecture in which the WLAN RF front end / MAC 504 (corresponding to Figure 4 processing 401) and the BTRF front end / MAC 508 (corresponding to Figure 4 processing 450) are coupled to the multi-threaded CPU 510 via respective digital interfaces 518 and 520, and the multi-threaded CPU 510 is itself coupled to the ROM / flash 512 and the RAM 514. Optionally, there is a thread mapping register 516 which provides the allocation of CPU cycles to Bluetooth or WLAN processes. In one example of the present invention, the number of process_id entries in the thread mapping register is fixed, and the number of increasing or decreasing thread_id values can appear in the thread mapping register to provide an increasing or decreasing number of process cycles to a specific process associated with each thread_id. For a pipelined multi-threaded processor that receives one instruction at a time as described above, the multi-threaded processor that receives each instruction of the thread determined by the thread mapping register (which issues the next thread_id for each instruction), the granularity of the control of the thread process allocation is instruction-by-instruction. Since the thread mapping register issues the thread_id in a repetitive and standardized manner, the allocation of process to thread has a very fine granularity, which is equal to the reciprocal of the number of values that the thread mapping register can support. In one example of the present invention, the thread management process can operate as one of the processes with a unique thread_id, and the thread management process checks the activities in other threads according to the activity level to increase or decrease the number of entries of the corresponding thread_id, and designates and deallocates the thread_id value from the thread mapping register. The activity level of the communication process associated with the communication processor can be determined, for example, by the number of data packets sent or received by the associated communication processor and processed by the thread, and a threshold can be established to indicate when there are more or fewer thread_id values in the thread mapping register for the specific thread. Examples of process threads with a unique thread_id (the thread_id having more or fewer entries dynamically placed in the thread mapping register by the thread management process) include link layer processes, network layer processes or application layer processes, where each link layer, network layer or application layer process can include multiple processes with unique threshold metrics, and each of these processes is associated with a specific communication processor (such as 401, 450, 504 or 508). The increased thread_id allocation in the thread mapping register can be made for a period of time when the threshold metric (such as, packet data rate, the number of packets remaining to be processed, thread load metric or percentage of thread process task completion) exceeds the threshold.

[0041] Figure 5AShows the allocation of memory (ROM / Flash 512 or RAM 514) to the various threads present. One thread can be the WLAN code corresponding to the task executed by the WLAN CPU 424 in Figure 4 and another thread can be the BT code corresponding to the task executed by the BT CPU 482 in Figure 4 . Additional threads can be specified to manage the thread mapping register, thus controlling the allocation of the bandwidth of various tasks to the previously described thread mapping register 103, and additional tasks can perform memory management of the packet buffer and other rarely executed functions of low priority. The thread mapping management task can periodically check the utilization of the BT and WLAN interfaces and change the CPU cycle allocation for each task as needed. In one aspect of the present invention, Bluetooth and WLAN operations are executed exclusively, and the CPU thread allocation for the interfaces ( Figure 5A for the BT and WLAN tasks) is dedicated to one interface or the other.

[0042] In another example of the present invention, the various threads can handle different parts of a particular communication protocol. For example, one thread can handle layer 2 and other operations, and another thread can handle layer 3 and the application aspects of a particular protocol. In one aspect of the present invention for any WLAN protocol, one thread can handle the basic communication aspects, which can be collectively referred to as the lower layer MAC functions. The lower layer MAC functions of WLAN and Bluetooth include packet transmission, packet reception, clear channel assessment (CCA), inter-frame spacing, rate control, request to send and clear to send (RTS / CTS) exchange, wireless packet acknowledgment DATA / ACK for WLAN and Bluetooth, or channel hopping specific to Bluetooth. The upper layer MAC functions perform other ISO (International Organization for Standardization) layer 2 functions at the data link layer that are not performed by the lower layer MAC functions. The upper layer MAC functions in this specification collectively refer to any of the following: WLAN supplicant (any protocol associated with joining or logging in to a wireless network access point), WLAN packet retransmission and acknowledgment, security functions such as those described in the standard WPA or WPA2 (Wi-Fi Protected Access). The ISO layer 3 (network layer) functions can be executed by a separate thread. The layer 3 functions include IP packet formation, TCP retransmission and acknowledgment, SSL encryption and connection management, and application layer operations such as packet encapsulation for a particular application layer process. In another example of the present invention for Bluetooth, one thread can be specified to handle the Bluetooth controller, stack, retry, and acknowledgment, and another thread can be specified to handle the application layer tasks. In this way, the two tasks for a particular protocol are separated and provided to separate threads, and a common interface (such as SRAM) can be used to transfer data from one thread to another.

[0043] In some applications, WLAN communication and Bluetooth communication can coexist and run concurrently. In this example configuration, CPU thread cycles can be dynamically allocated to the WLAN communication process when processing WLAN data packets, and to the BT thread cycle when processing Bluetooth data packets. Multiple processes associated with specific communication processors 401, 450, 504, or 508 can be created using unique thread_id values, each thread_id is placed in the thread mapping register 516 to provide processing bandwidth for each associated process, and these processes exit when the associated communication processor is not enabled, and the thread_id is removed from the thread mapping register 516. Concurrent communication can be performed by relying on the regular communication intervals of Bluetooth communication in which data packets are sent at regular time slot intervals, and the concurrent communication can be separated in time by larger intervals during which the channel is not used for BT communication. During these intervals, WLAN data packets can be sent and acknowledged without interfering with the BT communication window. The thread mapping register 103 can be dynamically changed to provide a larger percentage of CPU capacity for BT during Bluetooth data packet intervals, and then a larger percentage of CPU capacity for WLAN during WLAN data packet intervals, thereby reducing Figure 4 power consumption on the architecture.

[0044] Figure 4 and Figure 5 The examples shown are for specific different communication protocols for WLAN and Bluetooth, but it should be understood that these are for illustrative purposes only. Different communication protocols are sets of communication protocols that require completely different packet handling. Examples are any of Bluetooth, WLAN, Zigbee, Near Field Communication, and other examples are known to those skilled in the art of communication protocols.

Claims

1. A communication processor, comprising: A plurality of communication controllers operable to participate in wireless communication, each communication controller having a data interface for transmitting and receiving wireless data packets; A multi-threaded processor coupled to the plurality of communication controllers, the multi-threaded processor operable to execute at least one process thread with a unique thread_id for each communication controller; The multi-threaded processor has a thread mapping register that contains thread_id values and generates the thread_id values in a canonical sequence, the number of thread_id values in the canonical sequence being greater than the number of unique thread_id values; The multi-threaded processor executes each instruction of each thread according to the corresponding thread_id value generated by the thread mapping register; A thread management controller operable to modify the thread mapping register by: Determining a first activity metric of a process operable on a first communication controller and a second activity metric of a process operable on a second communication controller; Determining the greater activity metric of the first activity metric or the second activity metric; Adding an additional thread_id entry associated with the greater activity metric to the thread mapping register.

2. The communication processor according to claim 1, wherein at least one communication controller is operable to transmit and receive wireless local area network packets, i.e., WLAN packets, and the communication controller has a media access controller for transmitting and receiving data using the data interface.

3. The communication processor according to claim 1, wherein at least one communication controller is operable to transmit and receive Bluetooth packets.

4. The communication processor according to claim 1, wherein at least one thread has an associated thread_id value that is not in consecutive thread mapping register positions.

5. The communication processor according to claim 1, wherein the multi-threaded processor is coupled to the plurality of communication controllers through a single data interface.

6. The communication processor according to claim 5, wherein the single data interface is a serial data interface.

7. The communication processor according to claim 1, wherein at least one communication controller is operable to receive or transmit wireless local area network packets, i.e., WLAN packets.

8. The communication processor according to claim 1, wherein at least one communication controller is operable to receive or transmit Bluetooth packets.

9. The communication processor according to claim 1, wherein the thread management controller monitors at least one network parameter associated with the utilization of the Bluetooth interface or the utilization of the WLAN interface.

10. A thread management process operable on a plurality of process threads in a communication processor, the communication processor having: A plurality of communication controllers operable to participate in wireless communication, each communication controller having a data interface for transmitting and receiving wireless data packets; A multi-threaded processor coupled to the plurality of communication controllers, the multi-threaded processor operable to execute at least one process thread with a unique thread_id for each communication controller; The multi-threaded processor has a thread mapping register that contains thread_id values and generates the thread_id values in a canonical sequence, where the number of thread_id values in the canonical sequence is greater than the number of unique thread_id values; The multi-threaded processor executes each instruction of each thread according to the corresponding thread_id value generated by the thread mapping register; The thread management process includes: Determining a first activity metric of a process operable on a first communication controller and a second activity metric of a process operable on a second communication controller; Determining the greater activity metric of the first activity metric or the second activity metric; Adding an additional thread_id entry associated with the greater activity metric to the thread mapping register.

11. The thread management process according to claim 10, wherein at least one communication processor has a media access controller and an associated link layer communication process that performs transmission and reception of data packets, and the at least one communication processor also has an associated network layer process that performs network layer retransmission and acknowledgment of network layer data packets.

12. The thread management process according to claim 11, wherein the link layer communication process is executed by a first process thread having a unique thread_id, and the network layer process is executed by a different process having a unique thread_id different from the thread_id of the first process thread.

13. The thread management process according to claim 12, wherein when the number of data packets transmitted or received by the associated at least one communication processor increases beyond a threshold, the number of first process thread_id values in the thread mapping register increases, and the number of second process thread_ids in the thread mapping register increases.

14. The thread management process according to claim 10, wherein the communication processor has a security function executed by a security thread process having an associated unique thread_id, and during the security function, the number of security function unique thread_id entries in the thread mapping register increases.

15. The thread management process according to claim 14, wherein the security function is authentication.

16. The thread management process according to claim 10, wherein at least one communication processor includes a media access controller, i.e., MAC, and one process thread associated with the MAC disposes of at least one lower layer MAC function, and a different process thread associated with the MAC disposes of at least one upper layer MAC function.

17. The thread management process according to claim 16, wherein information sharing between the process thread disposing of the at least one lower layer MAC function and the process thread disposing of the at least one upper layer MAC function shares data in a random access memory associated with the multi-threaded processor.

18. The thread management process according to claim 10, wherein one of the multiple process threads is a thread management process that creates a new thread process and places an associated thread_id in the thread mapping register.

19. The thread management process according to claim 10, wherein a new thread process is created when a communication processor is enabled for a wireless communication protocol.

20. The thread management process according to claim 10, wherein at least one of the threads executes a memory management process, which includes allocating and deallocating packet buffer memory in a random access memory (RAM) accessible by the multi-threaded processor.

21. The thread management process according to claim 10, wherein at least one of the plurality of process threads is an interrupt handler process.

Citation Information

Patent Citations

  • Cross-thread interrupt controller for a multi-thread processor

    US7657683B2

  • Multiple thread in-order issue in-order completion DSP and micro-controller

    US7761688B1

  • Multi-thread media access controller with per-thread MAC address

    US8396063B1

  • System and method of controlling multiple program threads within a multithreaded processor

    CN101258465A

  • Method and apparatus for resource-based thread allocation in a multiprocessor computer system

    CN1955920A