Computer-implemented method, computation system, circuit, and machine-readable medium for orthogonalization-driven and architecturally comprehensive signal processing

TWI935632BActive Publication Date: 2026-08-11PEAK TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
TW114101527
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-09-12
Filing Date
2025-01-14
Publication Date
2026-08-11
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Conventional signal processing techniques are incompatible with modern trends of maximizing processing parallelism and reducing latency, particularly in computationally intensive tasks, and fail to efficiently process all possible permutations of input data vectors.

Method used

A novel orthogonalization-driven approach using a Symmetric Orthogonalization Cell (SOC) building block and N-input, 2N-output architecture, which includes a multi-cylinder parallel processing architecture to identify and locate mutually orthogonal q-vectors, reducing redundant hardware and enhancing operational symmetry.

Benefits of technology

This approach significantly enhances processing parallelism and reduces latency in signal processing applications, enabling efficient handling of complex scenarios like adaptive pulse Doppler processing and neuromagnetic signal enhancement, while minimizing physical space and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001905571_001
    Figure TWG2TB001905571_001
  • Figure TWG2TB001905571_002
    Figure TWG2TB001905571_002
  • Figure TWG2TB001905571_003
    Figure TWG2TB001905571_003
Patent Text Reader

Abstract

This paper discloses a novel orthogonalized driving and architecture-wide signal processing technique, encompassing modular architectures in computer implementation methods, computing systems, circuits, and machine-readable media. This significantly enhances processing parallelism and reduces processing latency for various computationally intensive signal processing tasks. This technique also plays a crucial role in the design of high-performance integrated circuit chips for these tasks. Developing this overall technique requires a unique N-input, 2N-output processing architecture as a baseline.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to the field of signal processing. It presents a novel orthogonalized driving and architecture-wide approach that can significantly enhance processing parallelism and reduce processing latency for a wide range of computationally intensive signal processing tasks. This approach also plays a crucial role in the design of high-performance integrated circuit chips for these tasks. [Previous Technology]

[0002] Machine learning, signal processing, wireless telecommunications, and semiconductor integrated circuit design techniques (such as orthogonalized networks and systolic array implementations) have been developed. Many technological advancements are essentially the result of applying fundamental mathematical principles at the level of basic solutions. A particularly crucial mathematical concept involves transforming a set of N linearly independent input vectors into a set of N mutually orthogonal output vectors. For over a century, this fundamental concept, along with other mathematical tools built upon it, has been used to solve a wide range of scientific and engineering problems. On the other hand, these mathematical principles and tools were already mature and widespread even decades before the advent of the computer age. Of course, the inherent characteristics of these mostly century-old concepts and algorithms may be incompatible with the latest technological trends of maximizing processing parallelism and reducing processing latency. Given this background, the present invention discloses a novel orthogonalization-driven approach for developing parallel processing architectures that are more efficient than known techniques and methods for a wide range of signal and data processing problems. [Summary of the Invention]

[0003] One object of this disclosure is to provide a method and system for processing input signal vectors based on an orthogonalization-driven and architecturally comprehensive approach. Using a two-input, two-output building block (i.e., a SOC) called a Symmetric Orthogonalization Cell (SOC), a new generation of parallel processing architecture solutions can be developed for a wide range of signal processing scenarios and applications. The entire innovation process is described through modular processing architecture diagrams and related data flow illustrations. This invention discloses a clear description of the processing architecture and technical level related to prior art. It also provides background information on relevant signal processing application areas. Furthermore, the specific breakthroughs and innovations of this invention can be summarized as follows.

[0004] (1) Using a SOC building unit with two inputs and two outputs and the resulting N-input and two-output baseline triangle architecture, the advantages of one-dimensional processing linear array architecture in certain common signal processing application scenarios involving overlapping adjacent input data windows are achieved by adopting an architecture compression process.

[0005] (2) Based on the baseline triangular architecture with N inputs and two outputs, an architecture with N inputs and 2N outputs is invented, which can be visualized as a hollow three-dimensional (3-D) cylindrical structure. In addition to the fact that conventional technology does not mention this concept of N inputs and 2N outputs, a technique is developed to concisely identify and locate multiple sets of mutually orthogonal q-vectors within the architecture. From a practical engineering perspective, the q-vector identification and tracking system must be applicable to any QR decomposition-driven signal processing system;

[0006] (3) In addition to the novel architecture with N inputs and 2N outputs, another key innovation disclosed in this invention is the establishment of a multi-cylinder parallel processing architecture that can utilize all the operational symmetry properties associated with all possible permutations of a given set of input data vectors. This, in turn, provides an analysis and architectural framework to handle the entire next-generation signal processing scenario in the most architecturally efficient way;

[0007] (4) The above-described cylindrical and multi-cylinder architectures are used for proper geometric articulation and effective visualization. In mapping these architectural ideas and descriptions to actual hardware or integrated circuit chip designs, a suitably defined idea is to roll up a "processor sheet" containing a two-dimensional planar or one-dimensional linear array processor architecture, thereby significantly reducing the unused physical space associated with the hollow cylindrical architecture also described in the architecture of this invention. This process of minimizing unused or hollow physical space will result in compact, latency-reduced, and energy-efficient chip designs or hardware implementations.

[0008] Therefore, the claims disclosed in this invention are based on the innovations and breakthroughs summarized above.

Implementation Method

[0010] To make the above and other objects, features and advantages of the present invention clearer and more understandable, preferred embodiments of the present invention will be described below in conjunction with the accompanying drawings. In order to establish clear continuity and objectively describe the original innovation, the following detailed description involves four distinct but interrelated parts: 1. Original innovation relative to prior art; 2. Combination of key mathematical principles and tools; 3. Practical application fields; 4. Detailed description of parallel-driven innovation.

[0011] I. Original innovation relative to conventional technology:

[0012] In the related art, an orthogonalization algorithm with N inputs and N outputs and its systolic array implementation are provided. This N-input and N-output processing architecture is carried out using a building block of two inputs and two outputs. In contrast, another N-input and N-output processing architecture for the Modified Gram-Schmidt (MGS) algorithm is constructed using a building block of two inputs and one output. In this document, N is the number of inputs or outputs, for example, N is a positive integer. The main competitive advantage of the N-input and N-output architecture, which is carried out using a building block of two inputs and two outputs, compared to the MGS-based architecture, is that (1) no data broadcasting line is required, and (2) any given processing node in the processing architecture only needs to communicate with its neighboring nodes. To the best of our knowledge, this related art clearly represents the prior art disclosed in this invention, which discloses a novel architectural and computational breakthrough relative to the prior art.

[0013] Specifically, conventional techniques fail to address the ultimate computational challenge of efficiently processing all possible permutations associated with a given set of input data vectors. The computational challenges in signal processing are particularly relevant to the development of 6G technology and Multiple-Input Multiple-Output (MIMO) systems in wireless communications. Therefore, using a scalable and modular processing architecture to process some or all of these possible permutations in parallel represents a significant opportunity for new technological development. To address the computational challenges associated with this "permutation space," this invention discloses that it is necessary to begin with a processing architecture of N inputs and 2N outputs as a basic solution (rather than the N-input and N-output architecture of conventional techniques). In other words, the N-input and N-output architecture of conventional techniques cannot solve or process every possible permutation of a given set of input vectors.

[0014] The connection between the contents of this invention and key mathematical principles of linear algebra, particularly regarding singular value decomposition (SVD) and QR decomposition (QRD), is described in the next section. However, a brief explanation of the direct coupling with QRD is necessary here because there is a particularly critical problem that has not been fundamentally solved by conventional techniques. It relates to the necessity of identifying and locating several mutually orthogonal output q vectors in the processing architecture. These q vectors are actually several column vectors of the Q matrix in QR decomposition. Accurately identifying the positions of these mutually orthogonal output q vectors is a critical task in any application scenario, and an innovative q vector identification and tracking technique is described in detail in Part IV.

[0015] In summary, the original innovations disclosed in this invention at the macro level relative to conventional technology include: (1) a solution to the idea of ​​using a processing architecture with N inputs and 2N outputs as the basis for various increasingly complex parallel processing, rather than the N inputs and N outputs architecture specifically used in conventional technology; (2) a general solution for developing a multi-tube architecture to simultaneously process all possible permutations associated with a given set of input data vectors; and (3) a novel "zigzagging" technique that can accurately identify many positions of many mutually orthogonal q vectors in a given N inputs and 2N outputs architecture.

[0016] In all the figures and diagrams accompanying this disclosure, the boundary between prior art and original innovations related to the present invention is defined as follows: Figures 1 to 4 illustrate key ideas and foundations established by prior art. On the other hand, original contributions and innovations related to the present invention that transcend prior art are expressed through the illustrations in Figures 5 to 14.

[0017] II. Combining relevant mathematical principles and tools:

[0018] In fields such as signal processing, machine learning, and wireless telecommunications, two particularly useful and popular mathematical tools for research and development are Singular Value Decomposition (SVD) and QR Decomposition (QRD). To establish clarity and continuity, the basic mathematical definitions of SVD and QRD are briefly provided below, along with their general development history and interrelationships in relation to the innovations disclosed in this invention.

[0019] Singular Value Decomposition (SVD) is considered a highlight of linear algebra because it is a powerful tool for solving a wide range of purely mathematical tasks as well as specific application problems. One of the key reasons why SVD is so useful is that it provides a way to decompose any data matrix into three simpler, more easily interpreted matrices: A = UΣV*, where U is an m×m identity matrix, Σ is an m×n diagonal matrix, V is an n×n identity matrix, and * is a superscript denoteing the conjugate transpose operation. The immense importance of SVD and its impact on applied mathematics and engineering are widely recognized. Articles reflecting this widespread recognition can be found in papers such as "On the Early History of the Singular Value Decomposition," which recounts the contributions of five different mathematicians living from the early 19th to the mid-20th century to the existence of SVD.

[0020] Although SVD is very important, it did not gain widespread attention in the scientific and engineering communities until the 1960s. In the past few decades, SVD has been particularly popular in signal processing and machine learning because it explicitly supports tasks such as dimensionality reduction, feature extraction, data compression, noise reduction, and pattern recognition. In research related to wireless telecommunications and future "sixth generation (6G)" capabilities, SVD also plays a crucial role in multiple-input multiple-output (MIMO) systems to improve communication channel capacity, signal quality, and signal detection.

[0021] On the other hand, QRD decomposes a matrix A into a product A = QR of an orthogonal matrix Q and an upper triangular matrix R. In general, QRD can be used to solve the same different problems in mathematics and engineering. However, unlike SVD, a single QRD does not explicitly compute numerous eigenvectors and eigenvalues.

[0022] It is worth noting that SVD always exists in any type of rectangular or square matrix, while eigenvalue decomposition can only exist in square matrices, and sometimes even in square matrices, eigenvalue decomposition does not exist. Therefore, for practical applications, it is generally recommended to use SVD to obtain eigenvalues ​​and eigenvectors.

[0023] It is worth noting that John Francis's "QR algorithm," developed in the late 1950s, bridged the gap between QRD and the operations on eigenvectors and eigenvalues. This "QR algorithm" is an iterative method that repeatedly applies QRD to transform a given matrix until it converges to the desired eigenvalues. Essentially, this "QR algorithm" revolutionized the operations on eigenvalues ​​and eigenvectors, representing a significant milestone in numerical linear algebra. Historically, SVD and QRD were independently developed linear algebra tools. Nevertheless, it is important to emphasize that the "QR algorithm" is a crucial bridge between the two.

[0024] QRD has been widely used, and various specific algorithms have been developed to generate the "Q" and "R" matrices. The so-called "classical Gram-Schmidt (CGS)" algorithm is generally considered the first method for performing QRD. It is well known that CGS suffers from numerical instability and is generally unsuitable for large matrices. The Householder Transformation, proposed by Alston Householder, improves the CGS algorithm by providing a more stable and efficient method. Givens' Rotation, developed by Wallace Givens, is another method that provides numerical stability and is particularly useful for sparse matrices. Finally, in an important experimental paper, John R. Rice discusses the theoretical and error analysis related to the Modified Gram-Schmidt (MGS) algorithm.

[0025] It is worth noting that both SVD and QRD have their "simplified" forms. The sigma (Σ) matrix in SVD is a square matrix and a diagonal matrix, while the Q matrix in QRD is not a square matrix. These "simplified" versions are used to more effectively solve or adapt to many practical application scenarios. Another common situation in many application scenarios is that the Q matrix may not need to be "orthonormal," that is, the "q" row vectors of Q are simply mutually orthogonal vectors and do not need to be orthonormalized. This is an assumption made for all the illustrative figures throughout this disclosure.

[0026] The above is a concise description of the coupling between SVD and QRD, and a brief history of the most recognized QRD algorithm. MGS, Householder Transformation, and Givens' Rotation remain active research areas, including studies on various tweaks to them, their unique advantages for certain applications and specific processing scenarios, and their mathematical interrelationships. After decades of analysis and comparative studies on computational parallelism, ease of hardware implementation, etc., the MGS algorithm has likely become the most popular method and is widely used in many fields and applications. This is evidenced by the various patents filed in the past two decades or so related to the hardware implementation of MGS.

[0027] To be precise, the original basic concept of Gram-Schmidt was formulated as early as the late 19th century, while the modified Gram-Schmidt algorithm was officially released in 1966. Mathematically, since its release, the MGS algorithm has essentially become a standard method for operating on a set of N mutually orthogonal output vectors based on a set of N linearly independent input vectors. It needs to be reiterated here that operating on a set of mutually orthogonal output q vectors and establishing a "foundational solution" are essentially the same task. In this disclosure, the key point about the equivalence between establishing a "foundational solution" and operating on a set of mutually orthogonal output q vectors is so crucial that it cannot be overstated. Understanding the mechanism for obtaining "Q" is also helpful; "Q" is the focus of QRD, and the "R" matrix is ​​a natural byproduct, at least in the context of the MGS algorithm. Given the above history and reasons, a particularly relevant key focus is to develop a novel technique for accurately identifying and locating these mutually orthogonal q vectors in a novel architecture with N inputs and 2N outputs, as disclosed in this invention. To reiterate, the prior art does not mention techniques for identifying these numerous mutually orthogonal positions on a novel architecture with N inputs and 2N outputs.

[0028] At a macro level, related to the mathematical development milestones briefly introduced above, it is sufficient to illustrate that computing a set of mutually orthogonal output vectors q based on a set of N linearly independent input vectors represents a key and necessary computational task for any scientific or engineering problem. In view of this fact and the ongoing trend, this invention discloses a novel method for reorganizing this computational task by applying and utilizing novel parallel processing architecture ideas, in order to significantly improve processing parallelism and reduce processing latency in many advanced signal processing applications. The next section will briefly describe some of these signal processing scenarios and applications in the context of parallel processing. The following section focuses on the connection between complex signal processing scenarios and the processing architecture concept of N inputs and 2N outputs, which is the central architectural theme disclosed in this invention.

[0029] III. Practical Application Areas:

[0030] Given the multidisciplinary nature and diversity of potential applications related to this invention, it is helpful to briefly explain the basic idea of ​​N inputs and N outputs, and the background of the novel N inputs and 2N outputs architectural arrangement mentioned earlier. Transforming these innovations into benefits and enhancing competitiveness in the context of hardware implementation is another important theme of this disclosure. In particular, issues such as physical compactness and processing latency related to integrated circuit chip design are explained in the final detailed description in Part IV.

[0031] For many filtering and signal processing scenarios, there may be as many output channels as input channels. A relatively simple and easily visualized application scenario is adaptive pulse Doppler processing in radar signal processing, where each specific Doppler channel (after a Fast Fourier Transform (FFT) operation) can be further processed using the remaining Doppler channels as auxiliary or support channels for signal enhancement and interference suppression. From this example, it can be clearly understood that the basic motivation of using the remaining N-1 input channels as several interference suppression support channels to process each of the N input channels leads to the need to deploy a processing architecture with N inputs and N outputs.

[0032] Similar N-input and N-output signal processing scenarios exist in various fields. For example, in biomedical engineering, one approach to studying information processing in the brain is to examine and monitor the brain's naturally generated magnetic fields. Measurements of these fields are called magnetoencephalography (MEG), and they are obtained non-invasively near the surface of the skull. This is a challenging task because neural magnetic fields are very weak. Currently, multi-sensor systems are used for neuromagnetic work. These sensors are placed at different locations on the outer side of the skull. Since the noise from each sensor can be modeled as additive noise that is uncorrelated between sensors, any measurement component that is correlated between sensors corresponds to a neuromagnetic signal. Therefore, the basic N-input and N-output idea shown in the prior art, along with the corresponding parallel signal processing architecture, is also used as a signal enhancement technique.

[0033] The two N-input and N-output signal processing examples described above paint a useful picture at the application level. While this basic N-input and N-output architecture is relatively easy to understand at the application level, deeper ideas and ultimate goals for applying the N-input and 2N-output architecture concept can only be realized in the context of developing general solutions for the entire "permutations space". As mentioned earlier, using a modular and scalable architecture to process some or all possible permutations of a given set of input vectors in parallel represents the focus of this invention. This can only be achieved by applying the novel N-input and 2N-output processing architecture disclosed in this invention.

[0034] Essentially, the coupling between the two application examples above, the architectural concept of N inputs and 2N outputs, and the development of a comprehensive solution for the "permutations space," aims to enhance the understanding of the interrelationships between the many practical applications, applied mathematics, and orthogonalization-driven methods described in this disclosure. Again, the prior art does not involve the concept of parallel processing of all possible permutations of a set of input vectors, nor does it involve the idea of ​​using an architecture with N inputs and 2N outputs, as shown in Section IV. This idea of ​​N inputs and 2N outputs will bring a novel multi-cylinder architectural solution to the aforementioned "permutations space."

[0035] IV. Detailed description of handling parallel-driven innovation:

[0036] The innovations and improvements presented in this disclosure are primarily explained in relation to the MGS algorithm, especially in the context of the basic processing architecture of MGS in Figure 3. Figure 3 shows a parallel processing architecture 30 for this modified Gram-Schmidt (MGS) algorithm, having a set of mutually orthogonal output vectors q, shown as q1, q2, q3, q4, ..., qm-1 and qm, which are located at several diagonal positions in each orthogonalization level. The input vectors / channels in Figure 3 are represented as Xi (i.e., X1, X2, X3, X4, ..., Xm-1 and Xm), and the building blocks are represented as several GSO processing units 3a. The GSO building blocks used for this MGS architecture and the SOC building blocks used for the example architecture with N inputs and two outputs in Figure 4 are illustrated in Figures 1 and 2, respectively, where w, v, v', w', s, and u are vectors. For SOC building blocks, achieving operational symmetry, thereby saving computational resources, is achieved solely by the definition of the inner product in linear algebra. The two-input, two-output SOC building block of Figure 2 can thus utilize this operational symmetry. Given this fact, all subsequent explanations and descriptions in the remainder of this disclosure can be easily understood simply by tracing the movement of all input vectors and intermediate and final output vectors within a given processing architecture. After all, every movement or transfer of each input or output vector is accomplished within a modular, scalable, and entirely conventional processing architecture. In other words, architecturally, the flow and movement of all input and output vectors are self-evident and easy to follow, thus requiring no explicit mathematical description. In fact, describing all the movements and flows of numerous vectors using purely mathematical descriptions would be extremely complex and practically useless for elucidating the details of the entire invention. This point regarding practical usefulness is particularly important when it comes to hardware implementation. Besides lacking the operational symmetry associated with the MGS building blocks of Figure 1, the input data broadcasting line across each orthogonalization layer of the MGS architecture in Figure 3 also negatively impacts the parallel processing capabilities of complex signal processing scenarios. It is worth noting that this will become increasingly clear in the following illustrations. Figure 4 shows an architecture 40 constructed using the aforementioned multiple SOC building blocks 4a. Although conventional art describes this basic example architecture with N inputs and two outputs, it fails to mention the numerous positions of the corresponding two sets of mutually orthogonal output q vectors generated in the processing architecture. Therefore, this specific missing information in the conventional art can be considered an innovative opportunity disclosed in this invention. In short, the orthogonalization-driven architecture shown in Figure 4, with its enhanced parallel processing and operational symmetry, can naturally be used to replace the basic MGS architecture of Figure 3.Figure 4 actually shows two final output channels, replacing the single final output channel shown in the basic MGS architecture of Figure 3. Five input vectors / channels Xi (i.e., X1, X2, X3, X4, and X5) and numerous SOC building blocks 4a arranged in four different tiers (i.e., tiers 1 to 4) are used for the illustrative example in Figure 4.

[0037] Figures 1 to 4 present key background information. This background information, along with the basic concept of an N-input and N-output architecture, is also included in the prior art. However, it must be reiterated that what the prior art does not address is the task of accurately locating the numerous mutually orthogonal output q vectors resulting from the architecture of N-input and two-output, and its extended versions of N-input and N-output. In the language of linear algebra, this particular task (again, this is omitted in the prior art) corresponds to the necessity of computing and identifying the numerous mutually orthogonal row vectors resulting from the Q matrix associated with QR decomposition. To reiterate the three specific deficiencies and unresolved problems associated with the prior art related to this disclosure, they are: (1) the lack of the aforementioned task of identifying numerous mutually orthogonal output q vectors in the processing architecture; (2) the complete absence of consideration for utilizing the basic N-input and 2N-output processing architecture; and (3) the prior art also fails to address the challenge of simultaneously processing all possible permutations associated with a given set of input data vectors. In light of this background, all charts and their contents other than Figure 4 in this disclosure present new ideas and original innovations.

[0038] Due to the importance of accurately locating the numerous mutually orthogonal output q vectors, the basic architectural concept of Figure 4 is reproduced as Figure 5 to include the positions of these q vectors. Additionally, an indexing system for identifying and tracking each possible vector entity represented is also developed and incorporated into Figure 5. Figure 5 illustrates an example 50, indicating a specific set of mutually orthogonal q vectors connected to one of the two final output vectors. The example in Figure 5 includes the following indices and information: an index indicating a specific orthogonalization level in the architecture; an index specifying the position information of a given output vector after a certain number of sequential orthogonalization steps; an index sequence specifying the precise order of the set of orthogonalization steps performed; and the precise positions of the numerous mutually orthogonal q vectors obtained, i.e., qi(x5; x4, x3, x2, x1), where i = 1 to 5 in this case. In this case, the index "i" of qi(x 5; x 4, x 3, x 2, x 1) refers to a specific vector among a given set of mutually orthogonal q vectors, and the first index within the parentheses refers to the "input channel to be processed," i.e., the primary input channel of focus. It is noteworthy that although the remaining indices within the same parentheses seem like a natural way of listing the remaining input channels, the order of these remaining input channels is crucial. Therefore, the sequence of indices within the parentheses actually specifies a specific permutation among all possible permutations associated with a given set of input channel vectors. In other words, in this specific case, x 5 refers to the "input channel to be processed," and the parentheses represent the ordered sequence of orthogonalization steps represented as x 4, x 3, x 2, x 1. Thus, all indices plus the ordered sequence within the parentheses represent the specific permutation of {5, 4, 3, 2, 1} in this particular example. Furthermore, using the indexing system employed here, q5(x5; x4, x3, x2, x1) is the same vector as (4,3,2,1) in Figure 5.

[0039] For completeness and to establish a clear image and an easy-to-follow pattern, in this specific case where the input channel X5 is regarded as the "input channel to be processed" in Figure 5, the five mutually orthogonal q vectors are as follows: q5(x5; x4, x3, x2, x1) =(4,3,2,1); q4(x5; x4, x3, x2, x1) =(2,3,4); q3(x5; x4, x3, x2, x1) =(3,2); q2(x5; x4, x3, x2, x1) =(3); q1(x5; x4, x3, x2, x1) =

[0040] The positions / locations of the aforementioned set of mutually orthogonal q vectors obtained in this architecture are identified based on the fundamental principle that each of the two outputs associated with any given SOC building block, as illustrated in Figure 2, is orthogonal to one of the two input vectors of the same building block. Based on this simple principle and by identifying the zigzagging pattern involved, these mutually orthogonal q vectors can be logically and easily located, as shown in Figure 5.

[0041] Figure 6 illustrates another example 60 that is essentially the same as Figure 5, except that it shows the positions of another set of mutually orthogonal q vectors for the sake of completeness. This is to reaffirm the practicality and usefulness of applying the "zigzagging" technique to track the positions of the multiple mutually orthogonal q vectors. Again, Figure 6 is provided to show the positions / locations of this set of mutually orthogonal output vectors associated with the "other" final output vector (i.e., (2,3,4,5)). It is worth noting that the convention and indexing system used here for this so-called "final output vector" also corresponds to the last vector in a given set of mutually orthogonal output vectors. Specifically, (2,3,4,5) can be understood as the result of processing the first input channel vector (i.e., X1 in this case) according to the ordered orthogonalized sequence of (2,3,4,5), which in turn corresponds to a specific permutation of the multiple input vectors {1, 2, 3, 4, 5}. There is some overlap in expressing all the necessary information and indexes for different vector entities throughout the architecture. This is to ensure consistency and to enable programmable cross-checking in any software development environment. Consistent with the above description and the zigzag vector positioning method, the set of mutually orthogonal q vectors associated with the final output vector (2,3,4,5) are as follows: q5(x1; x2, x3, x4, x5) =(2,3,4,5); q4(x1; x2, x3, x4, x5) =(4,3,2); q3(x1; x2, x3, x4, x5) =(3,4); q2(x1; x2, x3, x4, x5) =(3); q1(x1; x2, x3, x4, x5) = .

[0042] The above indexing system, as shown in Figures 5 and 6, is used to identify and locate numerous mutually orthogonal q vectors. It may appear somewhat complex, involving all indices. However, the entire indexing system is actually very structured. The fundamental orthogonality principle used to locate these mutually orthogonal q vectors in a zigzagging manner is easily understood. To further illustrate this fact, the fundamental orthogonality principle and the zigzag pattern involved are reproduced independently in Figures 5A and 5B as follows:

[0043] We begin with one of the two final output vectors in Figure 4, which should be understood as the last and final q vector of the corresponding set of mutually orthogonal q vectors. In the illustration of Figure 5A, it is designated as qf. As shown in Figure 5A, v' (=qf) is represented in this case as the last and final q vector, i.e., qf located at the bottom of this processing architecture.

[0044] Then the orthogonality characteristic of this building block (SOC) can be applied, that is, the right-hand output vector is orthogonal to the right-hand input vector. Therefore, the previous q vector before the final q vector qf is designated as qf-1, as shown in Figure 5A, and can be architecturally located and determined mathematically. Therefore, the identification and location of the zigzag pattern of the many remaining q vectors qf-1, qf-2, qf-3, ..., q1 performed in a bottom-up manner can be clearly understood and verified.

[0045] For completeness, Figure 5B illustrates, in a diagrammatic way, the same zigzag pattern of a specific set of q vectors for identification and localization starting from the other of the two final output vectors in Figure 4. As shown in Figure 5B, w' (=qf) in this case represents the last and final q vector, i.e., qf located at the bottom of the processing architecture.

[0046] To the best of our knowledge, the prior art does not mention the much-needed and practical method for locating a given set of mutually orthogonal q vectors in a parallel processing architecture, nor does it mention any public domain articles related to signal processing or parallel processing architectures.

[0047] It is worth noting again that, besides the fact that the prior art lacks the recognition of mutually orthogonal q-vectors, another important aspect of the prior art is that it essentially focuses only on comparing the architectural compactness between an MGS-based architecture constructed using the building units shown in Figure 1 and an architecture constructed using the building units shown in Figure 2. The MGS-based architecture is purely an N-input, N-output architecture. For the sake of comparison, the architecture constructed using the building units shown in Figure 2 of the prior art, as expected, is also an N-input, N-output architecture. On the other hand, the "baseline" architecture disclosed in this invention is an N-input, 2N-output architecture.

[0048] To reiterate the same key point: In order to solve any engineering problem or to carry out any innovative hardware implementation that conforms to the QR decomposition principle, it is necessary to clearly identify the positions of the many mutually orthogonal q vectors of the Q matrix in the processing architecture. In other words, before performing any appropriate operations on a given signal processing application scenario, these q vectors in the architecture must be identified and extracted. Therefore, the method for identifying and extracting these q vectors provided in this disclosure, as shown in Figures 5 and 6, and Figures 5A and 5B, presents an important and original innovation compared to the prior art.

[0049] Another point that should be emphasized is that the vector localization method used in Figures 5 and 6 is independent of the number of input channels involved. As long as a bottom-up, zigzag tracing path is followed, the correct positions of the resulting set of mutually orthogonal q vectors can be correctly identified. Through this vector localization method explained and illustrated above, the positions of numerous mutually orthogonal q vectors will therefore not need to be explicitly shown in all subsequent illustrations of this disclosure. It is worth noting that, since the architecture of Figure 4 and the corresponding building blocks of Figure 2 are completely different from those associated with the MGS algorithm in Figures 1 and 3, as expected, the positions of the numerous mutually orthogonal q vectors they each obtain are different. In any case, it is worth noting that the positions of these mutually orthogonal q vectors in this disclosure are identified purely by examining the architecture and following the basic orthogonalization principle, without any mathematical derivation.

[0050] The basic architecture shown in Figure 4 serves as the foundation for solving various signal processing scenarios and situations. The key point here is that the deeply entrenched MGS architecture in Figure 3 is simply insufficient to adequately address these scenarios and situations. One category of identified signal processing scenarios is "window-processing," a common application scenario involving input data in purely serial form. As an innovative contribution of this disclosure, Figures 7 and 8 illustrate, illustratively, how the basic two-dimensional (2-D) triangular architecture of Figure 4 can be compressed into a one-dimensional linear array architecture by utilizing a "window-processing" input data architecture. Figure 7 shows an example 70, indicating an input data overlapping window architecture associated with a "window-processing" scenario. Furthermore, to demonstrate the process of removing redundant hardware, numerous unnecessary SOCs are included to better visualize the idea of ​​hardware removal. Additionally, Figure 7 illustrate, illustratively, how numerous adjacent input data windows can significantly overlap each other. Therefore, the term "window-processing" is used to describe this processing situation. As shown in Figure 7, each triangular architecture surrounded by dashed lines is shown as being directly coupled to a specific input-data-overlapping window, such as the "first window" which uses multiple SOCs to process multiple input vectors X1-X5, the "second window" which uses multiple SOCs to process multiple vectors X2-X6, and the "third window" which uses multiple SOCs to process multiple vectors X3-X7. Furthermore, the three identical and overlapping triangular architectures (i.e., the first, second, and third windows) are grouped together in Figure 7, creating a visual effect of "hardware redundancy." In Figure 7, each input dataset or window enters the architecture one data window at a time. Because each specific input data window significantly overlaps with the windows preceding and following it, the term "input-data-overlapping window processing" is used. The "hardware redundancy" explicitly shown in Figure 7 is intended to illustrate, in a diagram, the possibility of eliminating "redundant hardware" even within a single two-dimensional (2-D) triangular architecture. This naturally leads to the feasibility of compressing the original two-dimensional (2-D) triangular architecture of Figure 4 into the one-dimensional linear array architecture shown in Figure 8.

[0051] Figure 8 illustrates an example 80, which indicates the result of compressing or flattening the architecture containing hardware redundancy in Figure 7 by utilizing the aforementioned input data overlapping window architecture and the storage of the processed results. In Figure 8, the continuous input datasets (i.e., the data windows including the "first window" of vectors X1-X5, the "second window" of vectors X2-X6, and the "third window" of vectors X3-X7) and how they are input into the processing architecture in the manner of "input data overlapping windows" to produce "first window outputs," "second window outputs," and "third window outputs," and the concept of storing the outputs of previous operations for use in the next input dataset, are shown as a data flow diagram 8D, indicating that data is transferred to adjacent memory after each processing of a given data window. In other words, the results stored in the memory associated with the previously processed input dataset or window are used to avoid reprocessing most of the operations that have already been completed. Again, this is to utilize the input data architecture associated with "input data overlapping window processing." The processing order is as follows: after initial processing (if necessary or optional, using the entire 2D triangular architecture for the first input data window), only one fully active computing column 8F, i.e., the right-hand diagonal processing column, is used subsequently. The key point here is that the MGS architecture in Figure 3 is simply incapable of utilizing this 2D-to-1D architecture compression and the "input data overlapping window" scenario. This is primarily due to the lack of architectural symmetry when using the two-input, one-output building blocks of the MGS architecture.

[0052] In addition to the aforementioned innovation in compressing the two-dimensional (2-D) to one-dimensional (1-D) architecture for the "input data overlapping window processing" scenario, the basic two-dimensional (2-D) architecture of N inputs and two outputs in Figure 4 can be extended to solve parallel processing of some or all possible permutations associated with a given set of input data vectors. This is a new challenge in signal processing. To address this "permutation space" in the context of parallel processing, an example involving four input vectors / channels is used for simplicity. In the example of Figure 9 with 4 input channels, "hardware redundancy" is explicitly illustrated, similar to that done in Figure 7, and is also adopted. Figure 9 shows an example 90 that uses the architectural concept of a system with N=4 inputs and 2N outputs. It also includes redundant hardware to emphasize the importance of leveraging architectural and computational symmetries and the idea of ​​removing these unnecessary hardware building blocks. Why the first three input vectors associated with the three input channels processed by the three SOC 9a in a group of 9Ds (i.e., X1, X2, and X3 in Figure 9) are repeated at the top may not be immediately obvious.

[0053] Figure 10 illustrates, in a diagrammatic way, an example 100 of removing the hardware redundancy contained in Figure 9, and presents a highly efficient parallel processing architecture with N inputs and 2N outputs using several SOCs 10a. It is noteworthy that it also implicitly reflects the characteristic of simultaneously processing two permutations of a set of input vectors. This is a subtle but crucial feature, representing the starting point for searching a "global optimal architecture solution" for parallel processing of all possible permutations of a given set of input data vectors. In Figure 10, the "redundant hardware" in the current architecture with N inputs and 2N outputs is removed, i.e., the numerous unnecessary SOCs explicitly drawn in Figure 9, as illustrated in the diagram. Figure 11 is similar to Figure 10, except that it also includes the aforementioned indexing system to illustrate the complete architecture and data flow diagram. Figure 11 shows an example 110 of inserting the aforementioned indexing and tracing system involving all entities into the architecture with N inputs and 2N outputs of Figure 10, similar to Figure 5, for a basic architecture with N inputs and two outputs.

[0054] In order to establish an effective representation with clear signs for any given set of mutually orthogonal q vectors, this disclosure provides a data identification and retrieval system (DIRS) as follows.

[0055] First, a specific final output vector is selected from a total of 2N available vectors as the last q vector in a given set of N mutually orthogonal q vectors. The remaining N-1 q vectors can be identified and retrieved in a bottom-up manner using the DIRS described in this disclosure. The notation used and the data flow concepts involved are consistent with those in Figure 11.

[0056] The entity Xij(m, n, o, ...) can represent any intermediate or final output vector in the entire architecture of N inputs and 2N outputs, as shown in the four-input and eight-output example in Figure 11, where "i" is the index of the number of orthogonalization levels / steps that have been completed, "j" is the index of which particular original input vector is being processed, and (m, n, o, ...) specifies an ordered sequence of several input vectors that have undergone orthogonalization with respect to the j-th input vector. For the sake of simplicity and consistency of notation, the set of input vectors X1, X2, X3, ..., XN is adjusted to X01, X02, X03, ..., X0N to include all original input vectors in the Xij(m, n, o, ...) representation in DIRS. In other words, appending the label "0" to the original input vectors means that all original input vectors have not yet undergone any orthogonalization steps. Therefore, when i = 0, since orthogonalization has not yet been performed, there is no need to append "(m, n, o, ...)".

[0057] In addition, two issues concerning the conventions and symbols for indexing are clarified as follows:

[0058] To reiterate, through a modular processing architecture with N inputs and 2N outputs, 2N final output vectors are generated from a given set of N input vectors. This architecture can be visualized as a hollow cylindrical structure. Therefore, two distinct final output vectors are concatenated with each specific input vector, which is processed to be orthogonal to the remaining N-1 input vectors. It should be noted that the key difference between two distinct final output vectors in a given pair lies in the order of the same set of orthogonalization steps involved; that is, one of the two final output vectors will undergo a specific sequence of N-1 orthogonalization steps in one direction, while the other in the pair will undergo the same set of orthogonalization steps in the reverse order. Furthermore, "i" for N-1 can be visualized as a given input vector being transformed into a final output vector. In Figure 11, for example, X32(3,4,1) and X32(1,4,3) form a pair of final output vectors that are concatenated with the same input vector of X02. Therefore, the order of the superscript "i" is as follows: 0, 1, 2, ..., N-1. Since any such pair of final output vectors is positioned adjacent to each other at the bottom of the architecture, as shown in Figure 11, each such pair of final output vectors can be uniquely labeled as the one on the right (i.e., "R") and the one on the left (i.e., "L"). This is useful for identifying and selecting a specific final output vector as a particular last q vector in a set of N q vectors, and for database design. For example, the pair of final output vectors X32(3,4,1) and X32(1,4,3) in Figure 11 can be labeled and represented as X32(L) and X32(R), respectively. It should be noted that, for the purposes of describing DIRS here, the notations "L" and "R" are strictly used to classify the 2N final output vectors and are used only to identify the last q vector in each group of N q vectors. Specifically, each architecture with N inputs and 2N outputs has 2N groups of N q vectors.

[0059] For example, as shown in FIG11, a method provided by this disclosure includes the steps of: providing a configuration of an architecture having N inputs and 2N outputs, wherein a plurality of building units (i.e., "SOCs") 11a are arranged in a plurality of layers (e.g., "layer 1", "layer 2" and "layer 3"), each layer having N building units 11a, each building unit 11a being capable of performing a symmetric orthogonalization operation, the orthogonalization being associated with two vectors input at two connection ports on both sides (such as X22(3, 4) and X21(4, 3)) and two vectors output at two connection ports on both sides (such as X31(4, 3, 2) and X32(3, 4, 1)). In this configuration, the two vectors (such as X 2 1 (4, 3) and X 3 2 (3, 4, 1)) input and output at the two connection ports on the same side of each building unit 11a are orthogonal. The two connection ports on different sides of adjacent two of the multiple building units 11a, or the two connection ports on different sides of the first and last of the multiple building units 11a, can input one of the N vectors in the first level (e.g., "level 1"). The N building units 11a can output one of the 2N vectors in the last level (e.g., "level 3"). In a higher level and a lower level adjacent to the first and last levels, the two connection ports on different sides of two adjacent building units 11a used for outputting data, or the two connection ports on different sides of the first and last building units 11a in the higher level (e.g., level 1 or level 2) used for outputting data, are linked to the two connection ports on different sides of one of the multiple building units 11a in the lower level (e.g., level 2 or level 3) used for inputting data.

[0060] In addition, as shown in FIG11, the method further includes the step of: retrieving a set of vectors in a linked-ports path formed from a connection port of a building unit 11a in the last layer (e.g., layer 3) to a connection port of a building unit 11a in the first layer (e.g., layer 1) in a zigzag manner from bottom to top.

[0061] In Figures 5A and 5B, the basic "zigzagging-from-bottom-up" approach is illustrated using direct expressions of qN, qN-1, qN-2, ..., q1 to list N q vectors of a given sequence in reverse order. To uniquely specify any particular q vector and map it directly to the generalized expression of Xij(m, n, o, ...) used in the example architecture of N inputs and 2N outputs illustrated in Figure 11, the more detailed and concise notation qs(j, p) is used, where s is a member of the {N, N-1, N-2, ..., 2, 1} sequence, meaning s is an index that operates in a back-to-foreign manner and represents the position of a q vector within a specific set of N q vectors; j indicates the index of the original input vector and the corresponding pair of final output vectors; "p" can be "L" or "R", as explained above.

[0062] Once a specific final output vector with a specific "j" and a specific "p" (i.e., p is "L" or "R", or alternatively, "0" or "1") is selected, the corresponding final q vector qN(j, p) is automatically determined. The retrieval process for the remaining N-1 q vectors of a given set is as follows.

[0063] If p = "L", then the bottom-up retrieval order of index j will follow the following order: j, j-1, j-1+2, j-1+2-3, ..., modulo N. In other words, for example, if j = 2 and N = 4, then the resulting sequence will correspond to: 2, 1, 3, 4 (because the actual counting sequence is as follows: 1, 2, 3, 4, 1, 2, 3, 4, 1, 2, ... for the case of N = 4). Furthermore, the orthogonal sequence (m, n, o, ...) will be "flipped and then with the first index deleted".

[0064] For example, if the final output vector is X 3 2(3,4,1) or X 3 2(L), then q 4(2, L) = X 3 2(3,4,1); q 3(2, L) = X 2 1(4,3); q 2(2, L) = X 1 3(4); q 1(2, L) = X 0 4.

[0065] If p = "R", then the bottom-up retrieval order of index j will follow the following order: j, j+1, j+1-2, j+1-2+3, ..., modulo N. In other words, for example, if j = 2 and N = 4, the resulting sequence will correspond to: 2, 3, 1, 4 (because the actual counting sequence is as follows: 1, 2, 3, 4, 1, 2, 3, 4, 1, 2, ... for the case of N = 4). Furthermore, the orthogonal sequence (m, n, o, ...) will be "flipped and the first index removed".

[0066] For example, if the final output vector is X 3 2(1, 4, 3) or X 3 2(R), then q 4(2, R) = X 3 2(1, 4, 3); q 3(2, R) = X 2 3(4, 1); q 2(2, R) = X 1 1(4); q 1(2, R) = X 0 4.

[0067] The above demonstrates the core framework of DIRS, and illustrates the key innovation of the present invention, the "zigzagging-from-bottom-up" technology, through analysis and logic and illustration.

[0068] To further clarify the relationship between the architecture of N inputs and 2N output types, as illustrated in Figure 11 for the case of N = 4, and to provide the following facts and patterns for the goal of processing all possible permutations of a given set of input data vectors in parallel.

[0069] Fact 1: Starting with a two-dimensional (2-D) parallelogram planar architecture 12a of Figure 12, similar to Example 110 of Figure 11 involving four input vectors / channels, a three-dimensional (3-D) cylindrical architecture 12b having numerous vectors Xa, Xb, Xc, and Xd, after being rolled up, the resulting top view 12c of a cylinder can be formed by connecting the left and right edges of the architecture 12a, as shown in the upper right corner of Example 120 of Figure 12. A key component of the "architecture generalization" method described in this disclosure is to develop a general parallel processing solution for each possible case corresponding to a specific permutation of a given set of input data vectors, as shown in Figure 12.

[0070] The main idea and steps illustrated in Figure 12 involve proposing a unique and powerful solution to address the "permutations space." It involves using an architecture-driven approach to discover and explore the parallelism and symmetry properties of the underlying operations. For example, the number of numerous independent cylindrical architectures that process all possible permutations of a given set of input vectors / channels is (N-1)! / 2, which is determined by examining the operational symmetry patterns on the architecture. Notably, placing the two edges together is equivalent to shortening the length of the three horizontal data broadcast lines of the parallelogram architecture in Figure 11 to "zero," thus forming a three-dimensional hollow cylindrical architecture.

[0071] Fact 2: Once in the form of a tube, one can obtain a top view 12c of processing a tube and visualize it as a circle with four specific entry points, representing the four input channels of the processing tube in that particular example.

[0072] Fact 3: Then all permutations of 12d can be listed. In this example involving four input vectors / channels, there are therefore 4! permutations of 12d, or twenty-four. For example, in the case of 4! permutations of a four-channel array: (1, 2, 3, 4) & (4, 3, 2, 1); (1, 2, 4, 3) & (3, 4, 2, 1); (1, 3, 2, 4) & (4, 2, 3, 1); (1, 3, 4, 2) & (2, 4, 3, 1); (1, 4, 2, 3) & (3, 2, 4, 1); (1, 4, 3, 2) & (2, 3, 4, 1); (2, 1, 3, 4) & (4, 3, 1, 2); (1, 2, 4, 3) & (3, 4, 2, 1); (2, 1, 3, 4) & (4, 3, 1, 2); (2, 1, 4, 3) & (3, 4, 1, 2); (2, 3, 1, 4) & (4, 1, 3, 2); (2, 4, 1, 3) & (3, 1, 4, 2); (3, 1, 2, 4) & (4, 2, 1, 3); (3, 2, 1, 4) & (4, 1, 2, 3).

[0073] Fact 4: In this case, each permutation is represented as a set of four ordered indices. A useful question to ask now is: given a particular permutation that has been processed with a given tubular architecture having four input vectors / channels, what other permutations can be processed in parallel by the same tubular architecture? This question naturally leads to examining the coupling between a tubular architecture with N inputs and 2N outputs, as shown in Figure 11, where N = 4 in this case, and the groups of permutations processed in parallel by the same tubular architecture. For example, if we start with the permutation {1,2,3,4}, then due to the circularly invariant property, the following permutations are actually processed in parallel by the same tubular architecture.

[0074] For example, starting with {1,2,3,4} in a forward loop, the following permutations will be generated: {2,3,4,1}, {3,4,1,2}, {4,1,2,3}; by looping in the other direction, i.e. the backward direction, starting with {4,3,2,1}, the following permutations will be generated: {3,2,1,4}, {2,1,4,3}, {1,4,3,2}.

[0075] Therefore, in this case, all eight different permutations of the above from a given set of four input vectors / channels are processed in parallel by the four-input and eight-output tubular architecture of Figure 11.

[0076] Similarly, starting from {1,3,2,4}, the loop proceeds in a forward direction: {3,2,4,1}, {2,4,1,3}, {4,1,3,2}; then, starting from {4,2,3,1}, the loop proceeds in the opposite direction: {2,3,1,4}, {3,1,4,2}, {1,4,2,3}. Finally, starting from {1,2,4,3}, the loop proceeds in a forward direction: {2,4,3,1}, {4,3,1,2}, {3,1,2,4}. Additionally, starting from {3,4,2,1}, the loop proceeds in the opposite direction: {4,2,1,3}, {2,1,3,4}, {1,3,4,2}. Therefore, with 4 input vectors / channels, all twenty-four permutations are considered.

[0077] Using the bidirectional loop technique and concept described above, the number of independent cylindrical processing architectures required to process all possible permutations in parallel can be systematically determined. Figure 12 illustrates the process and technique based on an example involving four input vectors / channels. It also provides an analytical representation 12e of (N-1)! / 2, showing the number of independent cylindrical processing architectures with N inputs and 2N outputs required to process all possible permutations associated with a given set of N input vectors / channels. In this example where N=4, the number of cylindrical architectures required would therefore be three independent processing cylinders, consistent with the bidirectional loop visualization technique described above.

[0078] As previously stated, parallel processing of all possible permutations may not be applicable or necessary in most practical signal processing scenarios. Nevertheless, establishing the aforementioned analytical relationships is necessary for the architecture-comprehensive approach described in this disclosure in order to develop a system for determining which permutations can or should be selected for a given application scenario.

[0079] Having defined and explained the architecture-driven innovations described above, the next challenge and opportunity for innovation is to define a mapping mechanism that will facilitate compact, latency-reduced, and energy-efficient chip designs or involve hardware implementations using GPUs, FPGAs, etc. One of the purposes of this disclosure is to translate the innovations in this disclosure into hardware implementation-related benefits.

[0080] The innovative mapping and compactness improvement process is illustrated in an example 130 shown in Figure 13. For simplicity, the process begins by drawing the basic N-input and 2N-output architecture 13A as a hollow two-dimensional (2-D) parallelogram configuration 13A of a rolled-up cylinder 13C. In other words, instead of viewing this basic N-input and 2N-output architecture as a hollow cylindrical architecture 13C, it is redrawn as a planar parallelogram similar to Figure 11. It is worth noting that for hardware implementation, a hollow three-dimensional (3-D) cylindrical architecture is clearly undesirable and impractical in terms of achieving spatial compactness and minimizing processing latency. In proposing a computationally efficient and space-compact hardware implementation, this means that an arrangement of (N-1)! / 2 (or less, depending on the final customized solution) number of planar processing parallelograms can first be chained together to form a "long sheet of processing parallelograms," and then this long sheet of processing parallelograms is rolled up. In other words, to process these arrangements in parallel using an efficient and compact arrangement involving minimal "integrated circuit real estate," multiple independent parallelograms are first chained together. Then, this "long sheet" of processing parallelograms is rolled up into a rolled-up cylindrical structure, as shown in Figure 13.

[0081] To restate and further elaborate, Figure 13 illustrates Example 130, which presents an innovation in hardware implementation 13A. An architecture 13A with N inputs and 2N outputs is viewed as a planar parallel processing parallelogram architecture (rather than visualized as a three-dimensional (3-D) cylindrical architecture 13C comprising multiple hardware implementations 13A), many of which can be considered as first being cascaded together, forming "long sheets of planar parallelogram shapes." These planar parallelogram shapes can then be rolled into a spatially compact architecture, eliminating the problem of large internal hollow spaces associated with each individual regularized cylindrical processing architecture. This has significant hardware implementation implications, including the ability and flexibility to minimize "chip real estate" in the IC chip design process. Compared to any MGS-based approach, this architecture-wide approach also reduces power consumption and processing latency.

[0082] As mentioned earlier, the concept of processing all possible permutations associated with a given set of input data vectors may be computationally impractical for many current applications. On the other hand, advancements and breakthroughs in high-tech fields such as accelerated computing, artificial intelligence chip design, and quantum computing are happening at an astonishing pace.

[0083] The final integration innovations and architectural breakthroughs related to this disclosure involve the popular "batch mode processing" application scenario, which is a common and widely used signal processing situation.

[0084] Typically, when using an architecture with N inputs and 2N outputs to process a given set of input vectors, as shown in Figure 11, for the case of N=4, a given set of input vectors is first processed by a number of SOCs in the first column of the processing architecture (defined in Figure 2), and then the number of output vectors generated by the number of SOCs in the top column becomes the input vectors of the number of SOCs in the second column, and so on. On the other hand, in the so-called "batch mode processing" scenario defined here, when a given column of SOCs performs a number of corresponding operations, the number of SOCs in all other columns are (or may be) idle, depending on the specific signal processing situation or the hardware implementation options involved. Considering this idle operation situation or the arrangement associated with this batch mode processing scenario, the technique of using feedback loops can be deployed to first compress the basic parallelogram into a linear array architecture, as shown in Figure 14. For example, Figure 14 shows an example 140, which proposes an innovation reflecting an ultra-compact hardware implementation arrangement 14A for the popular "batch mode processing" scenario. This corresponds to a repeatable module 14B where only one column of the SOC is operationally active while the remaining columns are idle or can be left idle. In this case, each output vector generated after orthogonalization of the first level of the architecture is fed back as an input vector to an adjacent SOC. Therefore, this feedback loop configuration compresses the initial architecture into one containing only a single column of SOC. In other words, a single SOC linear array replaces the entire parallelogram architecture with N inputs and 2N outputs. (Note: Again, if two opposite edges are connected, the "parallelogram" here can also be visualized as a "cylinder" 14C comprising multiple repeatable modules 14B). Since each architecture with N inputs and 2N outputs processes 2N of the total N! permutations, N! / 2N or (N-1)! / 2 independent SOC linear arrays are needed to cover all possible permutations in this case. Therefore, instead of rolling up a long parallelogram structure as shown in Figure 14, we complete the rolling up of (N-1)! / 2 end-to-end line arrays, as shown in Figure 14. To reiterate, this basic parallelogram structure with N inputs and 2N outputs becomes a single line array, which, after inserting all feedback loops, can be considered simply as the first column of the original parallelogram structure with N inputs and 2N outputs (SOC).

[0085] For tasks requiring multiple linear array architectures to handle all or part of the arrangement, the rolling action is the same as rolling up multiple parallelograms. This is also clearly shown in Figure 14. It is worth noting that instead of using a single column of SOCs and arranging multiple such single column SOCs before the rolling action, a group of such single column SOCs can be stacked before rolling. This will result in a "thicker" or "taller" rolled-up ring architecture.

[0086] It is worth noting that the basic two-dimensional (2-D) triangular parallel processing, as shown in Figure 4 for performing QR decomposition, is completely different from the MGS architecture refined in Figure 3. This disclosure uses illustrations throughout to explain the various innovations used to upgrade or optimize it into one-dimensional (1-D) linear array architectures and three-dimensional (3-D) cylindrical architectures for various application scenarios. These innovations and architectural breakthroughs are made possible largely by new insights gained from re-examining the advantages and disadvantages of the MGS algorithm and its corresponding parallel processing architecture from seventy years ago, as well as by discovering certain previously unexplored operational symmetries and parallel processing concepts related to the basic QR decomposition concept.

[0087] QR decomposition has been, and will continue to be, a powerful linear algebra tool applicable to a wide range of applications, particularly in key areas such as machine learning, signal processing, and the emerging "6G" future. The term "architecturally comprehensive methodology" was strategically and logically chosen as part of the title of this disclosure to reflect that all novel ideas and innovations are largely inspired and driven by architecture. Another reason for emphasizing the "comprehensiveness" factor is that this method can provide architecturally optimized solutions for a variety of signal processing and machine learning scenarios while mathematically adhering to the principles and theory of QR decomposition. To the best of our knowledge, this orthogonalization-driven and architecturally comprehensive methodology (which can be viewed as a well-defined system of interwoven and interconnected architecture-driven innovations) represents a groundbreaking technological milestone. Essentially, it integrates processing parallelism, QR decomposition, and key hardware implementation considerations in a cohesive and mathematically verifiable manner.

[0088] Alternatively, a suitable computing system can be used to perform the signal processing operations described above. For example, Figure 15 depicts an example of a computing device 150 that can implement the methods described herein, such as computer implementations of signal processing methods.

[0089] In some embodiments, the computing device 150 may include a processor 151 coupled to a memory 152 and configured to execute a plurality of program instructions stored in the memory 152 to perform a plurality of operations for implementing computer-implemented methods associated with signal processing.

[0090] For example, the processor 151 may include a microprocessor, an application-specific integrated circuit ("ASIC"), a state machine, or other processing device. The processor 151 may include one or more processing units. Such a processor may include or be in communication with a computer-readable medium storing or containing instructions that, when executed by the processor 151, cause the processor to perform the operations described herein. The memory 152 may include any suitable non-transitory computer-readable medium.

[0091] For example, the computer-readable medium may include any electronic, optical, magnetic, or other storage device capable of providing a plurality of computer-readable instructions or other program code to a processor. Non-limiting examples of a computer-readable medium include magnetic disks, memory chips, ROM, RAM, ASICs, configured processors, optical storage, magnetic tape or other magnetic storage, or any other medium in which a computer processor can read a plurality of instructions. These instructions may include a plurality of processor-specific instructions generated by a compiler and / or an interpreter from program code written in any suitable computer programming language, including, for example, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, and ActionScript.

[0092] Additionally, a suitable non-transitory machine-readable medium can be used to perform the above-described signal processing operations, as shown in Figures 5 to 6 and Figure 11. For example, a non-transitory machine-readable medium storing a plurality of instructions, which, when executed by the processor, cause the processor to perform a plurality of steps, including: providing a model having an architecture with N inputs and 2N outputs, wherein a plurality of building units are arranged in a plurality of layers, each layer having N building units, each building unit being able to perform a symmetric orthogonalization operation associated with two vectors input to two ports on both sides and two vectors output to two ports on both sides, wherein the two vectors input to and output to the two ports on the same side of each of the plurality of building units are orthogonal, and adjacent two of the N building units in the first layer are connected to ports on different sides or the first of the N building units in the first layer. Two ports on different sides of the last one can input one of N vectors. The N building units in the last layer can output 2N vectors. In a higher layer and a lower layer between the first layer and the last layer, two ports on different sides of two adjacent building units for outputting data, or two ports on different sides of the first and last building units in the higher layer for outputting data, are linked to two ports on different sides of one of the multiple building units in the lower layer for inputting data. A set of vectors is retrieved in a zigzag bottom-up manner from a port on one of the building units in the last layer to a port on one of the building units in the first layer.

[0093] Additionally, various signal processing operations described above can be performed using appropriate methods, as shown in Figures 5 to 6 and Figure 11. For example, Figure 16 depicts a method example 160, which can implement various methods described herein, such as computer-implemented methods for processing signals. Method example 160 indicates a computer-implemented method including steps 161 and 163, wherein in step 161, a configuration is provided having an architecture with N inputs and 2N outputs, wherein multiple building units are arranged in multiple layers, each layer having N building units, and each building unit is capable of performing a symmetric orthogonalization operation associated with two vectors input to two ports on both sides and two vectors output to two ports on both sides, wherein the two vectors input to and output by each of the multiple building units on the same side of the two ports are orthogonal, and the adjacent two in the N building units in the first layer are on different sides of the two ports, or the first and last in the N building units in the first layer are on different sides of the two ports. The port can input one of N vectors, and the N building units in the last layer can output 2N vectors. In a higher layer and a lower layer between the first layer and the last layer, two connection ports used for outputting data on different sides of two adjacent building units, or two connection ports used for outputting data on different sides of the first and last building units in the higher layer, are linked to two connection ports used for inputting data on different sides of one of the multiple building units in the lower layer. In step 163, a set of vectors in a link path formed from a connection port of a building unit in the last layer to a connection port of a building unit in the first layer is retrieved in a zigzag manner from bottom to top.

[0094] Furthermore, a suitable system can be used to perform the various operations of the above signal processing, as shown in Figures 5 to 6 and Figure 11. For example, a system includes a processor coupled to a memory storing a plurality of instructions to be executed in the processor, the processor being configured to perform a plurality of operations, including: providing a configuration having an architecture of N inputs and 2N outputs, wherein a plurality of building units are arranged in a plurality of layers, each layer having N building units, each building unit being capable of performing a symmetric orthogonalization operation associated with two vectors input to two connection ports on both sides and two vectors output to two connection ports on both sides, wherein the two vectors input to and output to the two connection ports on the same side of each of the plurality of building units are orthogonal, and adjacent two of the N building units in the first layer are connected to two connection ports on different sides or in the first layer of the N building units. Two ports on different sides of the first and last layer can input one of N vectors; the N building units in the last layer can output one of 2N vectors; and in a higher layer and a lower layer between the first layer and the last layer, two ports on different sides of two adjacent building units used for outputting data, or two ports on different sides of the first and last building units in the higher layer used for outputting data, are linked to two ports on different sides of one of the multiple building units in the lower layer used for inputting data; and a set of vectors is retrieved in a zigzag bottom-up manner from a port on one of the building units in the last layer to a port on one of the building units in the first layer.

[0095] Furthermore, a suitable circuit can be used to perform the above signal processing operation, as shown in Figures 5 to 6 and Figure 11. For example, a circuit implemented in a chip includes: a sub-circuit having an architecture of N inputs and 2N outputs, wherein multiple building units are arranged in multiple layers, each layer having N building units, and each building unit can perform a symmetric orthogonalization operation associated with two vectors input to two connection ports on both sides and two vectors output to two connection ports on both sides, wherein the two vectors input to and output by each of the multiple building units on the same side of the two connection ports are orthogonal, and the two connection ports on different sides of the N building units in the first layer or the two connection ports on different sides of the first and last of the N building units in the first layer can input one of the N vectors. The N building blocks in the last layer can output 2N vectors, and in a higher layer and a lower layer between the first layer and the last layer, two ports for outputting data on different sides of two adjacent building blocks, or two ports for outputting data on different sides of the first and last building blocks in the higher layer, are linked to two ports for inputting data on different sides of one of the plurality of building blocks in the lower layer; and a logic section is configured to retrieve a set of vectors in a zigzag bottom-up manner from a port of a building block in the last layer to a port of a building block in the first layer.

[0096] In some embodiments, the zigzag bottom-up retrieval of the set of vectors in the link path formed from the port of a building unit in the last layer to the port of a building unit in the first layer includes: setting a port of one of the plurality of building units in the last layer as a target port; and recursively performing a process until the target port is a port for input data in the first layer, the process including: in response to determining that the target port is a port for output data, setting the vector at the target port as one of the set of vectors, and setting the port for input data on the same side of the same building unit as the target port as the updated target port; in response to determining that the target port is a port for input data and is not in the first layer, setting the port for output data linked to the target port as the updated target port; and in response to determining that the target port is a port for input data in the first layer, setting the vector at the target port as one of the set of vectors.

[0097] In some embodiments, the N building blocks in each layer are input data in a cyclic manner; the plurality of building blocks in this configuration are wrapped in a cylindrical structure; the cylindrical structure is fabricated in a three-dimensional integrated circuit.

[0098] Although the present invention has been disclosed in preferred embodiments, they are not intended to limit the invention. Various modifications and variations can be made to the invention by any person skilled in the art without departing from its spirit and scope. Therefore, the scope of protection of the present invention should be determined by the scope of the appended patent application. [Simplified Explanation of the Diagram]

[0009] The reference diagrams provide a better understanding of this disclosure, and the concise descriptions of each diagram make the subsequent detailed descriptions significantly easier to understand. Figure 1 is a schematic diagram illustrating a Gram-Schmidt orthogonalization (GSO) building block relevant to this disclosure. Figure 2 is a schematic diagram illustrating a symmetric orthogonalization operation unit (SOC) building block employed and used in this disclosure. Figure 3 is a schematic diagram illustrating a parallel processing architecture using a modified Gram-Schmidt algorithm employing the GSO building block of Figure 1. Figure 4 is a schematic diagram illustrating a processing architecture constructed using the SOC building block of Figure 2. Figure 5 is a schematic diagram illustrating an indexing and data tracking system integrated with the processing architecture of Figure 4, showing the position of one of two sets of mutually orthogonal output q vectors. Figures 5A and 5B are schematic diagrams illustrating a unique and efficient method for locating q vectors in a bottom-up manner. Specifically, Figures 5A and 5B are inserted to further enhance the understanding of the techniques used to locate these q vectors in Figures 5 and 6, respectively. Figure 6 is identical to Figure 5, except that it alternatively indicates the positions of another set of mutually orthogonal q vectors. Figure 7 is a schematic diagram illustrating a data window processing scenario involving several overlapping windows of input data that can utilize the methods proposed in this disclosure. Figure 8 is a schematic diagram illustrating a data flow and architecture compression diagram associated with the data window processing scenario of Figure 7. Figure 9 is a schematic diagram illustrating an example of a processing architecture corresponding to a system with N inputs and 2N outputs in this disclosure. Figure 10 is a schematic diagram illustrating the concept and task of deduplicating redundant hardware as explicitly shown in the processing architecture of Figure 9. Figure 11 is a schematic diagram illustrating an indexing and data tracking system integrated with the efficient hardware processing architecture of Figure 10. Figure 12 is a schematic diagram illustrating the main ideas and steps involved in solving general problems involving the entire "permutation space" in conjunction with this disclosure. Figure 13 is a schematic diagram illustrating a hardware implementation of the processing architecture of Figure 11. Figure 14 is a schematic diagram illustrating an innovation in another efficient hardware implementation of this disclosure. Figure 15 is a schematic diagram illustrating a signal processing system that can be applied to this disclosure. Figure 16 is a schematic diagram illustrating a flowchart depicting a method for implementing the computer implementation of signal processing in this disclosure.

Claims

1. A computer-implemented method, comprising: Provided is a configuration with an architecture of N inputs and 2N outputs, wherein multiple building blocks are arranged in multiple layers, each layer having N building blocks. Each building block can perform a symmetric orthogonalization operation associated with two vectors input to and two vectors output from two connection ports on both sides. The input and output vectors from the connection ports on the same side of each of the multiple building blocks are orthogonal. In the first layer, adjacent units of the N building blocks on different sides, or the first and last units of the N building blocks on different sides, can input one of the N vectors to the connection ports on different sides. The N building blocks in each layer can output 2N vectors, and in a higher layer and a lower layer between the first layer and the last layer, two ports used for outputting data on different sides of two adjacent building blocks, or two ports used for outputting data on different sides of the first and last building blocks in the higher layer, are linked to two ports used for inputting data on different sides of one of the multiple building blocks in the lower layer, where N is a positive integer; and a set of vectors is retrieved in a zigzag bottom-up manner from a port of a building block in the last layer to a port of a building block in the first layer in a linked path.

2. The computer-implemented method as described in claim 1, wherein retrieving the set of vectors in the port path formed by connecting a port of a building unit in the last layer to a port of a building unit in the first layer in a bottom-up zigzag manner includes: Set one of the multiple building units in the last layer as a target connection port; The process recursively continues until the target port is a port used for input data in the first layer. This process includes: in response to determining that the target port is a port used for output data, setting the vector at the target port as one of the vectors in the set, and setting the input data port on the same side of the same building unit as the target port as the updated target port; in response to determining that the target port is a port used for input data and is not in the first layer, setting the output data port linked to the target port as the updated target port; and in response to determining that the target port is a port used for input data in the first layer, setting the vector at the target port as one of the vectors in the set.

3. The computer-implemented method as described in claim 1, wherein the N building blocks in each layer are input data in a cyclic manner.

4. The computer-implemented method as described in claim 1, wherein the plurality of building units in the configuration are rolled up in a cylindrical structure.

5. The computer-implemented method as described in claim 4, wherein the cylindrical structure is fabricated in a three-dimensional integrated circuit.

6. A computing system comprising a processor coupled to a memory storing a plurality of instructions to be executed in the processor, the processor being configured to perform a plurality of operations, including: Provided is a configuration with an architecture of N inputs and 2N outputs, wherein multiple building blocks are arranged in multiple layers, each layer having N building blocks. Each building block can perform a symmetric orthogonalization operation associated with two vectors input to and two vectors output from two connection ports on both sides. The input and output vectors from the connection ports on the same side of each of the multiple building blocks are orthogonal. In the first layer, adjacent units of the N building blocks on different sides, or the first and last units of the N building blocks on different sides, can input one of the N vectors to the connection ports on different sides. The N building blocks in each layer can output 2N vectors, and in a higher layer and a lower layer between the first layer and the last layer, two ports used for outputting data on different sides of two adjacent building blocks, or two ports used for outputting data on different sides of the first and last building blocks in the higher layer, are linked to two ports used for inputting data on different sides of one of the multiple building blocks in the lower layer, where N is a positive integer; and a set of vectors is retrieved in a zigzag bottom-up manner from a port of a building block in the last layer to a port of a building block in the first layer in a linked path.

7. The computing system as described in claim 6, wherein the retrieval of the set of vectors in the port path formed by the connection port of a building unit in the last layer to the connection port of a building unit in the first layer in a bottom-up zigzag manner includes: Set one of the multiple building units in the last layer as a target connection port; The process recursively continues until the target port is a port used for input data in the first layer. This process includes: in response to determining that the target port is a port used for output data, setting the vector at the target port as one of the vectors in the set, and setting the input data port on the same side of the same building unit as the target port as the updated target port; in response to determining that the target port is a port used for input data and is not in the first layer, setting the output data port linked to the target port as the updated target port; and in response to determining that the target port is a port used for input data in the first layer, setting the vector at the target port as one of the vectors in the set.

8. The computing system as described in claim 6, wherein the N building blocks in each layer are input data in a cyclic manner.

9. The computing system as described in claim 6, wherein the plurality of building blocks in the configuration are enclosed in a cylindrical structure.

10. The computing system as described in claim 9, wherein the cylindrical structure is fabricated in a three-dimensional integrated circuit.

11. A circuit implemented in a chip, comprising: A sub-circuit with an architecture of N inputs and 2N outputs, wherein multiple building blocks are arranged in multiple layers, each layer having N building blocks. Each building block can perform a symmetric orthogonalization operation associated with two vectors input to and output from two ports on either side. The input and output vectors from the same side ports of each of the multiple building blocks are orthogonal. In the first layer, adjacent units of the N building blocks on different sides, or the first and last units of the N building blocks on different sides, can input one of the N vectors. In the last layer, the N... Each building unit can output 2N vectors, and in a higher level and a lower level between the first level and the last level, two ports on different sides of two adjacent building units for outputting data, or two ports on different sides of the first and last building units in the higher level for outputting data, are linked to two ports on different sides of one of the plurality of building units in the lower level for inputting data, where N is a positive integer; and a logic section configured to retrieve a set of vectors in a zigzag bottom-up manner from a port on one of the building units in the last level to a port on one of the building units in the first level.

12. The circuit as claimed in claim 11, wherein the logic portion is further configured to: designate a port of one of the plurality of building blocks in the last layer as a target port; and recursively perform a process until the target port is a port used for inputting data in the first layer, the process comprising: In response to determining that the target connection port is a connection port for outputting data, the vector at the target connection port is set as one of the vectors in the group, and the connection port for inputting data on the same side of the same building unit as the target connection port is set as the updated target connection port; In response to determining that the target connection port is a connection port for inputting data and is not in the first layer, the connection port for outputting data linked to the target connection port is set as the updated target connection port; and In response to determining that the target connection port is a connection port for inputting data in the first layer, the vector at the target connection port is set as one of the vectors in the group.

13. The circuit as claimed in claim 11, wherein the N building blocks in each layer receive data in a cyclic manner.

14. The circuit as claimed in claim 11, wherein the plurality of building blocks in the configuration are wound in a cylindrical structure.

15. The circuit as described in claim 14, wherein the cylindrical structure is fabricated in a three-dimensional integrated circuit.

16. A non-transitory machine-readable medium storing a plurality of instructions, which, when executed by a processor, cause the processor to perform a plurality of steps, including: A model is provided with an architecture of N inputs and 2N outputs, wherein multiple building blocks are arranged in multiple layers, each layer having N building blocks. Each building block can perform a symmetric orthogonalization operation associated with two vectors input to and two vectors output from two connection ports on both sides. The input and output vectors from the connection ports on the same side of each of the multiple building blocks are orthogonal. In the first layer, adjacent units of the N building blocks on different sides, or the first and last units of the N building blocks on different sides, can input one of the N vectors to their respective connection ports. The N building blocks in each layer can output 2N vectors, and in a higher layer and a lower layer between the first layer and the last layer, two ports used for outputting data on different sides of two adjacent building blocks, or two ports used for outputting data on different sides of the first and last building blocks in the higher layer, are linked to two ports used for inputting data on different sides of one of the multiple building blocks in the lower layer, where N is a positive integer; and a set of vectors is retrieved in a zigzag bottom-up manner from a port of a building block in the last layer to a port of a building block in the first layer in a linked path.

17. The non-transitory machine-readable medium as described in claim 16, wherein the processor, while executing the stored plurality of instructions, is further configured to perform a plurality of steps, including: Set one of the multiple building units in the last layer as a target connection port; The process recursively continues until the target port is a port used for input data in the first layer. This process includes: in response to determining that the target port is a port used for output data, setting the vector at the target port as one of the vectors in the set, and setting the input data port on the same side of the same building unit as the target port as the updated target port; in response to determining that the target port is a port used for input data and is not in the first layer, setting the output data port linked to the target port as the updated target port; and in response to determining that the target port is a port used for input data in the first layer, setting the vector at the target port as one of the vectors in the set.

18. A non-transitory machine-readable medium as described in claim 16, wherein the N building blocks in each layer are input data in a cyclic manner.

Citation Information

Patent Citations

  • Large-scale matrix QR decomposition parallel computing structure

    CN111858465A

  • Method and system for performing a calculation operation and a device

    TW200414023A

  • Spatial locality transform of matrices

    TW202034191A

  • Method and system for orthogonalizing input signals

    US20120215482A1