Baseband processor based on total compatibility and control method
By dynamically allocating PE units and realizing MIMO balance and AI computing resource sharing based on a baseband processor based on universal computing integration, the problem of the inability to dynamically allocate computing resources of traditional baseband processors is solved, and computing efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202411486827.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-10-23
AI Technical Summary
Traditional baseband processors cannot achieve dynamic allocation of computing resources, cannot realize resource sharing for communication optimization and neural network computing, and have limited computing power and low resource utilization.
A baseband processor based on universal computing integration is adopted to realize dynamic allocation of PE units and sharing of computing resources through signal input module, signal preprocessing module, dynamic precision adjustable PE module, computing resource scheduling module, signal decoding module and neural network preprocessing module. The dynamic precision adjustable PE module includes a control unit, a storage unit and a computing unit. The computing unit contains several PE units, and resources are allocated through the computing resource scheduling module.
It realizes the shared allocation of PE units for MIMO equalization and AI computing, reduces the computing resource demand limit, improves computing resource utilization, optimizes computing efficiency and reduces hardware overhead.
Smart Images

Figure CN119364429B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of communication and AI, in particular to a baseband processor based on communication and AI computing. BACKGROUND
[0002] With the continuous evolution of communication, future 6G will cross people and things, and move towards intelligent connection of all things. In order to achieve this, the 6G end-to-end communication system needs to support AI (Artificial Intelligence) and ML (Machine Learning) computing well, so as to achieve the best efficiency. Therefore, in order to truly support universal intelligence, it is necessary to provide high-performance connection, high-capability computing and rich storage resources at the network edge.
[0003] 6G needs to provide high-performance wireless connection, which can be comparable to fiber in speed, specifically Tbps-level peak rate, 10-100 Gbps experience rate, sub-millisecond latency, and 10 times the connection density of 5G. As a result, the baseband signal processing of communication is required to achieve extremely high processing rate, extremely high parallelism, extremely large computing power and extremely high energy efficiency. However, one of the challenges currently faced is that, due to the advent of the post-Moore era, integrated circuits are facing performance bottlenecks. Therefore, it will be difficult to continue to design and implement future 6G baseband algorithms by simply stacking traditional resources.
[0004] With the development of 6G and AI, the boundaries between future communication and computing are gradually blurred, and communication networks and devices will support AI and ML computing. Baseband processors increasingly integrate AI / ML algorithms, which puts new requirements on the computing power of baseband processors.
[0005] Traditional baseband processors use separate and independent computing resources for communication and AI, have limited computing power, cannot complete intensive computing tasks in AI, and have low resource utilization. If the computing resources can be shared for communication and AI computing, the computing power can be greatly improved, and the utilization of hardware resources can be improved.
[0006] However, there are great challenges in implementing communication and AI computing resource sharing in baseband. On the one hand, traditional baseband processors are not flexible enough and the hardware is relatively fixed, which cannot adapt to flexible baseband processing algorithms for different communication quality. Secondly, implementing communication and AI computing resource sharing in baseband requires communication and AI computing to have a common computing paradigm, while the similarity of communication baseband algorithms and AI algorithms is limited.
[0007] Finally, traditional baseband processors do not have a computing resource allocation mechanism, and cannot schedule computing resources according to actual communication conditions and hardware constraints. SUMMARY
[0008] The application aims to provide a baseband processor based on communication and computing fusion, realize dynamic allocation of PE units (Processing Element), realize sharing of computing resources for communication computing and neural network computing, and solve the problems that the traditional baseband processor cannot realize dynamic allocation of computing resources according to communication requirements, and cannot realize sharing of computing resources for communication optimization and neural network computing.
[0009] In order to achieve the above-mentioned purpose of the application, the application provides the following technical solutions.
[0010] The baseband processor based on communication and computing fusion comprises:
[0011] A signal input module is configured to receive a communication signal.
[0012] A signal preprocessing module is configured to perform signal preprocessing on the communication signal received by the signal input module, and then transmit the preprocessed communication signal to a dynamic precision adjustable PE module.
[0013] A computing resource scheduling module is configured to schedule the dynamic precision adjustable PE module.
[0014] The dynamic precision adjustable PE module is configured to perform MIMO (Multiple Input Multiple Output) equalization and / or AI computing.
[0015] The dynamic precision adjustable PE module comprises a control unit, a storage unit and a computing unit, and the computing unit comprises a plurality of PE units; the control unit controls the computing unit to read data in the storage unit for computing and processing.
[0016] A signal decoding module is configured to receive the communication signal after MIMO equalization of the dynamic precision adjustable PE module, and perform decoding operation.
[0017] A neural network preprocessing module is configured to preprocess information content obtained through the decoding operation, and obtain a signal for neural network computing; then, the signal is input into the dynamic precision adjustable PE module for AI computing, and an AI computing result is obtained.
[0018] An AI result output module is configured to output the AI computing result.
[0019] The baseband processor fuses general computing common mode, realizes sharing and distribution of PE units required by MIMO equalization calculation and AI calculation, realizes sharing of the calculation capability of the baseband processor, makes dynamic distribution of the calculation resources for MIMO equalization and the resources for AI calculation, reduces the limitation of the calculation resource requirement, and improves the utilization rate of the calculation resources.
[0020] Further, the neural network preprocessing module pre-processes the information content of the decoding operation, that is, normalizes the data.
[0021] Further, the MIMO equalization is that the dynamic precision adjustable PE module equalizes the communication signal subjected to signal preprocessing by using a linear minimum mean square error algorithm.
[0022] The AI calculation is that the neural network calculates the information content subjected to the data normalization to obtain a digital sequence.
[0023] Further, the calculation resource scheduling module can control the dynamic precision adjustable PE module to distribute PE unit resources according to the communication quality, hardware constraints and task requirements to complete calculation.
[0024] Further, the PE unit at least includes a complex multiplier and an arithmetic logic unit, and the calculation precision of the PE unit is adjustable when the PE unit performs MIMO equalization and AI calculation.
[0025] Further, the AI result output module receives the AI calculation result, restores the AI calculation result into corresponding information content, and displays the information content.
[0026] Further, the calculation resource scheduling module can control the dynamic precision adjustable PE module to distribute the number of PE units for MIMO equalization and the number of PE units for AI calculation.
[0027] Further, the dynamic precision adjustable PE module dynamically distributes PE units to perform calculation according to the control signal given by the calculation resource scheduling module.
[0028] The application further provides a general computing common fusion control method using the baseband processor, which includes two stages, that is, a communication calculation stage and an AI calculation stage.
[0029] The first stage is the communication calculation stage.
[0030] The communication signal is received.
[0031] The input signal is pre-processed.
[0032] The pre-processed signal is sent to the dynamic precision adjustable PE module for MIMO equalization;
[0033] Decoding the communication signal after MIMO equalization;
[0034] The second stage, AI computing stage:
[0035] The decoded signal is sent to the neural network preprocessing module for preprocessing;
[0036] The signal after neural network preprocessing is sent to the dynamic precision adjustable PE module for AI calculation;
[0037] Output the results of AI calculation.
[0038] Furthermore, the channel capacity of the MIMO communication system is calculated, the time required for MIMO equalization is calculated according to the channel capacity, and the number of PE units allocated to MIMO equalization is determined.
[0039] Furthermore, the number of PE units allocated to AI calculation is determined according to the number of PE units allocated to MIMO equalization.
[0040] Preferably, the calculation formula of the channel capacity is as follows:
[0041]
[0042] Where C is the channel capacity; k1 and ρ are coefficient factors; N r is the number of columns of the channel matrix (i.e. the number of receiving antennas); is the quantization error matrix H q The average singular value of is the average eigenvalue of the covariance matrix Q; Q d is the quantization bit width of the calculation element; R R Normalize the receiving Hermit matrix to the unit diagonal element; is the variance of the additive white Gaussian noise.
[0043] Preferably, the calculation formula for the time required for MIMO equalization is as follows:
[0044]
[0045] Among them, T cycle is the time required for MIMO equalization; D is the amount of data input for communication; and C is the channel capacity.
[0046] Preferably, the calculation formula for the number of PE units allocated to MIMO equalization is as follows:
[0047]
[0048] P com is the number of PE units allocated to MIMO equalization; epsilon is the resource required by MIMO equalization; N t is the number of rows of the channel matrix, that is, the number of transmitting antennas; T cycle is the time required by MIMO equalization; and xi(Q) is the resource consumption of a single PE unit when performing MIMO equalization when the actual bit width of the calculation element is Q.
[0049] The calculation formula of the number of PE units allocated to AI calculation is as follows:
[0050] P nnc = P-P com
[0051] P nnc is the number of PE units allocated to AI calculation, P is the total number of PE units, and P com is the number of PE units allocated to MIMO equalization.
[0052] The general-purpose computing fusion of the application refers to the general-purpose allocation of comprehensive computing resources, the fusion of signal processing and AI calculation, the sharing of the same PE array and computing resources by the two, and the reduction of the development pressure of the baseband processor.
[0053] Compared with the prior art, the application has the following beneficial effects:
[0054] The application can use the same PE unit for MIMO equalization and AI calculation, solves the problem that the communication optimization and neural network calculation of the traditional baseband processor cannot share resources, and solves the problem that the traditional baseband processor cannot dynamically allocate computing resources according to communication requirements.
[0055] The baseband processor based on general-purpose computing fusion of the application adopts a dynamic PE resource scheduling algorithm, allocates computing resources according to communication quality, task requirements, hardware requirements and other constraints, so that the final computing efficiency is optimal and the hardware overhead is minimal, and the design pressure of the baseband processor is greatly reduced. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 is a schematic structural diagram of the baseband processor based on general-purpose computing fusion;
[0057] Figure 2 is a schematic structural diagram of the dynamic precision adjustable PE module of the baseband processor based on general-purpose computing fusion.
[0058] Figure 3 is a working timing diagram of the baseband processor based on general-purpose computing fusion. DETAILED DESCRIPTION
[0059] A baseband processor for communication and fusion, mainly comprising the following modules:
[0060] Signal input module: for receiving communication signal input.
[0061] Signal preprocessing module: for preprocessing the input signal; preferably, signal preprocessing includes one or more of precoding, modulation, serial-parallel conversion, parallel-serial conversion, etc.
[0062] Computing resource scheduling module: for scheduling the dynamic precision adjustable PE module, which generates corresponding control signals for PE resource allocation under the guidance of the optimal scheduling algorithm by accepting communication quality, task demand, hardware constraints, etc. from the signal preprocessing and neural network preprocessing modules. This makes the total computing time minimal, the computing efficiency highest, and the hardware overhead minimal.
[0063] Dynamic precision adjustable PE module: the module integrates a control unit, a computing unit, and a storage unit. The computing unit is composed of a series of PE units that can be flexibly configured. The control unit can flexibly configure the PE unit according to communication quality, hardware constraints, task demand, etc. to complete communication MIMO equalization and AI calculation. Under the guidance of the control signal from the computing resource scheduling module, the PE module performs MIMO equalization and AI calculation with appropriate resources, respectively.
[0064] Signal decoding module: for decoding the communication signal calculated by the PE module.
[0065] Neural network preprocessing module: for preprocessing the decoded communication information to adapt to the subsequent AI calculation.
[0066] AI result output module: for outputting the AI calculation result of the PE module.
[0067] The working method of the above baseband processor includes two stages:
[0068] The first stage is the communication calculation stage, which mainly receives external signals for signal processing. The processing includes at least one of encoding, filtering, modulation, and MIMO equalization; preferably, it includes filtering, modulation, and MIMO equalization. The main steps of the first stage are as follows:
[0069] 1. The signal input module accepts external communication signal input.
[0070] 2. The signal preprocessing module preprocesses the input signal, including encoding, modulation, filtering, etc.
[0071] 3. The preprocessed signal is sent to the dynamic precision adjustable PE module (i.e. PE array) for MIMO equalization. The correlation calculation of the communication is completed.
[0072] 4. The signal decoding module decodes the communication signal after MIMO equalization.
[0073] The second stage is mainly the AI calculation stage (i.e. neural network calculation), which mainly performs AI inference tasks on the signal after MIMO equalization. The main steps of the second stage are as follows:
[0074] 5. The decoded signal is sent to the neural network preprocessing module for preprocessing, preparing for the subsequent AI calculation.
[0075] 6. The neural network preprocessed signal is sent to the dynamic precision adjustable PE module (i.e. PE array) for AI operation.
[0076] 7. The signal output module outputs the AI calculated signal.
[0077] During the process, the dynamic precision adjustable PE module is controlled by the computing resource scheduling module to allocate resources to MIMO equalization and AI calculation.
[0078] The application also provides an optimal computing resource scheduling algorithm, which targets at task timeliness, and can allocate computing resources according to communication quality, task demand, hardware demand and other constraints, so that the total calculation time is minimized, the calculation efficiency is highest, and the hardware cost is minimized.
[0079] A baseband processor resource scheduling algorithm, comprising the following steps:
[0080] S1, estimate the channel capacity model of the MIMO communication system, determine the baseband processor communication calculation time and resource model according to the channel capacity, and then determine the number of PE units allocated to MIMO equalization according to the communication calculation time and resource model.
[0081] S2, determine the AI calculation time and resource model, and then determine the number of PE units for AI calculation according to the number of PE units allocated by MIMO equalization.
[0082] Further, the channel capacity model of the MIMO communication system is as follows,
[0083] has N t transmit antennas and N r receive antennas, N t ≤N r , and the received symbol vector y is expressed as follows:
[0084] y = Hx + n (1)
[0085] H is a N r ×N t complex channel matrix, x is a column vector containing complex modulation signals, n is additive white Gaussian noise with variance σ 2 = N0, the modulation symbols belong to a set of size 2 M , and M is the number of bits per modulation symbol.
[0086] For linear detectors, the MIMO channel is treated as a linear system. The estimate of the transmitted symbol vector x is achieved by directly multiplying the received symbol vector y with an estimate matrix G to counteract the effect of the channel. The estimate matrix G is usually based on the Minimum Mean Square Error (MMSE) criterion and can be computed by a large number of matrix operations involving intensive data path computation. In Linear Minimum Mean-Square Error Detection (LMMSE), the detector minimizes the mean square error of the equalized estimated symbols, defined as:
[0087]
[0088] The computation formula for estimating the transmitted symbol vector x using the LMMSE equalization method is:
[0089]
[0090] In the equalization stage, the N t ×N r matrix (H H H+N0I) needs to be inverted, and then multiplied twice with H H and the received symbol vector y.
[0091] In actual working environments, the computational complexity of a communication system is limited due to limited hardware resources. Therefore, to implement an efficient communication system, it is necessary to optimize the computational complexity of the communication system by quantizing the channel matrix while ensuring the quality of communication. The quantization of the channel matrix can be represented as:
[0092] H≈H q = Q(H) (4)
[0093] Q(·) is a quantization function, and quantization greatly reduces the complexity of the system. By quantizing an N-bit channel matrix to an N-M-bit matrix, the complexity of the multiplier is reduced to N(N-M). Without loss of generality, the quantization function is equivalent to the following equation based on approximate calculation:
[0094] H=Hq +E q (5)
[0095] H q is the quantization matrix, E q is the quantization error matrix.
[0096] For an ideal baseband MIMO communication system, the transmitted signal can be represented as equation (3), where the channel model H is given by R T R R and R T are the normalized receive and transmit correlation Hermitian matrices with unit diagonal elements, respectively. Due to the existence of lossless feedback link, i.e., CSIT and CSIR are the same, both the transmitter and receiver can obtain R ω R c . The elements of H H are independent and identically distributed as N 2 (0, 1). x represents the transmitted signal, which is given by:
[0097] x = Ws (6)
[0098] s is the transmitted symbol, and W is the precoding matrix. The precoding matrix is given by:
[0099] W = (H -1 H + σ H I) rev (7)
[0100] I is the identity matrix. The demodulated signal received at the receiver is given by:
[0101] s
[0102] For the channel matrix H, there exists precision error caused by quantization truncation in practical hardware implementation. For the convenience of theoretical analysis, these precision errors are equivalently represented as rounding noise equation. Therefore, the relationship between the quantized channel matrix H q with quantization error and the ideal channel matrix H can be represented as:
[0103]
[0104] where E ω is the rounding noise matrix. The rounding noise matrix E ω is a random matrix with mean μ q and variance , satisfying the probability density function as follows:
[0105]
[0106] The rounding noise Eω Also, the following is satisfied:
[0107]
[0108] Q represents the quantization bit-width of the elements.
[0109] From which the mean μ of the rounding noise E ω is derived as: q
[0110]
[0111] The variance σ2 of the rounding noise E ω is:
[0112]
[0113] where Ω is the length of the quantization interval, equal to 2 -(Q-1) Based on the channel quantization error model established by equation (1) and equation (5), the output of the channel can be written as:
[0114] y = H q x + E q x + n (14)
[0115] Therefore, the total noise of the system is represented as n total = E q x + n, with mean 0 and covariance matrix:
[0116]
[0117] In the equation: the expectation is related to x, n, E ω The elements of the matrix E ω are independent and identically distributed We use the following lemma:
[0118]
[0119] The mutual information between the received signal y, the transmitted signal x and the quantization channel H q can be constrained as:
[0120]
[0121] Here, Q = E(xx H ), and the upper bound of the mutual information can be expressed as
[0122]
[0123] Let be the eigenvalues of H , then Hq the singular values of According to the lemma, we can transform the lower bound expression of mutual information into
[0124]
[0125] κ1and ρ are coefficient factors, Q is a covariance matrix, is the singular value of the matrix (·), N r is the number of columns of the channel matrix (N r is the number of receive antennas, calculated as the number of columns of the channel matrix), Q represents the quantization bit width of the calculated elements, is the average singular value of the matrix (·), is the average eigenvalue of the matrix (·). The eigenvalue of the total noise covariance matrix is
[0126]
[0127] ρ is a coefficient factor, ρ = tr(R T E(xx H )).
[0128] The upper bound of mutual information can be expressed as:
[0129]
[0130] For Gaussian input, the gap between the lower bound and the upper bound is usually small, so the channel capacity of the system is represented by the following formula
[0131]
[0132] Substitute equation (20) into equation (23), the channel capacity of the system can be expressed as:
[0133]
[0134] is the average singular value of the matrix H q , and are the average eigenvalues of the matrices Q and R R , respectively.
[0135] κ1and ρ are coefficient factors, Q is a covariance matrix, is the singular value of the matrix (·), N r is the number of columns of the channel matrix, Q represents the quantization bit width of the calculated elements, is the average singular value of the matrix (·), is the average eigenvalue of the matrix (·).
[0136] 2.1 Communication computation time and resource model
[0137] The channel capacity of the system can be expressed as:
[0138]
[0139] The average singular value of the matrix H q , denoted as and are the average eigenvalues of the matrices Q and R R , respectively.
[0140] For a communication channel , the matrix operation of its conjugate matrix H H with itself H requires multiplication operations and addition operations. H The resources required for matrix multiplication H
[0141]
[0142] ξ(Q D ) represents the resource overhead consumed by a single PE unit. Q D is the bit width of the input data.
[0143] In the pulse array implementation of matrix multiplication, each PE unit needs to at least contain a complex multiplier and an arithmetic logic unit (ALU). The total number of PEs required for the communication process is:
[0144]
[0145] T cycle is the number of clock cycles. The computation time T cycle is constrained by the data volume and the channel capacity, which is specified as follows:
[0146]
[0147] D is the data volume of the input communication, and C is the channel capacity.
[0148] 2.2 AI computation time and resource model
[0149] Neural networks can be mapped to specific hardware architectures through data flow graphs, in which each node represents a neural network operation and each edge represents the direction of data flow in the neural network inference process. The inference time of neural network hardware can be expressed as:
[0150]
[0151] T represents the clock cycle, d represents the pipeline depth, and B is the batch size of the input data.
[0152] Γ is the topology matrix of the dataflow graph, where each column represents a node in the dataflow graph, each row represents an edge in the dataflow graph, and each element represents how much data the current node consumes or produces on the current edge, subject to the number of PEs in the hardware architecture. Each node in the dataflow graph represents an operation in the neural network, i.e., a layer, and the edges in the dataflow graph represent the connection order of different layers in the neural network.
[0153] W is the workload vector matrix, where the i-th element represents the dimension of the feature matrix of the hardware module corresponding to the i-th node in the input dataflow graph, determined by the data size input to the neural network and the structure of the neural network.
[0154] By element-wise division and rounding, the resulting matrix represents the number of clock cycles required for each element in the hardware completion pipeline to complete its work. The maximum value between the pipeline depth and the loop count represents the total number of cycles required to complete one prediction. Multiply the total number of cycles by the clock cycle length to get the total inference time.
[0155] To minimize the inference time, we need to determine the optimal workload vector W and topology matrix Γ.
[0156] For a specific neural network structure, given the dimension of the input data, the workload matrix can be completely determined, and the workload matrix can be directly derived
[0157]
[0158] For a specific neural network structure, the overall structure of the dataflow graph can be directly determined. Therefore, the dimension of the topology matrix Γ can be completely determined. The variable is the number of PEs allocated to each layer (convolution, pooling, and nonlinear activation).
[0159] To simplify the representation, the edges in the original matrix that are not related to each node are omitted. The original Γ can be represented as
[0160]
[0161] P j The number of PEs required for the j-th layer of the neural network is represented as
[0162]
[0163] The summation of the i-th part represents the number of clock cycles corresponding to the number of layers of the neural network included in the i-th stage of the pipeline.
[0164] For a certain neural network, the optimal pipeline partition strategy can be searched by heuristic algorithm to solve the above equation, and then The maximum element of the matrix can be expressed as
[0165]
[0166] The value of a depends on the specific neural network structure and hardware architecture.
[0167] D represents the number of elements in the neural network input data feature matrix, where W i represents the optimal i-th pipeline stage number.
[0168] P nn represents the number of PEs applied to the neural network. Then, the prediction time of the neural network can be calculated as
[0169]
[0170] The total time of the entire system to complete the communication and neural network calculation task can be expressed as
[0171]
[0172] C represents the channel capacity, which is a function of noise and quantization bit width Q, given by the previous channel capacity model, k represents a coefficient satisfying k < 1, P represents the number of PEs in the entire system, and kP represents the number of PEs allocated to neural network calculation.
[0173] The minimum time model of communication and neural network calculation is
[0174]
[0175] ε and ξ(Q) have been given, since the PE units are allocated to both neural network calculation and communication tasks, kP represents the number of PE units allocated to neural network calculation, and then (1-k)P represents the number of PEs allocated to communication. The constraint ensures that the number of PEs allocated to communication meets the computational resource requirement of the communication part.
[0176] In order to minimize the total time, the values of Q and k can be adjusted. k can have an upper bound
[0177]
[0178] The minimum value of the total time is
[0179]
[0180] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with test examples and specific embodiments. However, it should not be understood that the scope of the above-mentioned subject matter of the present application is limited to the following examples only, and any technology realized based on the content of the present application falls within the scope of the present application.
[0181] Embodiment 1
[0182] As Figure 1 shown, the embodiment of the present application provides a baseband processor based on general fusion, which comprises a signal input module, a signal preprocessing module, a PE resource scheduling module, a dynamic precision adjustable PE module, a signal decoding module, a neural network preprocessing module and an AI result output module.
[0183] The signal input module is configured to receive a communication signal.
[0184] Specifically, the signal input module is connected with the signal preprocessing module and transmits the received communication signal to the signal preprocessing module.
[0185] The signal preprocessing module performs signal preprocessing on the communication signal and transmits the communication signal after the signal preprocessing to the dynamic precision adjustable PE module in a MIMO transmission mode.
[0186] Specifically, the signal preprocessing includes encoding, modulation and filtering.
[0187] Specifically, the MIMO transmission mode is a MIMO transmission system with N t transmit antennas and N r receive antennas and (N t ≤ N r ); N t , N r are positive integers, N t represents the number of transmit antennas, and N r represents the number of receive antennas.
[0188] Specifically, the signal preprocessing module is connected with the PE resource scheduling module and transmits the communication signal after the signal preprocessing to the PE resource scheduling module.
[0189] The dynamic precision adjustable PE module realizes dynamic allocation of PE units and is used for MIMO equalization or neural network calculation; as Figure 2 shown, the dynamic precision adjustable PE module comprises a control module, a storage module and a calculation unit; the calculation unit comprises a plurality of PE units; the MIMO equalization and the neural network calculation share the PE units.
[0190] Specifically, the MIMO equalization is used for communication optimization of the communication signal after the signal preprocessing; the MIMO equalization equalizes the communication signal after the signal preprocessing by using a linear minimum mean square error algorithm; the neural network calculation calculates the information content after the data normalization to obtain a digital sequence; the MIMO equalization and the neural network calculation have a common calculation paradigm, and the PE units can be time-multiplexed; and the communication optimization mainly aims to reduce the inter-symbol interference.
[0191] Specifically, the MIMO equalization and the neural network calculation have a common calculation paradigm, and the PE units can be time-multiplexed, including that the MIMO equalization is essentially a matrix multiplication operation, the neural network calculation is also essentially a matrix multiplication operation, the MIMO equalization and the neural network calculation have a common calculation paradigm, i.e., a matrix multiplication operation, and the MIMO equalization and the neural network calculation can both use PE units to perform the matrix multiplication operation.
[0192] Specifically, the PE unit at least includes a complex multiplier and an arithmetic logic unit, and the calculation precision of the multiplication calculation and the addition calculation can be adjusted when the PE unit performs the MIMO equalization and the neural network calculation.
[0193] Specifically, as shown in Figure 2 The calculation precision of the multiplication calculation and the addition calculation can be adjusted, including that the precision of the multiplication calculation is dynamic, 2 bits and 2 bits are multiplied to be 4 bits, 4 bits and 4 bits are multiplied to be 8 bits, and so on, 8 bits and 8 bits are multiplied to be 16 bits, and the precision of the addition calculation is also dynamic. The addition may overflow, so two 16-bit additions are temporarily stored by using 20 bits, so the PE unit uses 20 bits to save the intermediate results when performing partial sum accumulation, and then 16 bits are cut back, and the calculation precision adjustment can save the overhead in the case of not requiring high-precision calculation. Figure 2 The control module is connected with the storage module, indicating that the control module can transmit a control signal to the storage module and control the resource allocation of the PE, and the two storage modules store the elements of two matrix multiplications to be performed, respectively. The storage module is connected with the PE array, indicating that the storage module transmits data to the specified PE unit.
[0194] Specifically, the dynamic precision adjustable PE module is connected with the signal decoding module, and the communication signal after the MIMO equalization is transmitted to the signal decoding module.
[0195] Specifically, the dynamic precision adjustable PE module is connected with the neural network output module, and the calculation result after the neural network calculation is transmitted to the neural network output module.
[0196] The signal decoding module decodes the communication signal completed by the MIMO equalization.
[0197] Specifically, the signal decoding module is connected with the neural network preprocessing module, and the communication signal completed by the decoding operation is transmitted to the neural network preprocessing module.
[0198] The neural network preprocessing module performs data normalization processing on the information content carried by the communication signal completed by the decoding operation, and transmits the information content completed by the data normalization processing to the dynamic precision adjustable module.
[0199] Specifically, the neural network preprocessing module is connected with the dynamic precision adjustable PE module, and the information content completed by the data normalization processing is transmitted to the dynamic precision adjustable PE module.
[0200] The AI result output module receives the digital sequence obtained by the neural network calculation, restores the digital sequence into corresponding information content, and displays the information content.
[0201] The PE resource scheduling module is configured to receive the communication signal completed by the signal preprocessing, obtain the number of PE units allocated to the MIMO equalization and the number of PE units allocated to the neural network calculation according to an optimal scheduling algorithm, and transmit the number of PE units allocated to the MIMO equalization and the number of PE units allocated to the neural network calculation to the dynamic precision adjustable PE module.
[0202] Specifically, the number of PE units allocated to the MIMO equalization and the number of PE units allocated to the neural network calculation are used for dynamic allocation of PE units to the dynamic precision adjustable PE module.
[0203] Specifically, the number of PE units allocated to the MIMO equalization and the number of PE units allocated to the neural network calculation are used for dynamic allocation of PE units to the dynamic precision adjustable PE module, including that the dynamic precision adjustable PE module receives the number of PE units allocated to the MIMO equalization and the number of PE units allocated to the neural network calculation; and a control module of the dynamic precision adjustable PE module performs dynamic allocation of PE units according to the number of PE units allocated to the MIMO equalization and the number of PE units allocated to the neural network calculation.
[0204] Specifically, the control module of the dynamic precision adjustable PE module performs dynamic allocation of PE units according to the number of PE units allocated to MIMO equalization and the number of PE units allocated to neural network calculation, including: the control module allocates the same number of PE units as the number of PE units allocated to MIMO equalization or the number of PE units allocated to neural network calculation for MIMO equalization or neural network calculation.
[0205] Specifically, the number of PE units allocated to MIMO equalization and the number of PE units allocated to neural network calculation according to the optimal scheduling algorithm includes: calculating the channel capacity under MIMO transmission mode; calculating the time required for MIMO equalization according to the channel capacity; and obtaining the number of PE units allocated to MIMO equalization and the number of PE units allocated to neural network calculation according to the time required for MIMO equalization and the resources required for MIMO equalization.
[0206] Specifically, the calculation of the channel capacity under MIMO transmission mode includes:
[0207] For a MIMO transmission system with N t transmit antennas and Nr receive antennas and (N t ≤N r ), the received symbol vector y can be represented as:
[0208] y=Hx+n (1)
[0209] where y is the received symbol vector, H is the channel matrix of N r ×N t , x is the transmitted symbol vector, and n is the additive white Gaussian noise with variance The modulation symbol belongs to a set of size 2 M , and M is the number of bits per modulation symbol.
[0210] In the actual working environment, due to limited hardware resources, the computational complexity of the communication system is limited. Therefore, to realize an efficient communication system, it is necessary to optimize the computational complexity of the communication system by quantizing the channel matrix while ensuring the communication quality.
[0211] The quantization of the channel matrix can be represented as:
[0212] H=H q +E q (2)
[0213] where H q is the quantized matrix, and E q is the quantization error matrix.
[0214] The quantized matrix H qand quantization error matrix E q may be expressed as:
[0215]
[0216] where R R is a unit-diagonal normalized receive Hermitian matrix, R T is a unit-diagonal normalized transmit Hermitian matrix, H w is a channel precoding matrix, E w is a rounding noise matrix.
[0217] where rounding noise matrix E w is a random matrix with mean μ q and variance , since it satisfies the probability density function as follows:
[0218]
[0219] and rounding noise matrix E w also satisfies:
[0220]
[0221] where Q d is the quantization bit-width of the computational elements, which are the elements of the channel matrix.
[0222] Then the mean μ w of rounding noise matrix E q is:
[0223]
[0224] The variance of rounding noise matrix E w is:
[0225]
[0226] where Ω is the length of the quantization interval, which is equal to
[0227] From (1) and (2), the received symbol vector y can be expressed as:
[0228] y = H q x + E q x + n
[0229] Then the total noise can be expressed as:
[0230] n total = E q x + n
[0231] The covariance matrix of the total noise can be expressed as:
[0232]
[0233] R T is the covariance matrix of the unit diagonal normalized transmit Hermitian matrix, R R is the covariance matrix of the unit diagonal normalized receive Hermitian matrix, is an N r ×N r unit matrix.
[0234] The eigenvalues of the covariance matrix of the total noise are:
[0235]
[0236] ρ is a coefficient factor, ρ = tr(R T (xx H ).
[0237] The mutual information between the received symbol vector y, the transmitted symbol vector x and the quantization matrix H q can be constrained as:
[0238]
[0239] σ i (·) denotes the eigenvalues of the matrix (·), H q denotes the quantization matrix, is the eigenvalues of the covariance matrix Q, is the eigenvalues of the total noise covariance matrix.
[0240] For Gaussian input, the gap between the lower bound and the upper bound is usually small, and thus the channel capacity can be expressed as:
[0241]
[0242] The eigenvalues of the total noise covariance matrix are brought into the equation to obtain:
[0243]
[0244] The calculation formula of the channel capacity is as follows:
[0245]
[0246] wherein C is the channel capacity; k1 and ρ are coefficient factors; N r is the number of rows of the channel matrix, i.e. the number of receive antennas; is the average singular value of the quantization error matrix H q . is the average eigenvalue of the covariance matrix Q; Q d is the quantization bit width of the calculation element; R R is the unit diagonal element normalized receiving Hermit matrix; is the variance of the additive white Gaussian noise.
[0247] Specifically, according to the channel capacity, the time required for MIMO equalization is calculated, including:
[0248] The time required for MIMO equalization is related to the amount of data input for communication and the channel capacity, and the specific calculation formula is as follows:
[0249]
[0250] wherein T cycle is the time required for MIMO equalization; D is the amount of data input for communication; C is the channel capacity.
[0251] Specifically, according to the time required for MIMO equalization and the resources required for MIMO equalization, the number of PE units allocated to MIMO equalization and the number of PE units allocated to neural network calculation are obtained, including:
[0252] Suppose that the channel matrix H obtained under the MIMO transmission mode, wherein the channel matrix H is a matrix with N r rows and N t columns, the channel matrix H and its own conjugate matrix H H perform matrix multiplication operation needs to make times multiplication operation and times addition operation, the calculation formula of the resources required for MIMO equalization is as follows:
[0253]
[0254] wherein ε is the resources required for MIMO equalization, N t is the number of columns of the channel matrix H, N r is the number of rows of the channel matrix, ξ(Q d ) is the resource consumption of a single PE unit when performing MIMO equalization when the quantization bit width of the calculation element is Q d ; the calculation element is all elements for MIMO equalization.
[0255] The calculation formula of the number of PE units allocated to MIMO equalization is as follows:
[0256]
[0257] wherein P comN is the number of PE units allocated to MIMO equalization; ε is the resource required by MIMO equalization; N t T is the number of rows of the channel matrix, i.e., the number of transmit antennas; T cycle is the time required for MIMO equalization; ξ(Q) is the resource overhead consumed by a single PE unit when performing MIMO equalization when the actual bit width of the calculation element is Q; the calculation element is all elements performing MIMO equalization.
[0258] The number of PE units allocated to neural network calculation is subject to the number of PE units required by MIMO equalization, because the number of PE units allocated to MIMO equalization must meet the demand of the number of PE units required by MIMO equalization, so the calculation formula of the number of PE units allocated to neural network calculation is as follows:
[0259] P nnc = P - P com
[0260] Where, P nnc is the number of PE units allocated to neural network calculation, P is the total number of PE units; P com is the number of PE units required by MIMO equalization.
[0261] As Figure 3 is the working timing diagram of the baseband processor based on the overall fusion, the communication time and the neural network prediction time are performed in sequence. By optimizing the number of PE units allocated to MIMO equalization and neural network algorithm calculation, the overall time is minimized, the resource allocation of the two is optimized, and the best efficiency is achieved.
Claims
1. A baseband processor based on universal computing and integration, characterized by: include: A signal input module, used for receiving communication signals; A signal preprocessing module performs signal preprocessing on the communication signal received by the signal input module, and then transmits the preprocessed communication signal to the dynamic precision adjustable PE module; Computing resource scheduling module, used to schedule dynamic precision adjustable PE modules; The computing resource scheduling module can control the dynamic precision adjustable PE module to allocate PE unit resources and complete the calculation according to the communication quality, hardware constraints and task requirements; The formula for calculating the number of PE units allocated to MIMO equalization is as follows: Among them, P com is the number of PE units allocated to MIMO equalization; is the resource required for MIMO equalization; Nt is the number of columns in the channel matrix, i.e. the number of transmitting antennas; T cycle The time required for MIMO equalization; It is the resource overhead consumed by a single PE unit when performing MIMO equalization when the actual bit width of the calculation element is Q; Dynamic precision adjustable PE module for MIMO equalization and / or AI calculation; The dynamic precision adjustable PE module includes: a control unit, a storage unit and a calculation unit, wherein the calculation unit includes a plurality of PE units; the control unit controls the calculation unit to read the data of the storage unit for calculation processing; A signal decoding module is used to receive the communication signal after MIMO equalization by the dynamic precision adjustable PE module and perform decoding operations; A neural network preprocessing module is used to preprocess the information content obtained by the decoding operation to obtain a signal for neural network calculation; then, the signal is input into the dynamic precision adjustable PE module for AI calculation to obtain an AI calculation result; The AI result output module is used to output the AI calculation results.
2. The baseband processor based on universal computing and integration according to claim 1, characterized in that: The MIMO equalization is as follows: the dynamic precision adjustable PE module equalizes the communication signal after signal preprocessing using a linear minimum mean square error algorithm; The AI calculation is to perform neural network calculation on the information content of the completed data normalization process to obtain a digital sequence.
3. The baseband processor based on universal computing and integration according to claim 1, characterized in that: The computing resource scheduling module can control the number of PE units allocated by the dynamic precision adjustable PE module to MIMO equalization and the number of PE units allocated to AI calculation.
4. The baseband processor based on universal computing and integration according to claim 1, characterized in that: The PE unit includes at least a complex multiplier and an arithmetic logic unit; when the PE unit performs MIMO equalization and AI calculation, the calculation accuracy can be adjusted.
5. A method for performing integrated computing and control using the baseband processor according to any one of claims 1 to 4, characterized in that: Including a communication computing stage and / or an AI computing stage; Communication and computing phase: receiving communication signal input; Preprocess the input signal; The pre-processed signal is sent to the dynamic precision adjustable PE module for MIMO equalization; Decoding the communication signal after MIMO equalization; AI computing stage: The decoded signal is sent to the neural network preprocessing module for preprocessing; The signal after neural network preprocessing is sent to the dynamic precision adjustable PE module for AI calculation; Output the results of AI calculation.
6. The method for performing integrated computing and control of a baseband processor according to claim 5, wherein: Calculate the channel capacity of the MIMO communication system, determine the time required to calculate MIMO equalization based on the channel capacity, and determine the number of PE units allocated to MIMO equalization.
7. The method for performing integrated computing and control of a baseband processor according to claim 6, wherein: The calculation formula of the channel capacity is as follows: Where C is the channel capacity; and is the coefficient factor; is the number of rows of the channel matrix, i.e. the number of receiving antennas; is the quantization error matrix The average singular value of is the average eigenvalue of the covariance matrix Q; Qd is the quantization bit width of the calculation element; RR is the unit diagonal element normalized receiving Hermit matrix; is the variance of additive white Gaussian noise; The calculation formula for the time required for MIMO equalization is as follows: Among them, T cycle is the time required for MIMO equalization; D is the amount of data input for communication; and C is the channel capacity.
8. The method for performing integrated computing and control of a baseband processor according to claim 7, wherein: The number of PE units allocated to AI calculation is determined based on the number of PE units allocated to MIMO equalization.
9. The method for performing integrated computing and control of a baseband processor according to claim 8, wherein: The calculation formula for the number of PE units allocated to AI calculation is as follows: Among them, P nnc is the number of PE units allocated to AI calculation, P is the total number of PE units; P com is the number of PE units allocated to MIMO equalization.
Citation Information
Patent Citations
Baseband processor based on baseband reconstruction and processing share and communication method thereof
CN101860496A
Processing communications signals using a machine-learning network
US20200343985A1