Linear pre-correction method, device, equipment, medium and program product

By using a coarse-grained reconstructible computing architecture to perform LU decomposition and orthogonalization of the basis function matrix in large bandwidth or large-scale MIMO scenarios, the problem of insufficient real-time performance of linearized pre-correction in the prior art is solved, and a more efficient linearized pre-correction effect is achieved.

CN120567239APending Publication Date: 2025-08-29CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410220268.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing digital predistortion algorithms are difficult to adapt to the rapidly changing amplifier characteristics in large bandwidth or large-scale MIMO scenarios, resulting in insufficient real-time performance of linearized pre-correction. The existing closed-loop adjustment mechanism cannot effectively ensure the real-time performance of linearized pre-correction in this scenario.

Method used

The basis function matrix is ​​decomposed by a reconstructible calculation architecture based on coarse-grainedness. By obtaining the cross-correlation matrix of the basis function matrix and judging its orthogonalization processing conditions, the decomposition matrix is ​​used for orthogonalization processing, and the correction parameters are obtained using the minimum root mean square algorithm to improve the orthogonalization processing efficiency of the basis function matrix.

Benefits of technology

In large bandwidth or large-scale MIMO scenarios, real-time performance improvement of linear pre-correction is achieved, reducing scenario limitations and improving the processing efficiency of the basis function matrix.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567239A_ABST
    Figure CN120567239A_ABST
Patent Text Reader

Abstract

The invention discloses a linear pre-correction method, device and equipment, a medium and a program product, and the method comprises the steps: obtaining a cross-correlation matrix of a primary function matrix based on the primary function matrix of a current input signal, and judging whether the primary function matrix meets a preset orthogonalization processing condition or not; when the primary function matrix meets the orthogonalization processing condition, performing LU decomposition on the primary function matrix based on a coarse-grained reconfigurable computing architecture to obtain a decomposition matrix; performing orthogonalization processing on the primary function matrix by using the decomposition matrix to obtain a primary function orthogonal matrix; and on the basis of the primary function orthogonal matrix, a correction parameter used for linearization pre-correction is obtained by using a minimum root mean square algorithm. According to the method, the processing efficiency of orthogonalization of the primary function matrix can be improved by utilizing the coarse-grained reconfigurable computing architecture, so that the real-time performance of linear pre-correction in a large-bandwidth or large-scale MIMO scene can be ensured, and the scene limitation is relatively small.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of signal processing technology, and in particular to a linearization pre-correction method, apparatus, terminal equipment, computer-readable storage medium, and computer program product. Background Art

[0002] In wide-bandwidth or massive MIMO scenarios, current digital pre-distortion algorithms are difficult to apply to rapidly changing power amplifier characteristics due to their non-real-time learning nature, necessitating linearization pre-correction. Existing technologies typically address the issue of correcting power amplifier nonlinearities through closed-loop regulation, but this approach cannot guarantee real-time performance in these scenarios, significantly limiting their use. Summary of the Invention

[0003] The present invention provides a linearization pre-correction method, apparatus, device, medium and program product, which performs LU decomposition on the basis function matrix of the current input signal based on a coarse-grained reconfigurable computing architecture, so as to use the decomposition matrix to orthogonalize the basis function matrix, thereby improving the processing efficiency of the orthogonalization of the basis function matrix, thereby ensuring the real-time performance of linearization pre-correction in large bandwidth or large-scale MIMO scenarios, and having fewer scenario limitations.

[0004] In order to solve the above technical problems, a first aspect of an embodiment of the present invention provides a linearization pre-correction method, including:

[0005] Based on the basis function matrix of the current input signal, obtaining a cross-correlation matrix of the basis function matrix and determining whether the basis function matrix meets a preset orthogonal processing condition;

[0006] When the basis function matrix satisfies the orthogonalization processing condition, performing LU decomposition on the basis function matrix based on a coarse-grained reconfigurable computing architecture to obtain a decomposition matrix;

[0007] orthogonalizing the basis function matrix using the decomposition matrix to obtain a basis function orthogonal matrix;

[0008] Based on the basis function orthogonal matrix, a least mean square algorithm is used to obtain correction parameters for linearization pre-correction.

[0009] As a preferred solution, performing LU decomposition on the basis function matrix based on the coarse-grained reconfigurable computing architecture to obtain a decomposition matrix specifically includes:

[0010] Based on the LU decomposition algorithm, determining a set of multiplication operations of several groups of vectors and scalars of the basis function matrix in the LU decomposition process;

[0011] Mapping the processing values ​​of several groups of the multiplication operation sets to the processing unit array in the coarse-grained reconfigurable computing architecture and executing the LU decomposition algorithm to obtain the decomposition matrix; wherein the processing values ​​include vector values ​​and scalar values.

[0012] As a preferred solution, mapping the processing values ​​of the plurality of groups of the multiplication operation sets to a processing unit array in a coarse-grained reconfigurable computing architecture and executing the LU decomposition algorithm to obtain the decomposition matrix specifically includes steps S21 to S26:

[0013] Step S21, determining whether, among the scalar values ​​corresponding to the kth row in the current basis function matrix, there are any two scalar values ​​corresponding to a vector in a multiplication operation set whose length is less than a preset vector length threshold; wherein k is an integer whose initial value is 1; and the vector length threshold is set based on the number of processing units in the processing unit array;

[0014] Step S22: If present, mapping the processing values ​​of the multiplication operation set corresponding to the at least two scalar values ​​to the processing unit array based on a preset parallel mapping strategy, so as to execute the currently mapped multiple sets of multiplication operations in parallel; wherein the multiplication operation set includes multiple multiplication operations between any scalar value corresponding to the kth row and multiple vector values ​​corresponding to the kth column;

[0015] When the execution of the sets of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column is obtained according to the execution results of the elements of the current i-th row and j-th column in the basis function matrix and the multiplication operations of the current mapping; wherein i is the row corresponding to the vector value in the multiplication operation of the current mapping, and its value is k+1 to N; j is the column corresponding to the scalar value in the multiplication operation of the current mapping, and its value is k+1 to N; N represents the length of the basis function matrix;

[0016] Step S23: If not, mapping the processing value of the multiplication operation set corresponding to any scalar value to the processing unit array based on a preset serial mapping strategy to execute the currently mapped set of multiplication operations;

[0017] When the set of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column is obtained according to the execution results of the element of the current i-th row and j-th column in the basis function matrix and the multiplication operations of the current mapping;

[0018] Step S24, repeatedly executing steps S21 to S23 until all multiplication operation sets corresponding to the multiple scalar values ​​corresponding to the current k-th row are completed;

[0019] Step S25, increase the value of k by one, and re-execute steps S21 to S24 until the value of k is equal to N;

[0020] Step S26: Obtain the decomposition matrix according to the current updated elements.

[0021] As a preferred solution, the processing unit array includes a plurality of processing units, each of which supports interconnection access with processing units within a preset relative unit distance; then, the method maps the processing values ​​of the multiplication operation set corresponding to at least two scalar values ​​to the processing unit array based on a preset parallel mapping strategy, specifically comprising:

[0022] Selecting a first processing unit and a second processing unit from a plurality of the processing units in the processing unit array;

[0023] Mapping any two scalar values ​​to the first processing unit and the second processing unit respectively;

[0024] Based on a plurality of vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values, mapping processing is performed on a plurality of to-be-mapped processing units that are within the preset relative unit distance from the first processing unit or the second processing unit;

[0025] When the current number of processing units to be mapped is greater than or equal to L+1, the new scalar value corresponding to the current k-th row is mapped to any processing unit to be mapped, and based on the several vector values ​​in the multiplication operation set corresponding to the new scalar value, mapping processing is performed on several processing units to be mapped that are within the preset relative unit distance from the any processing unit to be mapped; repeat this step until the current number of processing units to be mapped is less than L+1; where L represents the length of the vector in the multiplication operation set; and the new scalar value is a scalar value that is not currently mapped.

[0026] As a preferred solution, selecting the first processing unit and the second processing unit from a plurality of the processing units in the processing unit array specifically includes:

[0027] Select any one processing unit from the processing unit array as the first processing unit;

[0028] According to a selection strategy with the least shared processing units, the second processing unit is selected from the remaining processing units in the processing unit array except the first processing unit; wherein the shared processing unit is a processing unit whose relative distance from the first processing unit and the second processing unit is within the preset relative unit distance.

[0029] As a preferred solution, the mapping of the plurality of to-be-mapped processing units within the preset relative unit distance from the first processing unit or the second processing unit based on the plurality of vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values ​​specifically includes:

[0030] Based on several vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values, mapping processing is performed on several to-be-mapped processing units that are within the preset relative unit distance from the first processing unit or the second processing unit, and the total distance between the remaining to-be-mapped processing units after the mapping processing is the smallest.

[0031] As a preferred solution, the processing unit array includes a plurality of processing units, each of which supports interconnection access with processing units within a preset relative unit distance; then, mapping the processing value of the multiplication operation set corresponding to any scalar value to the processing unit array based on a preset serial mapping strategy specifically includes:

[0032] Select any one processing unit from the processing unit array as the third processing unit;

[0033] Mapping any scalar value to the third processing unit;

[0034] Based on a number of vector values ​​in the multiplication operation set corresponding to the arbitrary scalar value, mapping processing is performed on a number of to-be-mapped processing units that are within the preset relative unit distance from the third processing unit.

[0035] As a preferred solution, obtaining the updated element in the current i-th row and j-th column according to the execution results of each multiplication operation of the element in the current i-th row and j-th column in the basis function matrix and the current mapping specifically includes:

[0036] According to the execution results of the multiplication operations of the element in the current i-th row and j-th column in the basis function matrix and the current mapping, the updated element in the current i-th row and j-th column is obtained by the following expression:

[0037] A ′ [i][j]=A[i][j]-A[i][k]×A[k][j];

[0038] Wherein, A[i][j] represents the element in the current i-th row and j-th column of the basis function matrix; A[i][k] represents the element in the current i-th row and k-th column of the basis function matrix, which is any vector value corresponding to the k-th column; A[k][j] represents the element in the current k-th row and j-th column of the basis function matrix, which is any scalar value corresponding to the k-th row; the product of A[i][k] and A[k][j] is the execution result of the multiplication operation; A′ [i][j] represents the updated element in the current i-th row and j-th column.

[0039] As a preferred solution, the method specifically performs the multiplication operation set of the current mapping through the following steps:

[0040] The processing unit mapped with the vector value accesses the processing unit mapped with the scalar value belonging to the same group of multiplication operation sets as the vector value, and performs the multiplication operation between the vector value and the accessed scalar value to complete the execution of the currently mapped multiplication operation set.

[0041] As a preferred solution, the method specifically determines whether the basis function matrix meets the preset orthogonal processing conditions through the following steps:

[0042] Obtaining the maximum eigenvalue and the minimum eigenvalue of the cross-correlation matrix, and calculating the ratio between the maximum eigenvalue and the minimum eigenvalue;

[0043] When the absolute value of the ratio is greater than a preset ratio threshold, it is determined that the basis function matrix meets the orthogonalization processing condition.

[0044] As a preferred solution, the processing unit array includes 16 processing units, and the 16 processing units form a 4×4 processing unit array.

[0045] A second aspect of an embodiment of the present invention provides a linearization pre-correction device, including:

[0046] An orthogonalization processing condition judgment module is used to obtain a cross-correlation matrix of the basis function matrix based on the current input signal and judge whether the basis function matrix meets the preset orthogonalization processing condition;

[0047] a basis function matrix decomposition module, configured to perform LU decomposition on the basis function matrix based on a coarse-grained reconfigurable computing architecture to obtain a decomposition matrix when the basis function matrix satisfies the orthogonalization processing condition;

[0048] An orthogonalization processing module, configured to perform an orthogonalization process on the basis function matrix using the decomposition matrix to obtain a basis function orthogonal matrix;

[0049] The correction parameter acquisition module is used to obtain correction parameters for linearization pre-correction based on the basis function orthogonal matrix using a least mean square algorithm.

[0050] A third aspect of an embodiment of the present invention provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the linearization pre-correction method described in any one of the first aspects when executing the computer program.

[0051] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the linearization pre-correction method described in any one of the first aspects.

[0052] A fifth aspect of an embodiment of the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the linearization pre-correction method described in any one of the first aspects.

[0053] Compared with the existing technology, the beneficial effect of the embodiments of the present invention is that the basis function matrix of the current input signal is LU decomposed based on the coarse-grained reconfigurable computing architecture, and the basis function matrix is ​​orthogonalized using the decomposition matrix, which can improve the processing efficiency of the orthogonalization of the basis function matrix, thereby ensuring the real-time performance of linearization pre-correction in large bandwidth or large-scale MIMO scenarios, and the scenario limitations are relatively small. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 1 is a flow chart of a linearization pre-correction method according to an embodiment of the present invention;

[0055] Figure 2 is a schematic structural diagram of a processing unit array in an embodiment of the present invention;

[0056] Figure 3 Schematic diagram of mapping scalar values ​​in an embodiment of the present invention;

[0057] Figure 4 Schematic diagram of mapping vector values ​​in an embodiment of the present invention;

[0058] Figure 5 Schematic diagram of the calculation process of each update element in an embodiment of the present invention;

[0059] Figure 6 1 is a schematic diagram of the LU decomposition cycle flow in an embodiment of the present invention;

[0060] Figure 7 1 is a schematic structural diagram of a linearization pre-correction device according to an embodiment of the present invention;

[0061] Figure 8 It is a schematic structural diagram of a terminal device in an embodiment of the present invention. DETAILED DESCRIPTION

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0063] See also Figure 1 , Figure 1 1 is a flow chart of a linearization pre-correction method according to an embodiment of the present invention. A first aspect of an embodiment of the present invention provides a linearization pre-correction method, comprising the following steps S1 to S4:

[0064] Step S1, based on the basis function matrix of the current input signal, obtaining the cross-correlation matrix of the basis function matrix and determining whether the basis function matrix meets the preset orthogonal processing conditions;

[0065] Step S2: When the basis function matrix satisfies the orthogonalization processing condition, performing LU decomposition on the basis function matrix based on a coarse-grained reconfigurable computing architecture to obtain a decomposition matrix;

[0066] Step S3, performing orthogonal processing on the basis function matrix using the decomposition matrix to obtain a basis function orthogonal matrix;

[0067] Step S4: Based on the basis function orthogonal matrix, a least mean square algorithm is used to obtain correction parameters for linearization pre-correction.

[0068] Specifically, the basis function matrix of the current input signal can be expressed as follows:

[0069]

[0070] Among them, each matrix element is represented as a set of memory depth and polynomial order. The specific expression is as follows:

[0071] x[nm]|x[nm]| p-1 ;

[0072] Where m represents the memory depth and p represents the polynomial order.

[0073] Furthermore, the cross-correlation matrix of the basis function matrix is ​​calculated by the following expression:

[0074]

[0075] Among them, ψ represents the basis function matrix; ψ H represents the conjugate matrix of the basis function matrix; N represents the number of elements in the basis function matrix.

[0076] The cross-correlation matrix based on the basis function matrix can determine whether the basis function matrix has an instability problem during the inversion calculation process, and thus can determine whether the basis function matrix meets the preset orthogonal processing conditions.

[0077] Furthermore, when the basis function matrix meets the orthogonalization processing conditions, in order to improve the execution efficiency of the orthogonalization processing of the basis function matrix, this embodiment performs LU decomposition on the basis function matrix based on a coarse-grained reconfigurable computing architecture. The coarse-grained reconfigurable computing architecture has the advantages of flexible programming and high computing speed, which helps to improve the LU decomposition efficiency of the basis function matrix.

[0078] Furthermore, the decomposition matrix output by the LU decomposition is multiplied with the basis function matrix to complete the orthogonalization of the basis function matrix and obtain the basis function orthogonal matrix. Then, the least mean square algorithm is used to learn and obtain the correction parameters for linearization pre-correction. Specifically, the correction parameters are used to multiply with the basis function matrix and then participate in the linearization pre-correction link in large bandwidth or large-scale MIMO scenarios.

[0079] The linearization pre-correction method provided in an embodiment of the present invention performs LU decomposition on the basis function matrix of the current input signal based on a coarse-grained reconfigurable computing architecture, so as to orthogonalize the basis function matrix using the decomposition matrix. This can improve the processing efficiency of the orthogonalization of the basis function matrix, thereby ensuring the real-time performance of linearization pre-correction in large bandwidth or large-scale MIMO scenarios, and has fewer scenario limitations.

[0080] As a preferred solution, performing LU decomposition on the basis function matrix based on the coarse-grained reconfigurable computing architecture to obtain a decomposition matrix specifically includes:

[0081] Based on the LU decomposition algorithm, determining a set of multiplication operations of several groups of vectors and scalars of the basis function matrix in the LU decomposition process;

[0082] Mapping the processing values ​​of several groups of the multiplication operation sets to the processing unit array in the coarse-grained reconfigurable computing architecture and executing the LU decomposition algorithm to obtain the decomposition matrix; wherein the processing values ​​include vector values ​​and scalar values.

[0083] Specifically, the execution flow of the LU decomposition algorithm is as follows:

[0084]

[0085] Therefore, based on the LU decomposition algorithm, it is possible to determine several sets of vector and scalar multiplication operations of the basis function matrix in the LU decomposition process; further, the vector values ​​and scalar values ​​in the several sets of multiplication operations are mapped to the processing unit array in the coarse-grained reconfigurable computing architecture to execute the process of the LU decomposition algorithm, thereby obtaining the decomposition matrix.

[0086] As a preferred solution, mapping the processing values ​​of the plurality of groups of the multiplication operation sets to a processing unit array in a coarse-grained reconfigurable computing architecture and executing the LU decomposition algorithm to obtain the decomposition matrix specifically includes steps S21 to S26:

[0087] Step S21, determining whether, among the scalar values ​​corresponding to the kth row in the current basis function matrix, there are any two scalar values ​​corresponding to a vector in a multiplication operation set whose length is less than a preset vector length threshold; wherein k is an integer whose initial value is 1; and the vector length threshold is set based on the number of processing units in the processing unit array;

[0088] Step S22: If present, mapping the processing values ​​of the multiplication operation set corresponding to the at least two scalar values ​​to the processing unit array based on a preset parallel mapping strategy, so as to execute the currently mapped multiple sets of multiplication operations in parallel; wherein the multiplication operation set includes multiple multiplication operations between any scalar value corresponding to the kth row and multiple vector values ​​corresponding to the kth column;

[0089] When the execution of the sets of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column is obtained according to the execution results of the elements of the current i-th row and j-th column in the basis function matrix and the multiplication operations of the current mapping; wherein i is the row corresponding to the vector value in the multiplication operation of the current mapping, and its value is k+1 to N; j is the column corresponding to the scalar value in the multiplication operation of the current mapping, and its value is k+1 to N; N represents the length of the basis function matrix;

[0090] Step S23: If not, mapping the processing value of the multiplication operation set corresponding to any scalar value to the processing unit array based on a preset serial mapping strategy to execute the currently mapped set of multiplication operations;

[0091] When the set of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column is obtained according to the execution results of the element of the current i-th row and j-th column in the basis function matrix and the multiplication operations of the current mapping;

[0092] Step S24, repeatedly executing steps S21 to S23 until all multiplication operation sets corresponding to the multiple scalar values ​​corresponding to the current k-th row are completed;

[0093] Step S25, increase the value of k by one, and re-execute steps S21 to S24 until the value of k is equal to N;

[0094] Step S26: Obtain the decomposition matrix according to the current updated elements.

[0095] Specifically, as shown in the execution process of the above-mentioned LU decomposition algorithm, the LU decomposition algorithm is executed starting from the initial value of k: 1, and it is determined whether there are any two scalar values ​​corresponding to the several scalar values ​​in the current basis function matrix. The length of the vector in the multiplication operation set is less than the preset vector length threshold. It can be understood that the number of processing units in the processing unit array determines the computing power scale of the current coarse-grained reconfigurable computing architecture. Since each processing unit is only responsible for performing the multiplication operation of a single vector value and a scalar value, this embodiment first needs to consider the length of the vector in the multiplication operation set and the number of processing units in the processing unit array, so as to provide a basis for the mapping of subsequent multiplication operation sets. The length of the vector is the number of vector values ​​contained in the vector. Based on this, it is possible to clarify the number of multiplication operations between vector values ​​and scalar values ​​that need to be performed in a set of multiplication operation sets.

[0096] Furthermore, if such a set exists, the processing values ​​of the multiplication operation set corresponding to at least two scalar values ​​are mapped to the processing unit array based on a preset parallel mapping strategy to execute several sets of multiplication operation sets currently mapped in parallel. It can be understood that when the length of the vector in the multiplication operation set corresponding to any two scalar values ​​is less than the vector length threshold, it is determined that at least two sets of vector and scalar multiplication operation sets can be executed in parallel, thereby achieving the purpose of parallel optimization and helping to improve the LU decomposition efficiency of the basis function matrix. It is worth noting that, as shown in the execution flow of the above-mentioned LU decomposition algorithm, when the for loop flow of the value i is executed, the value j is unchanged, that is, in the current multiplication operation set, A[k][j] is a fixed value, equivalent to any scalar value corresponding to the kth row, and forms a vector (A[k+1][k], A[k+2][k], ..., A[N][k]), which is equivalent to several vector values ​​corresponding to the kth column.

[0097] When the execution of several sets of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column, i.e. A[i][j] in the execution process of the above LU decomposition algorithm, is obtained based on the execution results of each multiplication operation of the current mapping, i.e. A[i][k]*A[k][j] in the execution process of the above LU decomposition algorithm. ′ [i][j], thereby refreshing the value of the element in the current i-th row and j-th column.

[0098] In addition, if it does not exist, the processing value of the multiplication operation set corresponding to any scalar value is mapped to the processing unit array based on the preset serial mapping strategy to execute the currently mapped set of multiplication operations. It can be understood that when there is no vector length in the multiplication operation set corresponding to any two scalar values ​​that is less than the vector length threshold, it indicates that the computing power scale of the current coarse-grained reconfigurable computing architecture does not support the parallel execution of at least two sets of vector and scalar multiplication operation sets, but can execute a set of vector and scalar multiplication operation sets.

[0099] When the execution of a set of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column is obtained according to the element of the current i-th row and j-th column in the basis function matrix, that is, A[i][j] in the execution process of the above LU decomposition algorithm, and the execution results of each multiplication operation of the current mapping, that is, the execution results of A[i][k]*A[k][j] in the execution process of the above LU decomposition algorithm, that is, A[i][k]*A[k][j] in the execution process of the above LU decomposition algorithm. ′ [i][j], thereby refreshing the value of the element in the current i-th row and j-th column.

[0100] Furthermore, steps S21 to S23 are repeatedly executed until the multiplication operation sets corresponding to the several scalar values ​​corresponding to the current k-th row are all completed, that is, as shown in the execution process of the above-mentioned LU decomposition algorithm, under the current value k, the for loop process of the value j is completed.

[0101] Furthermore, as shown in the execution process of the above-mentioned LU decomposition algorithm, the value of k is increased by one to start a new round of loop execution process, that is, steps S21 to S24 are re-executed until the value of k is equal to N, indicating that the LU decomposition process of the current basis function matrix has been completed, and the output decomposition matrix is ​​determined based on the current updated elements of the i-th row and j-th column.

[0102] The linearization pre-correction method provided by an embodiment of the present invention can dynamically map the multiplication operation set in the processing unit matrix based on the length of the vector in the current multiplication operation set during the LU decomposition process of the basis function matrix, and when the length of the vector in the multiplication operation set corresponding to any two scalar values ​​is less than a preset vector length threshold, at least two groups of multiplication operation sets can be mapped in the processing unit matrix, thereby achieving the purpose of parallel execution and significantly improving the LU decomposition efficiency of the basis function matrix.

[0103] As a preferred solution, the processing unit array includes a plurality of processing units, each of which supports interconnection access with processing units within a preset relative unit distance; then, the method maps the processing values ​​of the multiplication operation set corresponding to at least two scalar values ​​to the processing unit array based on a preset parallel mapping strategy, specifically comprising:

[0104] Selecting a first processing unit and a second processing unit from a plurality of the processing units in the processing unit array;

[0105] Mapping any two scalar values ​​to the first processing unit and the second processing unit respectively;

[0106] Based on a plurality of vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values, mapping processing is performed on a plurality of to-be-mapped processing units that are within the preset relative unit distance from the first processing unit or the second processing unit;

[0107] When the current number of processing units to be mapped is greater than or equal to L+1, the new scalar value corresponding to the current k-th row is mapped to any processing unit to be mapped, and based on the several vector values ​​in the multiplication operation set corresponding to the new scalar value, mapping processing is performed on several processing units to be mapped that are within the preset relative unit distance from the any processing unit to be mapped; repeat this step until the current number of processing units to be mapped is less than L+1; where L represents the length of the vector in the multiplication operation set; and the new scalar value is a scalar value that is not currently mapped.

[0108] Specifically, this embodiment assumes that each processing unit in the processing unit array supports interconnection access with processing units within a relative unit distance. For example, any processing unit supports interconnection access with processing units within a relative unit distance of 2, and follows a cyclic access mechanism, such as Figure 2 The figure is a schematic diagram of the structure of the processing unit array in an embodiment of the present invention. Taking the processing unit in the upper left corner as an example, its accessible range includes all processing units within a relative unit distance of 2 on all sides. The processing units represented by the dotted box are folded into the 4×4 processing unit array due to the existence of the cyclic access mechanism. Therefore, Figure 2 The processing units marked with "PE" are all processing units that can be interconnected and accessed by the processing unit in the upper left corner.

[0109] Furthermore, when it is determined that the length of the vector in the multiplication operation set corresponding to any two scalar values ​​is less than the vector length threshold, it indicates that the processing unit array supports mapping processing data of at least two sets of multiplication operation sets. Therefore, this embodiment first selects a first processing unit and a second processing unit from several processing units in the processing unit array to map the selected any two scalar values ​​to the first processing unit and the second processing unit respectively.

[0110] It is worth noting that, in addition to the two scalar values ​​currently mapped, the remaining scalar values ​​corresponding to the current k-th row can be simultaneously mapped to the remaining processing units that support interconnection access with the first processing unit / second processing unit for local storage, so that when the first processing unit and the second processing unit complete the execution of the current two sets of multiplication operations, if there are still any two scalar values ​​corresponding to the multiplication operation set whose vector length is less than the vector length threshold, then the other two scalar values ​​can be read from the remaining processing units that locally store the scalar values ​​through interconnection access, so as to continue to execute the other two sets of multiplication operations in parallel; in addition, after the first processing unit and the second processing unit complete the execution of the current two sets of multiplication operations, if there are still any two scalar values ​​corresponding to the multiplication operation set whose vector length is less than the vector length threshold, then the other two scalar values ​​can be mapped to the first processing unit and the second processing unit respectively, so as to continue to execute the other two sets of multiplication operations in parallel. This embodiment is not specifically limited here.

[0111] Furthermore, in order to achieve parallel optimization of any two selected sets of multiplication operations, for the multiplication operation set corresponding to the scalar value mapped by the first processing unit, based on a number of vector values ​​in the set of multiplication operations, mapping processing is performed on a number of processing units to be mapped that are within a preset relative unit distance from the first processing unit; for the multiplication operation set corresponding to the scalar value mapped by the second processing unit, based on a number of vector values ​​in the set of multiplication operations, mapping processing is performed on a number of processing units to be mapped that are within a preset relative unit distance from the second processing unit, and each processing unit to be mapped can only map one vector value.

[0112] Furthermore, after completing the mapping process for any two selected sets of multiplication operations, if the current number of processing units to be mapped is greater than or equal to N+1, it indicates that another new set of multiplication operations is supported, that is, the new scalar value corresponding to the current k-th row is mapped to any processing unit to be mapped, and based on the several vector values ​​in the multiplication operation set corresponding to the new scalar value, mapping processing is performed on several processing units to be mapped that are within a preset relative unit distance from any selected processing unit to be mapped, and this step is repeated until the current number of processing units to be mapped is less than L+1. It can be understood that since the length L of the vector in a set of multiplication operations corresponds to the need to map L processing units to be mapped, and in addition, one processing unit to be mapped is required to complete the mapping of the scalar value, the addition of another new set of multiplication operations can only be supported when the current number of processing units to be mapped is greater than or equal to L+1.

[0113] The linearization pre-correction method provided by the embodiment of the present invention can fully consider the computing power scale of the current coarse-grained reconfigurable computing architecture, and perform parallel mapping of at least two sets of vector and scalar multiplication operation sets, which helps to achieve parallel optimization of vector and scalar multiplication operation sets, improve the execution efficiency of the multiplication operation set, and optimize the group delay of the LU decomposition algorithm.

[0114] As a preferred solution, selecting the first processing unit and the second processing unit from a plurality of the processing units in the processing unit array specifically includes:

[0115] Select any one processing unit from the processing unit array as the first processing unit;

[0116] According to a selection strategy with the least shared processing units, the second processing unit is selected from the remaining processing units in the processing unit array except the first processing unit; wherein the shared processing unit is a processing unit whose relative distance from the first processing unit and the second processing unit is within the preset relative unit distance.

[0117] It can be understood that this embodiment first selects any one processing unit from the processing unit array as the first processing unit, and then traverses the remaining processing units. If there is any one processing unit as the second processing unit, the number of shared processing units between the first processing unit and the second processing unit is the least, then the processing unit is the second processing unit. It is worth noting that if there are many shared processing units between the first processing unit and the second processing unit, it indicates that the range of processing units that the first processing unit and the second processing unit can access through interconnection has a large degree of overlap, resulting in low processing unit utilization of the processing unit array. If there are many vector values ​​in any two selected sets of multiplication operation sets, there may be a problem that the vector values ​​cannot be fully mapped, thereby failing to achieve parallel optimization of at least two sets of multiplication operation sets. Therefore, this embodiment determines the first processing unit and the second processing unit according to the selection strategy of the least shared processing units, ensuring that parallel optimization of at least two sets of vector and scalar multiplication operation sets can be achieved in the processing unit array.

[0118] As a preferred solution, the mapping of the plurality of to-be-mapped processing units within the preset relative unit distance from the first processing unit or the second processing unit based on the plurality of vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values ​​specifically includes:

[0119] Based on several vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values, mapping processing is performed on several to-be-mapped processing units that are within the preset relative unit distance from the first processing unit or the second processing unit, and the total distance between the remaining to-be-mapped processing units after the mapping processing is the smallest.

[0120] Specifically, this embodiment takes into account that each processing unit can only access the remaining processing units within a preset relative unit distance. If the spacing between the processing units to be mapped remaining after the current mapping processing is large, it is possible that some of the processing units to be mapped cannot be interconnected and accessed. Even if the number of remaining processing units to be mapped is greater than or equal to L+1, another new set of multiplication operations may not be executed. Therefore, in the process of executing the mapping of any two selected sets of vector and scalar multiplication operations, this embodiment calculates the sum of the spacing between the processing units to be mapped remaining after the mapping processing, and executes the mapping method based on the minimum sum of spacing. This ensures that the topological positions of the processing units to be mapped remaining after the mapping processing are as concentrated as possible, thereby providing maximum feasibility for mapping another new set of multiplication operations.

[0121] As a preferred solution, the processing unit array includes a plurality of processing units, each of which supports interconnection access with processing units within a preset relative unit distance; then, mapping the processing value of the multiplication operation set corresponding to any scalar value to the processing unit array based on a preset serial mapping strategy specifically includes:

[0122] Select any one processing unit from the processing unit array as the third processing unit;

[0123] Mapping any scalar value to the third processing unit;

[0124] Based on a number of vector values ​​in the multiplication operation set corresponding to the arbitrary scalar value, mapping processing is performed on a number of to-be-mapped processing units that are within the preset relative unit distance from the third processing unit.

[0125] Specifically, when it is determined that there is no vector in the multiplication operation set corresponding to any two scalar values ​​whose length is less than the vector length threshold, it indicates that only one set of processing data of the multiplication operation set can be mapped in the current processing unit array. Therefore, this embodiment selects any one processing unit from the processing unit array as the third processing unit to map any one scalar value to the third processing unit.

[0126] Furthermore, in order to ensure that the mapped processing units can be interconnected and accessed to perform multiplication operations of each vector value and scalar value, this embodiment maps several processing units to be mapped that are within a preset relative unit distance from the third processing unit based on several vector values ​​in the multiplication operation set corresponding to any selected scalar value, thereby completing the mapping of a selected set of vector and scalar multiplication operations in the processing unit array.

[0127] As a preferred solution, obtaining the updated element in the current i-th row and j-th column according to the execution results of each multiplication operation of the element in the current i-th row and j-th column in the basis function matrix and the current mapping specifically includes:

[0128] According to the execution results of the multiplication operations of the element in the current i-th row and j-th column in the basis function matrix and the current mapping, the updated element in the current i-th row and j-th column is obtained by the following expression:

[0129] A ′ [i][j]=A[i][j]-A[i][k]×A[k][j];

[0130] Wherein, A[i][j] represents the element in the current i-th row and j-th column of the basis function matrix; A[i][k] represents the element in the current i-th row and k-th column of the basis function matrix, which is any vector value corresponding to the k-th column; A[k][j] represents the element in the current k-th row and j-th column of the basis function matrix, which is any scalar value corresponding to the k-th row; the product of A[i][k] and A[k][j] is the execution result of the multiplication operation; A ′ [i][j] represents the updated element in the current i-th row and j-th column.

[0131] As a preferred solution, the method specifically performs the multiplication operation set of the current mapping through the following steps:

[0132] The processing unit mapped with the vector value accesses the processing unit mapped with the scalar value belonging to the same group of multiplication operation sets as the vector value, and performs the multiplication operation between the vector value and the accessed scalar value to complete the execution of the currently mapped multiplication operation set.

[0133] It is worth noting that this embodiment uses a processing unit mapped with a scalar value by having several vector values ​​in the same set of multiplication operations share this processing unit, thereby making the most of the computing power of the coarse-grained reconfigurable computing architecture and mapping as many vector and scalar multiplication operation sets as possible to the same processing unit array for parallel processing, which helps to improve the LU decomposition efficiency of the basis function matrix.

[0134] As a preferred solution, the method specifically determines whether the basis function matrix meets the preset orthogonal processing conditions through the following steps:

[0135] Obtaining the maximum eigenvalue and the minimum eigenvalue of the cross-correlation matrix, and calculating the ratio between the maximum eigenvalue and the minimum eigenvalue;

[0136] When the absolute value of the ratio is greater than a preset ratio threshold, it is determined that the basis function matrix meets the orthogonalization processing condition.

[0137] It is worth noting that when the absolute value of the ratio between the maximum eigenvalue and the minimum eigenvalue of the cross-correlation matrix is ​​greater than the preset ratio threshold, it indicates that there will be instability problems in the basis function matrix during the inversion calculation process, and the basis function matrix needs to be orthogonalized, that is, it is determined that the basis function matrix meets the orthogonalization processing conditions.

[0138] As a preferred solution, the processing unit array includes 16 processing units, and the 16 processing units form a 4×4 processing unit array.

[0139] For example, based on the 4×4 processing unit array composed of 16 processing units proposed in this embodiment, the vector length threshold is correspondingly set to 7, so that when the length of the vector in the multiplication operation set corresponding to any two scalar values ​​is less than 7, at least two groups of multiplication operation sets containing these two groups of multiplication operation sets can be mapped to the above-mentioned 4×4 processing unit array to achieve parallel execution. It should be noted that, according to actual application requirements, the vector length threshold can also be set to 4, 5, 6, etc., as long as it is ensured that at least two groups of multiplication operation sets can be mapped in parallel when the length of the vector in the multiplication operation set is less than the set vector length threshold. This embodiment does not make specific limitations here.

[0140] In order to better demonstrate the effect of improving the efficiency of basis function matrix decomposition in the embodiments of the present invention, a specific embodiment is described below.

[0141] Taking the 7×7 basis function matrix as an example, based on the execution process of the above LU decomposition algorithm, when k=1, the corresponding scalar values ​​include A[1][2] to A[1][7], such as Figure 3 As shown, the third processing unit in the first row and the first processing unit in the third row are the first processing unit and the second processing unit determined based on the parallel mapping strategy. The first processing unit, the second processing unit, and the first processing unit / the second processing unit that can be interconnected are selected with a relative position of a unit distance of 4 to complete the mapping of the scalar value A[1][2] to A[1][7].

[0142] Further, if Figure 4As shown, for the multiplication operation set corresponding to the scalar value A[1][4] mapped by the first processing unit, based on several vector values ​​in the multiplication operation set, i.e. A[2][1] to A[7][1], several processing units to be mapped that are within a preset relative unit distance from the first processing unit are mapped, thereby completing the multiplication operation set of the vector and the scalar around the scalar value A[1][4]; for the multiplication operation set corresponding to the scalar value A[1][6] mapped by the second processing unit, based on several vector values ​​in the multiplication operation set, several processing units to be mapped that are within a preset relative unit distance from the second processing unit are mapped, thereby completing the multiplication operation set of the vector and the scalar around the scalar value A[1][6]. By Figure 4 It can be seen that there are no idle processing units in this calculation cycle, achieving maximum utilization of the processing units. After completing the access to the current scalar value, the first processing unit and the second processing unit access the scalar values ​​in the first and second processing units in the first row: A[1][3] and A[1][2] through interconnection, preparing for the next parallel operation.

[0143] Furthermore, if Figure 5 As shown, the execution results of the multiplication operations of each processing unit are temporarily stored locally, and the expression is completed by reading the element of the i-th row and j-th column in the corresponding basis function matrix: ′ [i][j]=A[i][j]-A[i][k]*A[k][j], and derive the result; Figure 5 The RF in the figure is a storage unit inside the processing unit, which is used to temporarily store the execution result of the multiplication operation.

[0144] Furthermore, the above steps are repeated until all the multiplication operations corresponding to the scalar values ​​A[1][2] to A[1][7] are completed, such as Figure 6 shown.

[0145] Furthermore, the value of k is increased by one, and the above steps are executed again until the LU decomposition process is completed.

[0146] See also Figure 7 According to a second aspect of an embodiment of the present invention, a linearization pre-correction device 100 is provided, comprising:

[0147] An orthogonalization processing condition judgment module 11 is used to obtain a cross-correlation matrix of the basis function matrix based on the current input signal and judge whether the basis function matrix meets the preset orthogonalization processing condition;

[0148] A basis function matrix decomposition module 12 is configured to perform LU decomposition on the basis function matrix based on a coarse-grained reconfigurable computing architecture to obtain a decomposition matrix when the basis function matrix satisfies the orthogonalization processing condition;

[0149] An orthogonalization processing module 13 is used to perform an orthogonalization process on the basis function matrix using the decomposition matrix to obtain a basis function orthogonal matrix;

[0150] The correction parameter acquisition module 14 is used to obtain correction parameters for linearization pre-correction based on the basis function orthogonal matrix using a least mean square algorithm.

[0151] As a preferred solution, the basis function matrix decomposition module 12 is used to perform LU decomposition on the basis function matrix based on a coarse-grained reconfigurable computing architecture to obtain a decomposition matrix, specifically including:

[0152] Based on the LU decomposition algorithm, determining a set of multiplication operations of several groups of vectors and scalars of the basis function matrix in the LU decomposition process;

[0153] Mapping the processing values ​​of several groups of the multiplication operation sets to the processing unit array in the coarse-grained reconfigurable computing architecture and executing the LU decomposition algorithm to obtain the decomposition matrix; wherein the processing values ​​include vector values ​​and scalar values.

[0154] As a preferred solution, the basis function matrix decomposition module 12 is used to map the processing values ​​of several groups of the multiplication operation sets to the processing unit array in the coarse-grained reconfigurable computing architecture and execute the LU decomposition algorithm to obtain the decomposition matrix, specifically including executing steps S21 to S26:

[0155] Step S21, determining whether, among the scalar values ​​corresponding to the kth row in the current basis function matrix, there are any two scalar values ​​corresponding to a vector in a multiplication operation set whose length is less than a preset vector length threshold; wherein k is an integer whose initial value is 1; and the vector length threshold is set based on the number of processing units in the processing unit array;

[0156] Step S22: If present, mapping the processing values ​​of the multiplication operation set corresponding to the at least two scalar values ​​to the processing unit array based on a preset parallel mapping strategy, so as to execute the currently mapped multiple sets of multiplication operations in parallel; wherein the multiplication operation set includes multiple multiplication operations between any scalar value corresponding to the kth row and multiple vector values ​​corresponding to the kth column;

[0157] When the execution of the sets of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column is obtained according to the execution results of the elements of the current i-th row and j-th column in the basis function matrix and the multiplication operations of the current mapping; wherein i is the row corresponding to the vector value in the multiplication operation of the current mapping, and its value is k+1 to N; j is the column corresponding to the scalar value in the multiplication operation of the current mapping, and its value is k+1 to N; N represents the length of the basis function matrix;

[0158] Step S23: If not, mapping the processing value of the multiplication operation set corresponding to any scalar value to the processing unit array based on a preset serial mapping strategy to execute the currently mapped set of multiplication operations;

[0159] When the set of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column is obtained according to the execution results of the element of the current i-th row and j-th column in the basis function matrix and the multiplication operations of the current mapping;

[0160] Step S24, repeatedly executing steps S21 to S23 until all multiplication operation sets corresponding to the multiple scalar values ​​corresponding to the current k-th row are completed;

[0161] Step S25, increase the value of k by one, and re-execute steps S21 to S24 until the value of k is equal to N;

[0162] Step S26: Obtain the decomposition matrix according to the current updated elements.

[0163] As a preferred solution, the processing unit array includes a plurality of processing units, each of which supports interconnection access with processing units within a preset relative unit distance; then, the basis function matrix decomposition module 12 is used to map the processing values ​​of the multiplication operation set corresponding to at least two scalar values ​​to the processing unit array based on a preset parallel mapping strategy, specifically including:

[0164] Selecting a first processing unit and a second processing unit from a plurality of the processing units in the processing unit array;

[0165] Mapping any two scalar values ​​to the first processing unit and the second processing unit respectively;

[0166] Based on a plurality of vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values, mapping processing is performed on a plurality of to-be-mapped processing units that are within the preset relative unit distance from the first processing unit or the second processing unit;

[0167] When the current number of processing units to be mapped is greater than or equal to L+1, the new scalar value corresponding to the current k-th row is mapped to any processing unit to be mapped, and based on the several vector values ​​in the multiplication operation set corresponding to the new scalar value, mapping processing is performed on several processing units to be mapped that are within the preset relative unit distance from the any processing unit to be mapped; repeat this step until the current number of processing units to be mapped is less than L+1; where L represents the length of the vector in the multiplication operation set; and the new scalar value is a scalar value that is not currently mapped.

[0168] As a preferred solution, the basis function matrix decomposition module 12 is used to select a first processing unit and a second processing unit from a plurality of processing units in the processing unit array, specifically including:

[0169] Select any one processing unit from the processing unit array as the first processing unit;

[0170] According to a selection strategy with the least shared processing units, the second processing unit is selected from the remaining processing units in the processing unit array except the first processing unit; wherein the shared processing unit is a processing unit whose relative distance from the first processing unit and the second processing unit is within the preset relative unit distance.

[0171] As a preferred solution, the basis function matrix decomposition module 12 is configured to map a plurality of to-be-mapped processing units that are within the preset relative unit distance from the first processing unit or the second processing unit based on a plurality of vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values, specifically including:

[0172] Based on several vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values, mapping processing is performed on several to-be-mapped processing units that are within the preset relative unit distance from the first processing unit or the second processing unit, and the total distance between the remaining to-be-mapped processing units after the mapping processing is the smallest.

[0173] As a preferred solution, the processing unit array includes a plurality of processing units, each of which supports interconnection access with processing units within a preset relative unit distance; then, the basis function matrix decomposition module 12 is used to map the processing value of the multiplication operation set corresponding to any scalar value to the processing unit array based on a preset serial mapping strategy, specifically including:

[0174] Select any one processing unit from the processing unit array as the third processing unit;

[0175] Mapping any scalar value to the third processing unit;

[0176] Based on a number of vector values ​​in the multiplication operation set corresponding to the arbitrary scalar value, mapping processing is performed on a number of to-be-mapped processing units that are within the preset relative unit distance from the third processing unit.

[0177] As a preferred solution, the basis function matrix decomposition module 12 is used to obtain the updated element of the current i-th row and j-th column according to the execution results of each multiplication operation of the element of the current i-th row and j-th column in the basis function matrix and the current mapping, specifically including:

[0178] According to the execution results of the multiplication operations of the element in the current i-th row and j-th column in the basis function matrix and the current mapping, the updated element in the current i-th row and j-th column is obtained by the following expression:

[0179] A ′ [i][j]=A[i][j]-A[i][k]×A[k][j];

[0180] Wherein, A[i][j] represents the element in the current i-th row and j-th column of the basis function matrix; A[i][k] represents the element in the current i-th row and k-th column of the basis function matrix, which is any vector value corresponding to the k-th column; A[k][j] represents the element in the current k-th row and j-th column of the basis function matrix, which is any scalar value corresponding to the k-th row; the product of A[i][k] and A[k][j] is the execution result of the multiplication operation; A ′ [i][j] represents the updated element in the current i-th row and j-th column.

[0181] As a preferred solution, the basis function matrix decomposition module 12 is used to perform a set of multiplication operations of the current mapping, specifically including:

[0182] The processing unit mapped with the vector value accesses the processing unit mapped with the scalar value belonging to the same group of multiplication operation sets as the vector value, and performs the multiplication operation between the vector value and the accessed scalar value to complete the execution of the currently mapped multiplication operation set.

[0183] As a preferred solution, the orthogonalization processing condition judgment module 11 is used to judge whether the basis function matrix meets the preset orthogonalization processing conditions, specifically including:

[0184] Obtaining the maximum eigenvalue and the minimum eigenvalue of the cross-correlation matrix, and calculating the ratio between the maximum eigenvalue and the minimum eigenvalue;

[0185] When the absolute value of the ratio is greater than a preset ratio threshold, it is determined that the basis function matrix meets the orthogonalization processing condition.

[0186] As a preferred solution, the processing unit array includes 16 processing units, and the 16 processing units form a 4×4 processing unit array.

[0187] It should be noted that the linearization pre-correction device 100 provided in an embodiment of the present invention can implement all the processes of the linearization pre-correction method described in any of the above embodiments. The functions of each module in the device and the technical effects achieved are respectively the same as the functions and technical effects achieved by the linearization pre-correction method described in the above embodiments, and will not be repeated here.

[0188] See also Figure 8 , is a schematic diagram of the structure of a terminal device 200 provided in an embodiment of the present invention. A third aspect of the embodiments of the present invention provides a terminal device 200, comprising a memory 22, a processor 21, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, it implements the linearization pre-correction method described in any embodiment of the first aspect.

[0189] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 22 and executed by the processor 21 to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device 200.

[0190] The terminal device 200 may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art will appreciate that the schematic diagram is merely an example of the terminal device 200 and does not limit the terminal device 200. The terminal device 200 may include more or fewer components than shown, or may combine certain components or different components. For example, the terminal device 200 may also include input and output devices, network access devices, buses, and the like.

[0191] The processor 21 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor 21 is the control center of the terminal device 200, and connects various parts of the entire terminal device 200 using various interfaces and lines.

[0192] The memory 22 can be used to store the computer programs and / or modules. The processor 21 implements the various functions of the terminal device 200 by running or executing the computer programs and / or modules stored in the memory 22 and calling the data stored in the memory 22. The memory 22 can mainly include a program storage area and a data storage area. The program storage area can store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 22 can include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0193] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the linearization pre-correction method described in any embodiment of the first aspect.

[0194] A fifth aspect of the embodiments of the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the linearization pre-correction method described in any embodiment of the first aspect.

[0195] Wherein, if the module / unit integrated in the terminal device 200 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0196] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A linearization pre-correction method, characterized in that: include: Based on the basis function matrix of the current input signal, obtaining a cross-correlation matrix of the basis function matrix and determining whether the basis function matrix meets a preset orthogonal processing condition; When the basis function matrix satisfies the orthogonalization processing condition, performing LU decomposition on the basis function matrix based on a coarse-grained reconfigurable computing architecture to obtain a decomposition matrix; orthogonalizing the basis function matrix using the decomposition matrix to obtain a basis function orthogonal matrix; Based on the basis function orthogonal matrix, a least mean square algorithm is used to obtain correction parameters for linearization pre-correction.

2. The linearization pre-correction method according to claim 1, wherein: The step of performing LU decomposition on the basis function matrix based on the coarse-grained reconfigurable computing architecture to obtain a decomposition matrix specifically includes: Based on the LU decomposition algorithm, determining a set of multiplication operations of several groups of vectors and scalars of the basis function matrix in the LU decomposition process; Mapping the processing values ​​of several groups of the multiplication operation sets to the processing unit array in the coarse-grained reconfigurable computing architecture and executing the LU decomposition algorithm to obtain the decomposition matrix; wherein the processing values ​​include vector values ​​and scalar values.

3. The linearization pre-correction method according to claim 2, wherein: Mapping the processing values ​​of the plurality of multiplication operation sets to the processing unit array in the coarse-grained reconfigurable computing architecture and executing the LU decomposition algorithm to obtain the decomposition matrix specifically includes steps S21 to S26: Step S21, determining whether, among the scalar values ​​corresponding to the kth row in the current basis function matrix, there are any two scalar values ​​corresponding to a vector in a multiplication operation set whose length is less than a preset vector length threshold; wherein k is an integer whose initial value is 1; and the vector length threshold is set based on the number of processing units in the processing unit array; Step S22: If present, mapping the processing values ​​of the multiplication operation set corresponding to the at least two scalar values ​​to the processing unit array based on a preset parallel mapping strategy, so as to execute the currently mapped multiple sets of multiplication operations in parallel; wherein the multiplication operation set includes multiple multiplication operations between any scalar value corresponding to the kth row and multiple vector values ​​corresponding to the kth column; When the execution of the sets of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column is obtained according to the execution results of the elements of the current i-th row and j-th column in the basis function matrix and the multiplication operations of the current mapping; wherein i is the row corresponding to the vector value in the multiplication operation of the current mapping, and its value is k+1 to N; j is the column corresponding to the scalar value in the multiplication operation of the current mapping, and its value is k+1 to N; N represents the length of the basis function matrix; Step S23: If not, mapping the processing value of the multiplication operation set corresponding to any scalar value to the processing unit array based on a preset serial mapping strategy to execute the currently mapped set of multiplication operations; When the set of multiplication operations of the current mapping is completed, the updated element of the current i-th row and j-th column is obtained according to the execution results of the element of the current i-th row and j-th column in the basis function matrix and the multiplication operations of the current mapping; Step S24, repeatedly executing steps S21 to S23 until all multiplication operation sets corresponding to the multiple scalar values ​​corresponding to the current k-th row are completed; Step S25, increase the value of k by one, and re-execute steps S21 to S24 until the value of k is equal to N; Step S26: Obtain the decomposition matrix according to the current updated elements.

4. The linearization pre-correction method according to claim 3, wherein: The processing unit array includes a plurality of processing units, each of which supports interconnection access with processing units within a preset relative unit distance. The method maps processing values ​​of a multiplication operation set corresponding to at least two scalar values ​​to the processing unit array based on a preset parallel mapping strategy, specifically comprising: Selecting a first processing unit and a second processing unit from a plurality of the processing units in the processing unit array; Mapping any two scalar values ​​to the first processing unit and the second processing unit respectively; Based on a plurality of vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values, mapping processing is performed on a plurality of to-be-mapped processing units that are within the preset relative unit distance from the first processing unit or the second processing unit; When the current number of processing units to be mapped is greater than or equal to L+1, the new scalar value corresponding to the current k-th row is mapped to any processing unit to be mapped, and based on the several vector values ​​in the multiplication operation set corresponding to the new scalar value, mapping processing is performed on several processing units to be mapped that are within the preset relative unit distance from the any processing unit to be mapped; repeat this step until the current number of processing units to be mapped is less than L+1; where L represents the length of the vector in the multiplication operation set; and the new scalar value is a scalar value that is not currently mapped.

5. The linearization pre-correction method according to claim 4, wherein: The selecting a first processing unit and a second processing unit from a plurality of the processing units in the processing unit array specifically includes: Select any one processing unit from the processing unit array as the first processing unit; According to a selection strategy with the least shared processing units, the second processing unit is selected from the remaining processing units in the processing unit array except the first processing unit; wherein the shared processing unit is a processing unit whose relative distance from the first processing unit and the second processing unit is within the preset relative unit distance.

6. The linearization pre-correction method according to claim 4, wherein: The mapping of the plurality of to-be-mapped processing units within the preset relative unit distance from the first processing unit or the second processing unit based on the plurality of vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values ​​specifically includes: Based on several vector values ​​in the multiplication operation set corresponding to the arbitrary two scalar values, mapping processing is performed on several to-be-mapped processing units that are within the preset relative unit distance from the first processing unit or the second processing unit, and the total distance between the remaining to-be-mapped processing units after the mapping processing is the smallest.

7. The linearization pre-correction method according to claim 3, wherein: The processing unit array includes a plurality of processing units, each of which supports interconnection access with processing units within a preset relative unit distance; then, mapping the processing value of the multiplication operation set corresponding to any scalar value to the processing unit array based on a preset serial mapping strategy specifically includes: Select any one processing unit from the processing unit array as the third processing unit; Mapping any scalar value to the third processing unit; Based on a number of vector values ​​in the multiplication operation set corresponding to the arbitrary scalar value, mapping processing is performed on a number of to-be-mapped processing units that are within the preset relative unit distance from the third processing unit.

8. The linearization pre-correction method according to claim 3, wherein: The step of obtaining the updated element in the current i-th row and j-th column according to the execution results of the multiplication operations of the element in the current i-th row and j-th column in the basis function matrix and the current mapping specifically includes: According to the execution results of the multiplication operations of the element in the current i-th row and j-th column in the basis function matrix and the current mapping, the updated element in the current i-th row and j-th column is obtained by the following expression: A ′ [i][j]=A[i][j]-A[i][k]×A[k][j]; Wherein, A[i][j] represents the element in the current i-th row and j-th column of the basis function matrix; A[i][k] represents the element in the current i-th row and k-th column of the basis function matrix, which is any vector value corresponding to the k-th column; A[k][j] represents the element in the current k-th row and j-th column of the basis function matrix, which is any scalar value corresponding to the k-th row; the product of A[i][k] and A[k][j] is the execution result of the multiplication operation; A ′ [i][j] represents the updated element in the current i-th row and j-th column.

9. The linearization pre-correction method according to claim 4, wherein: The method specifically performs the multiplication operation set of the current mapping through the following steps: The processing unit mapped with the vector value accesses the processing unit mapped with the scalar value belonging to the same group of multiplication operation sets as the vector value, and performs the multiplication operation between the vector value and the accessed scalar value to complete the execution of the currently mapped multiplication operation set.

10. The linearization pre-correction method according to claim 1, wherein: The method specifically determines whether the basis function matrix meets the preset orthogonal processing conditions through the following steps: Obtaining the maximum eigenvalue and the minimum eigenvalue of the cross-correlation matrix, and calculating the ratio between the maximum eigenvalue and the minimum eigenvalue; When the absolute value of the ratio is greater than a preset ratio threshold, it is determined that the basis function matrix meets the orthogonalization processing condition.

11. The linearization pre-correction method according to any one of claims 2 to 9, characterized in that: The processing unit array includes 16 processing units, and the 16 processing units form a 4×4 processing unit array.

12. A linearization pre-correction device, characterized in that: include: An orthogonalization processing condition judgment module is used to obtain a cross-correlation matrix of the basis function matrix based on the current input signal and judge whether the basis function matrix meets the preset orthogonalization processing condition; a basis function matrix decomposition module, configured to perform LU decomposition on the basis function matrix based on a coarse-grained reconfigurable computing architecture to obtain a decomposition matrix when the basis function matrix satisfies the orthogonalization processing condition; An orthogonalization processing module, configured to perform an orthogonalization process on the basis function matrix using the decomposition matrix to obtain a basis function orthogonal matrix; The correction parameter acquisition module is used to obtain correction parameters for linearization pre-correction based on the basis function orthogonal matrix using a least mean square algorithm.

13. A terminal device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the linearization pre-correction method according to any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the linearization pre-correction method according to any one of claims 1 to 11.

15. A computer program product, characterized in that The method comprises a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the linearization pre-correction method according to any one of claims 1 to 11 is implemented.