Signal detection method, device, equipment, medium and product

By using the number of visits to child nodes for shift processing in Monte Carlo tree search and combining the reward value and exploration value to select the optimal child node, the high complexity problem caused by the division operation in M-MIMO technology is solved, and more efficient signal detection is achieved.

CN120750459APending Publication Date: 2025-10-03PURPLE MOUNTAIN LAB
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510916364.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In existing M-MIMO technology, the division operation in the Monte Carlo tree search process leads to high complexity in signal detection hardware implementation, making it difficult to balance computational complexity and detection performance.

Method used

By performing shift processing based on the number of visits to child nodes during the Monte Carlo tree search process to determine the exploration value, the optimal child node is selected by combining the reward value and the exploration value, thus avoiding division calculations and reducing the complexity of hardware implementation.

Benefits of technology

The complexity of signal detection hardware implementation is effectively reduced, the efficiency and performance of signal detection are improved, and the consumption of computing resources is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120750459A_ABST
    Figure CN120750459A_ABST
Patent Text Reader

Abstract

The invention relates to a signal detection method, device and equipment, a medium and a product. The method comprises the following steps: acquiring a target receiving signal to be detected and a Monte Carlo tree corresponding to a transmitting signal; under the condition that the current node in the Monte Carlo tree is completely expanded, for each child node of the current node, shifting processing is carried out based on the number of access times of the child node, the exploration value of the child node is determined, and the optimal child node of the current node is determined based on the reward value and the exploration value of the child node. Adding the optimal child node into a current search path corresponding to the current iterative search process, taking the optimal child node as a new current node until the current search path is searched, and determining a reward value corresponding to each node in the current search path; and determining a target transmitting signal corresponding to the target receiving signal based on the reward value of each node in the Monte Carlo tree when a preset iteration end condition is satisfied. By adopting the method, the hardware implementation complexity can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wireless communication technology, and in particular to a signal detection method, apparatus, device, medium and product. Background Art

[0002] Massive multiple-input multiple-output (M-MIMO) technology is a key technology for future 5G wireless communications. This technology utilizes multiple transmit antennas and multiple receive antennas at the transmitter and receiver, respectively, without increasing bandwidth. This allows signals to be transmitted and received via multiple antennas at both ends, thereby improving communication quality and fully utilizing spatial resources.

[0003] To balance computational complexity and detection performance, related M-MIMO technologies use a Monte Carlo tree search-based signal detection method on the receiving end to recover the transmitted signal. However, the Monte Carlo tree search involves division operations when calculating the node's online confidence value, which leads to high hardware complexity in the signal detection method. Summary of the Invention

[0004] Based on this, it is necessary to provide a signal detection method, device, equipment, medium and product that can reduce the complexity of hardware implementation in order to address the above technical problems.

[0005] In a first aspect, the present application provides a signal detection method, comprising:

[0006] Obtaining a target received signal to be detected and a Monte Carlo tree corresponding to the transmitted signal;

[0007] The current iterative search process is executed according to the following search operation: when the current node in the Monte Carlo tree is fully expanded, for each child node of the current node, a shift process is performed based on the number of visits to the child node, and the exploration value of the child node is determined. The optimal child node of the current node is determined based on the reward value and the exploration value of the child node, and the optimal child node is added to the current search path corresponding to the current iterative search process. The optimal child node is used as the new current node until the current search path is searched. The reward value corresponding to each node in the current search path is determined based on the current search path and the target received signal;

[0008] The target transmit signal corresponding to the target receive signal is determined based on the reward value of each node in the Monte Carlo tree when the preset iteration end condition is met.

[0009] In one embodiment, performing a shift process based on the number of visits to a child node to determine the exploration value of the child node includes:

[0010] Get the first access count corresponding to the current node and the second access count corresponding to the child node;

[0011] Determine the data to be moved according to the first access number, and determine the number of shifted bits according to the second access number;

[0012] Shift the shifted data right by the number of shift bits to obtain the exploration value corresponding to the current node.

[0013] In one embodiment, determining a reward value corresponding to each node in the current search path based on the current search path and the target received signal includes:

[0014] determining a current score corresponding to the current search path based on a partial Euclidean distance between the current search path and the target received signal;

[0015] Based on the larger value of the current score and the reward value currently corresponding to the target node, the reward value corresponding to the target node is updated. The target node is any node in the current search path.

[0016] In one embodiment, during the current iterative search process according to the following search operation, the method further includes:

[0017] When the current node is not fully expanded, the expansion value corresponding to each unexpanded child node is obtained according to the partial Euclidean distance between each unexpanded child node of the current node and the target received signal;

[0018] The unexpanded child node with the smallest median expansion value is determined as the optimal child node of the current node;

[0019] Until the current search path is searched, including:

[0020] Until the new current node is a leaf node of the Monte Carlo tree.

[0021] In one embodiment, determining the unexpanded child node with the smallest median expansion value as the optimal child node of the current node includes:

[0022] Mark the unexpanded child nodes whose expansion value is greater than or equal to the preset distance threshold as unavailable nodes, which are nodes that are not visited during the iterative search process;

[0023] From the unexpanded child nodes whose expansion values ​​are less than a preset distance threshold, the unexpanded child node with the smallest expansion value is determined as the optimal child node of the current node.

[0024] In one embodiment, the method further comprises:

[0025] When all child nodes of the current node are unavailable nodes, the current node is marked as an unavailable node.

[0026] In a second aspect, the present application further provides a signal detection device, comprising:

[0027] An acquisition module, configured to acquire a target received signal to be detected and a Monte Carlo tree corresponding to the transmitted signal;

[0028] A search module is configured to execute a current iterative search process according to the following search operation: when a current node in a Monte Carlo tree is fully expanded, for each child node of the current node, perform a shift process based on the number of visits to the child node, determine the exploration value of the child node, determine the optimal child node of the current node based on the reward value and the exploration value of the child node, add the optimal child node to the current search path corresponding to the current iterative search process, and use the optimal child node as the new current node until the current search path is searched, and determine the reward value corresponding to each node in the current search path based on the current search path and the target received signal;

[0029] The determination module is used to determine the target transmission signal corresponding to the target reception signal based on the reward value of each node in the Monte Carlo tree when a preset iteration end condition is met.

[0030] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method described in the first aspect when executing the computer program.

[0031] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when the computer program is executed by a processor.

[0032] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.

[0033] The signal detection method, apparatus, device, medium, and product described above obtain a target received signal to be detected and a Monte Carlo tree corresponding to the transmitted signal; perform a current iterative search process according to the following search operation: when the current node in the Monte Carlo tree is fully expanded, perform a shift process on each child node of the current node based on the number of visits to the child node, determine the exploration value of the child node, determine the optimal child node of the current node based on the reward value and the exploration value of the child node, add the optimal child node to the current search path corresponding to the current iterative search process, and use the optimal child node as the new current node until the current search path is searched. Determine the reward value corresponding to each node in the current search path based on the current search path and the target received signal; and determine the target transmitted signal corresponding to the target received signal based on the reward value of each node in the Monte Carlo tree when a preset iteration end condition is met. In this way, in the iterative search process based on the Monte Carlo tree, the method of performing a shift process based on the number of visits to the child node when calculating the exploration value of the child node of the current node avoids the high hardware implementation complexity caused by the method of performing a division calculation based on the number of visits to the child node in the related art. The hardware implementation complexity of the signal detection method described above is relatively low. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 A diagram showing an application environment of a signal detection method in one embodiment;

[0036] Figure 2 1 is a flow chart of a signal detection method according to an embodiment;

[0037] Figure 3 is a schematic diagram of an exemplary structure of a Monte Carlo tree in one embodiment;

[0038] Figure 4 Schematic diagram of a flow chart of the steps of obtaining the exploration value corresponding to the target node in one embodiment;

[0039] Figure 5 Schematic diagram of a flow chart of the steps for obtaining a reward value corresponding to a target node in one embodiment;

[0040] Figure 6 is a flow chart of a signal detection method in another embodiment;

[0041] Figure 7is a structural block diagram of a signal detection device in one embodiment;

[0042] Figure 8 FIG. 4 is a diagram showing the internal structure of a communication device in one embodiment. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0044] The signal detection method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the mobile terminal 102 may also be referred to as user equipment (UE) or mobile user (MU). The mobile terminal 102 may be various mobile devices, such as a mobile phone (or "cellular" phone), a computer with a mobile terminal, etc. It may also be a portable, pocket-sized, handheld, computer-built-in or vehicle-mounted mobile device, or various Internet of Things devices. The base station 104 may be a device for communicating with the mobile terminal 102, for example, a next generation Node B (gNB) in a fifth generation new radio (5G NR) system, or a base station used in a future 6G system; a base station (BTS) in a global system for mobile communications (GSM) system or a code division multiple access (CDMA) system, a base station (NodeB) in a wideband code division multiple access (WCDMA) system, or an evolved Node B (eNB or eNodeB) in a long term evolution (LTE) system. The present application does not limit the mobile terminal 102 and the base station 104.

[0045] The signal detection method provided in the embodiment of the present application can be applied to a receiving end in a MIMO system, which can be a mobile terminal 102 or a base station 104.

[0046] The embodiment of the present application is a transmit antennas and Taking the spatial multiplexing MIMO system with multiple receiving antennas as an example, the implementation process of the provided signal detection method is explained. In this MIMO system, the transmitted signal is represented by a vector , the received signal is represented as a vector ,in, Indicates that A quadrature amplitude modulation (QAM) constellation of possible symbols, where Indicates taking the norm; represents the complex domain; the signal transmission of the MIMO system can be expressed as Formula 1:

[0047] , Formula 1

[0048] Among them, the matrix represents the channel matrix, which is assumed to be composed of independent and identically distributed (iid) complex Gaussian random variables with zero mean and unit variance, and is assumed to be perfectly estimated at the receiver; the vector Represents the additive white Gaussian noise (AWGN) vector, where each noise component obey distribution, where The mean is 0 and the variance is The complex Gaussian distribution of .

[0049] Real-Valued Decomposition (RVD) is used to convert the signal transmission in the complex domain into a real-valued system. In the equivalent real-valued system, , , , ,in, and Respectively represent plural The real and imaginary parts of ; therefore, the signal transmission in a real-valued system can be expressed as Equation 2:

[0050] , Formula 2

[0051] Among them, the matrix represents the real-valued channel matrix, the vector Represents a real-valued AWGN noise vector. represents the real-valued transmitted signal, where Is a RVD converts the original complex constellation into a real-valued modulation constellation. And the size is The MIMO system is converted to a real-valued modulation constellation And the size is Real-valued MIMO systems.

[0052] In an exemplary embodiment, the process of recovering the transmitted signal from the received signal uses the Maximum Likelihood Estimation (ML) algorithm to find the vector x from all possible combinations of vector signs. and The symbol combination with the smallest Euclidean distance (Euclidean distance for short) between them is the optimal solution , which is the recovered transmitted signal. The ML algorithm can be expressed as Formula 3:

[0053] , Formula 3

[0054] The channel matrix in formula 3 Can be decomposed into , where the matrix is a unitary matrix, the matrix is an upper triangular matrix. Using QR decomposition, Formula 3 can be reformulated as shown in Formula 4:

[0055] , Formula 4

[0056] Among them, the vector , since the matrix The upper triangular structure of and The Euclidean distance between can be decomposed into the form shown in Formula 5:

[0057] , Formula 5

[0058] Among them, the elements is a vector No. elements, elements is a matrix Middle Rank Elements of a column.

[0059] By expanding and rearranging Equation 5, the ML algorithm can be expressed as shown in Equation 6:

[0060] , Formula 6

[0061] in, ,element Represents the nth last element of vector x, element Represents the nth last element of vector z, element refers to the matrix The last Row and last row The elements of the column. In the embodiment of the present application, vector is defined and vector , respectively represent vectors and vector The End elements. By introducing the partial Euclidean distance (PED), Formula 5 can be expressed as shown in Formula 7:

[0062] , Formula 7

[0063] in, , , Represented by vector The End The partial Euclidean distance contributed by the elements, and The difference between the two vectors is given by The last element determines.

[0064] This cumulative property of PED allows the signal detection problem to be formulated as a tree search problem. In the tree search detection process based on QR decomposition, the number of levels of the search tree (excluding the root node) is equal to the length of the transmitted signal vector. Each level of the search tree corresponds to a symbol in the transmitted signal vector, advancing from the last element to the first element in the transmitted signal vector. The number of child nodes of each node is equal to the order of the real-valued modulation constellation, expressed as .

[0065] The search process starts from the root node of the search tree, which is connected to nodes, that is, the first level of the search tree has Each node in the search tree represents a possible real-valued symbol for the last element of vector x. For example, in a MIMO system using 4QAM modulation, each node in the search tree includes two child nodes, +1 and -1, representing the possible real-valued symbols for the last element of vector x. Based on the selected search algorithm, paths to different nodes are gradually filled from the end of vector x. When all elements of vector x are filled, the path from the root node to a leaf node represents a possible solution for the transmitted signal.

[0066] In an exemplary embodiment, please refer to Figure 2 , provides a signal detection method, which is applied to Figure 1 The mobile terminal 102 in FIG is taken as an example. It is understandable that the method can also be applied to the base station 104. Figure 2 As shown, the method includes the following steps 202 to 206. In which:

[0067] Step 202: Obtain the Monte Carlo tree corresponding to the target received signal and transmitted signal to be detected.

[0068] Wherein, based on the complex vector signal acquired by the receiving end and requiring MIMO detection, a target received signal to be detected is obtained.

[0069] For example, the target received signal may be a complex vector of received signals acquired by the mobile terminal. In the subsequent search process, when calculating the partial Euclidean distance, the complex vector The conversion calculation is performed on the real-valued data vector z converted by RVD decomposition and QR decomposition.

[0070] The number of layers corresponding to the Monte Carlo tree of the transmitted signal is determined based on the vector length of the transmitted signal in the MIMO system, wherein the vector length of the transmitted signal is determined based on the number of antennas at the transmitting end, and the number of child nodes of each node in the Monte Carlo tree is determined based on the real-valued constellation order corresponding to the transmitted signal, and the numerical value represented by the child node of each node is the possible real-valued symbol value of each element in the transmitted signal.

[0071] like Figure 3As shown, the Monte Carlo tree corresponding to the transmitted signal with a modulation mode of 4QAM and a vector length of 2 is shown. Except for the root node 1, which is an empty node, nodes 2 and 3 of the first layer represent +1 and -1, respectively, corresponding to the possible values ​​of the last element in the real-valued vector x. Nodes 4 and 5 of the second layer are child nodes of node 1, representing the possible values ​​of the second-to-last element in the real-valued vector x when the last element in the real-valued vector x is +1. For example, node 4 represents +1 and node 5 represents -1. Nodes 6 and 7 are child nodes of node 3, representing the possible values ​​of the second-to-last element in the real-valued vector x when the last element in the real-valued vector x is -1. The search path {node 1, node 2, node 4} represents that the value of the real-valued vector x is {+1, +1}; the search path {node 1, node 2, node 5} represents that the value of the real-valued vector x is {+1, -1}; the search path {node 1, node 3, node 6} represents that the value of the real-valued vector x is {-1, +1}, and the search path {node 1, node 3, node 7} represents that the value of the real-valued vector x is {-1, -1}.

[0072] Step 204, execute the current iterative search process according to the following search operations: when the current node in the Monte Carlo tree is fully expanded, for each child node of the current node, shift processing is performed based on the number of visits to the child node, and the exploration value of the child node is determined. The optimal child node of the current node is determined based on the reward value and the exploration value of the child node, and the optimal child node is added to the current search path corresponding to the current iterative search process. The optimal child node is used as the new current node until the current search path is searched. The reward value corresponding to each node in the current search path is determined based on the current search path and the target received signal.

[0073] Among them, the reward value of the node is used to characterize the performance of the node in the previous iterative search process, and the exploration value of the node is used to characterize the number of visits to the node in the previous iterative search process; the optimal child node of the current node is determined based on the reward value and the exploration value, wherein the reward value part is used to encourage the selection of child nodes that performed relatively well in the previous iterative search process during the search process, and the exploration value part is used to encourage the selection of child nodes that were visited relatively fewer times in the previous iterative search process during the search process; determining the optimal child node of the current node based on the reward value and the exploration value can balance the utilization and exploration in the Monte Carlo tree search process, and guide the search process to efficiently select the most promising child nodes.

[0074] In an embodiment of the present application, the process of determining the optimal child node of the current node based on the reward value and the exploration value can be expressed as determining the optimal child node of the current node based on the upper confidence bound (UCT) value of the node.

[0075] In one possible embodiment, the process of determining the optimal child node of the current node based on the reward value and exploration value of the child node includes: adding the reward value and exploration value of the child node to obtain the upper limit confidence value of the child node, and determining the child node with the largest upper limit confidence value among the child nodes of the current node as the optimal child node of the current node.

[0076] Among them, in a search process of the Monte Carlo tree, starting from the root node, the root node is used as the current node. If the current node is fully expanded, the optimal child node of the current node is determined based on the upper confidence value of each child node of the current node, and added to the current search path corresponding to the current iterative search process, and the optimal child node is updated to the current node, and the current node is continued to be judged whether it is fully expanded. If the current node is fully expanded, the optimal child node of the current node is continued to be determined based on the upper confidence value of each child node until the current search path is searched. In the case that the current node is not fully expanded, the optimal child node is determined from the unexpanded child nodes of the current node and added to the current search path corresponding to the current iterative search process. According to different strategies, it is judged whether the current search path is searched.

[0077] The current node is fully expanded if all of its child nodes have been visited. Visited means that the child node has been added to the search path as the optimal child node, has been assigned a reward, and has had its visit count increased by 1. Therefore, the current node is fully expanded if all of its child nodes can be calculated for an upper confidence limit. The current node is not fully expanded if there are unvisited child nodes, meaning there are child nodes for which an upper confidence limit cannot be calculated.

[0078] Exemplarily, when a new current node is determined, the number of visits to the new current node is increased by 1. Exemplarily, when the current search path is searched, the number of visits to each node in the current search path is increased by 1.

[0079] In one possible implementation, if the current node is not fully expanded, after determining the optimal child node from the current node's unexpanded child nodes and adding the optimal child node to the current search path, one strategy is to directly terminate the current search process. In this implementation, once a new unexpanded child node is added to the current search path, the current iterative search process ends and the search is restarted from the root node. Based on the values ​​corresponding to each node in the current search path, the partial Euclidean distance between the current search path and the target received signal is calculated to determine the reward value for each node in the current search path.

[0080] Exemplarily, when the current node is not fully expanded and the number of unexpanded child nodes of the current node is 1, the unexpanded child node is determined as the optimal child node and added to the current search path corresponding to the current iterative search process.

[0081] Exemplarily, when the current node is not fully expanded, the number of unexpanded child nodes of the current node is multiple, and the process of determining the optimal child node from the unexpanded child nodes of the current node includes: calculating the partial Euclidean distance between each unexpanded child node and the target received signal, obtaining the expansion value corresponding to each unexpanded child node, and determining the unexpanded child node with the smallest median value of each expansion value as the optimal child node.

[0082] As another example, the process of determining the best child node from the unexpanded child nodes of the current node includes: randomly selecting one unexpanded child node from each unexpanded child node as the best child node.

[0083] In this embodiment, the conditions for the end of the current iterative search process, that is, the conditions for the completion of the current search path search, include: a new unexpanded child node is added to the current search path, or the optimal child node is a leaf node in the Monte Carlo tree.

[0084] In one possible implementation, if the current node is not fully expanded and the current node has multiple unexpanded child nodes, a child node is randomly selected from the unexpanded child nodes of the current node as the optimal child node and added to the current search path. One strategy is to determine the optimal child node as the new current node and continue searching using a random selection method until a leaf node in the Monte Carlo tree is reached. The current search path is then considered complete, and the partial Euclidean distance between the current search path and the target received signal is calculated to determine the reward value for each node in the current search path. In this implementation, the condition for terminating the current iterative search process, i.e., the condition for completing the current search path, includes: the optimal child node being a leaf node in the Monte Carlo tree.

[0085] As the number of iterative searches increases, more and more nodes are expanded, and the number of node visits also increases with the number of iterative searches. The exploration value of each node is calculated to determine the new current node. At the same time, the number of node visits changes and increases in real time with the increase in the number of iterative searches. Therefore, the exploration value of the node needs to be calculated continuously during the iterative search process of the Monte Carlo tree. According to formula 8,

[0086] , Formula 8

[0087] Where c is a constant, Indicates the current node Number of visits, Refers to the child nodes of the current node The number of visits.

[0088] As can be seen from Formula 8, in the related art, it is necessary to continuously perform division calculations during the iteration process of the Monte Carlo tree, and the division calculation is relatively complex when implemented in hardware, resulting in high complexity in the hardware implementation of the signal detection method using this exploration value calculation method.

[0089] The exploration value part is used to encourage the selection of child nodes that have been accessed relatively fewer times in the previous iterative search process during the search process; the exploration value of a child node can be measured by the proportion of the number of visits to the child node in the total number of visits to the parent node corresponding to the child node. Unlike the method in the related art that determines the exploration value by the ratio of the logarithm of the number of visits to the current node v to the number of visits to the child nodes of the current node, in this embodiment, the number of visits to the child node is shifted to obtain the exploration value of the child node. Among them, the data processing process in the hardware system is in a binary manner, and shifting a data refers to dividing or multiplying the data by a power of 2. The shift processing process can reduce the complexity of implementation compared to the division processing process.

[0090] In this embodiment, shift processing is performed based on the number of visits to the child nodes. If the number of visits to the child node A is relatively large and the number of visits to the child node B at the same layer is relatively small, for example, the shift direction is right shift, that is, according to the division direction of division by the power of 2, then the exploration value obtained after shifting the shifted data based on the number of visits to the child node A is smaller than the exploration value obtained after shifting the shifted data based on the number of visits to the child node A.

[0091] In one possible implementation, the data to be moved can be determined based on the number of visits to the parent node corresponding to the child node. Figure 4In this embodiment, the shift processing is performed based on the number of visits to the child node, and the process of determining the exploration value of the child node includes steps 402 to 406, wherein:

[0092] Step 402: Obtain a first access count corresponding to the current node and a second access count corresponding to the child node.

[0093] Step 404: Determine the data to be shifted based on the first access count, and determine the number of shifted bits based on the second access count.

[0094] Exemplarily, the first access count is directly determined as the shifted data, and the second access count is directly determined as the number of shifted bits.

[0095] Exemplarily, the logarithm of the first access count is determined as the shifted data, and the integer part of the logarithm of the second access count is determined as the number of shifted bits.

[0096] Exemplarily, the first access number is first subjected to logarithm processing and then to square root processing to obtain the shifted data, and the integer part of the logarithm of the second access number is determined as the number of shifted bits.

[0097] Step 406: right-shift the shifted data by the shift bit number to obtain the exploration value corresponding to the child node.

[0098] In an exemplary implementation of this embodiment, the process of determining the exploration value of a child node can be expressed as Formula 9:

[0099] , Formula 9

[0100] in, Indicates that Shift right Bit. express The integer part of . is a constant used to adjust the ratio of exploration and reward in UCT. Indicates the current node The corresponding first visit number, Represents a child node The corresponding second visit number.

[0101] In this implementation, the first access count is first processed by taking the logarithm and then taking the square root, which can slow down the decay rate of the exploration item: as the first access count of the parent node increases, Increase, The growth rate of is much slower than the linear function, ensuring that the exploration item will not disappear prematurely; at the same time, even if the first visit number corresponding to the parent node is large, the exploration value It can still have a significant impact on child nodes with fewer visits, preventing the algorithm from converging to the local optimum too early.

[0102] In one possible embodiment, the shifted data can be determined based on the number of visits to the root node in the Monte Carlo tree. In this embodiment, the shifting process is performed based on the number of visits to the child nodes of the current node, and the process of determining the exploration value of the child node includes: obtaining the third number of visits to the root node in the Monte Carlo tree and the second number of visits corresponding to the child nodes of the current node, determining the shifted data based on the third number of visits, and right-shifting the shifted data by the number of shifted bits corresponding to the second number of visits to obtain the exploration value corresponding to the target node. In this embodiment, the shifted data is fixed to be determined based on the number of visits to the root node, and the number of shifted bits is determined based on the number of visits to the child nodes of the current node. The exploration value can measure the visits of the child nodes of the current node during the overall search process of the Monte Carlo tree.

[0103] Step 206 : Determine the target transmit signal corresponding to the target receive signal based on the reward value of each node in the Monte Carlo tree when the preset iteration end condition is met.

[0104] In one possible implementation, a larger reward value indicates a better performance of the node during the iterative search process. In this implementation, when a preset termination condition is met, starting from the root node of the Monte Carlo tree, the node with the largest reward value in each layer is selected layer by layer until the last layer of the Monte Carlo tree is reached. The symbols represented by all nodes passed along the way are used as the output result, and the target transmission signal is determined based on this output result.

[0105] In one possible implementation, a smaller reward value indicates a better performance of the node during the iterative search process. In this implementation, when a preset iteration termination condition is met, starting from the root node of the Monte Carlo tree, the node with the smallest reward value in each layer is selected layer by layer until the last layer of the Monte Carlo tree is reached. The symbols represented by all nodes passed along the way are used as the output result, and the target transmission signal is determined based on this output result.

[0106] Exemplarily, the target transmission signal is a complex vector of the transmission signal The above output result is converted into a complex form to obtain the target transmission signal.

[0107] In one possible implementation, the preset iteration termination condition includes a preset number of iterations. Exemplarily, the preset number of iterations is determined based on the scale of transmit and receive antennas in the MIMO system and prior testing. In this implementation, the search operation is terminated when the preset number of iterations is reached.

[0108] In one possible implementation, the preset iteration termination condition includes convergence of reward values ​​corresponding to nodes in the current search path. Exemplarily, a convergence threshold is set, and if the reward values ​​corresponding to nodes in the current search path obtained through multiple consecutive iterative search processes are less than the convergence threshold, the search operation is terminated.

[0109] In the above-mentioned signal detection method, a target received signal to be detected and a Monte Carlo tree corresponding to the transmitted signal are obtained; the current iterative search process is performed according to the following search operation: when the current node in the Monte Carlo tree is fully expanded, for each child node of the current node, a shift process is performed based on the number of visits to the child node to determine the exploration value of the child node; the optimal child node of the current node is determined based on the reward value and the exploration value of the child node; the optimal child node is added to the current search path corresponding to the current iterative search process, and the optimal child node is used as the new current node until the current search path is searched. The reward value corresponding to each node in the current search path is determined based on the current search path and the target received signal; and the target transmitted signal corresponding to the target received signal is determined based on the reward value of each node in the Monte Carlo tree when a preset iteration end condition is met. In this way, in the iterative search process based on the Monte Carlo tree, the method of performing a shift process based on the number of visits to the child node when calculating the exploration value of the child node of the current node avoids the problem of high hardware implementation complexity caused by the method of performing a division calculation based on the number of visits to the child node in the related art. The signal detection method provided in this embodiment has low hardware implementation complexity.

[0110] In an exemplary embodiment, a larger reward value of a node indicates that the node performed better in the previous iterative search process. Figure 5 This embodiment involves a process of determining the reward value corresponding to each node in the current search path based on the current search path and the target received signal, including steps 502 to 504. Among them:

[0111] Step 502: Determine a current score corresponding to the current search path based on a partial Euclidean distance between the current search path and the target received signal.

[0112] According to the formula , calculating the partial Euclidean distance between the current search path and the target received signal. In this embodiment, n represents the number of layers corresponding to leaf nodes, or the number of layers corresponding to the node farthest from the root node after the current search path is searched. It should be noted that the number of layers involved in this embodiment does not include the layer where the root node is located, and the number of layers is counted starting from the layer where the child nodes of the root node are located.

[0113] The smaller the partial Euclidean distance, the closer the current search path is to the target transmission signal, and the larger the corresponding current score. There is a negative correlation between the partial Euclidean distance and the current score.

[0114] In a possible implementation, the inverse of the partial Euclidean distance between the current search path and the target received signal is taken as the current score corresponding to the current search path.

[0115] In a possible implementation, the reciprocal of a partial Euclidean distance between the current search path and the target received signal is used as the current score corresponding to the current search path.

[0116] Step 504 : Based on the larger value of the current score and the reward value currently corresponding to the target node, update the reward value corresponding to the target node, where the target node is any node in the current search path.

[0117] If the current score is greater than the current reward value of the target node, the reward value of the target node is updated to the current score; if the current score is less than the current reward value of the target node, the reward value of the target node remains unchanged. For example, the initial value of the reward value corresponding to each node in the Monte Carlo tree is zero.

[0118] In related art, a node's reward value is determined by the average of the scores corresponding to the search paths it participated in during previous iterative searches. This means that a good node may perform poorly in a single iterative search, requiring multiple iterative searches to correct the situation. This iterative process relies on a large number of searches to converge, leading to low signal detection efficiency. In the signal detection method provided in this embodiment, a node's reward value is maintained at the maximum value of the scores corresponding to the search paths it participated in during previous iterative searches. This reduces the number of iterations and improves signal detection efficiency.

[0119] In an exemplary embodiment, based on Figure 2 In the embodiment shown, in the signal detection method provided, the "performing the current iterative search process according to the following search operation" described in step 204 also includes: when the current node is not fully expanded, obtaining partial Euclidean distances between each unexpanded child node of the current node and the target received signal, obtaining the expansion value corresponding to each unexpanded child node, and determining the unexpanded child node with the smallest median value of each expansion value as the optimal child node of the current node.

[0120] For example, according to the formula , calculate the partial Euclidean distance between each unexpanded child node of the current node and the target received signal, wherein, in this process, n represents the number of layers corresponding to the unexpanded child nodes of the current node.

[0121] Based on the cumulative property of the partial Euclidean distance, the partial Euclidean distance between the target received signal and the partial search path corresponding to any layer can be calculated.

[0122] In this embodiment, after determining the optimal child node of the current node, the optimal child node is continued to be used as the new current node until the current search path is searched. In this embodiment, until the current search path is searched, includes: until the new current node is a leaf node of the Monte Carlo tree.

[0123] In this embodiment, the current iterative search process is executed according to the following search operations: when the current node is fully expanded, for each child node of the current node, shift processing is performed based on the number of visits to the child node, the exploration value of the child node is determined, and the optimal child node of the current node is determined based on the reward value and the exploration value of the child node; when the current node is not fully expanded, the expansion value corresponding to each unexpanded child node is obtained based on the partial Euclidean distance between each unexpanded child node of the current node and the target received signal, and the unexpanded child node with the smallest median of each expansion value is determined as the optimal child node of the current node; the optimal child node is determined as the new current node, and the above process is repeated until the new current node is a leaf node of the Monte Carlo tree.

[0124] Exemplarily, a portion of the Euclidean distance between each non-expanded sub-node and the target received signal is used as the expansion value corresponding to each non-expanded sub-node.

[0125] In this embodiment, in each iterative search process, starting from the root node, according to the optimal child node determined based on the exploration value and the reward value or the optimal child node determined according to the expansion value, a greedy search is performed from the root node to the leaf node, so that the current search path in each iterative search process is as close as possible to the true solution of the transmitted signal, thereby improving the convergence speed of the iterative search.

[0126] In one possible implementation of this embodiment, the process of determining the unexpanded child node with the smallest median expansion value as the optimal child node of the current node includes: marking unexpanded child nodes with expansion values ​​greater than or equal to a preset distance threshold as unavailable nodes; and determining the unexpanded child node with the smallest expansion value among the unexpanded child nodes with expansion values ​​less than the preset distance threshold as the optimal child node of the current node. The unavailable node is a node that is not visited during the iterative search process.

[0127] The preset distance threshold is a distance value preset based on the communication environment or hardware resources. If a node's expansion value is greater than or equal to the preset distance threshold, the node is considered unlikely to represent a possible solution for transmitting a signal. The node is marked as unavailable and will not be visited in subsequent searches, saving computing resources and improving search speed.

[0128] For example, the preset distance threshold is , where r is a constant, is the number of transmit antennas and σ is the variance of the noise estimate.

[0129] In a possible implementation manner of this embodiment, the provided signal detection method further includes: when all child nodes of the current node are unavailable nodes, marking the current node as an unavailable node.

[0130] For example, please refer to Figure 3 If the expansion values ​​corresponding to nodes 4 and 5 are both greater than the preset distance threshold, nodes 4 and 5 are marked as unavailable nodes; although the expansion value corresponding to node 2 is less than the preset distance threshold, node 2 is also marked as an unavailable node. In the subsequent search process, the branch represented by node 2 will not be considered, and the search value and reward value of node 2 will no longer be updated, saving computing resources and improving search speed.

[0131] In one embodiment, please refer to Figure 6 , provides a signal detection method, which is applied to Figure 1 The mobile terminal 102 in FIG is taken as an example. It is understandable that the method can also be applied to the base station 104. Figure 6 As shown, the method includes the following steps 602 to 606. In which:

[0132] Step 602: Obtain the target received signal to be detected and the Monte Carlo tree corresponding to the transmitted signal.

[0133] Step 604: When the current node in the Monte Carlo tree is fully expanded, a shift process is performed on each child node of the current node based on the number of visits to the child node to determine the exploration value of the child node, and the optimal child node of the current node is determined based on the reward value and exploration value of the child node.

[0134] Among them, the shift processing is performed based on the number of visits to the child node, and the process of determining the exploration value of the child node includes: obtaining the first number of visits corresponding to the current node and the second number of visits corresponding to the child node; determining the shifted data according to the first number of visits, and determining the number of shifts according to the second number of visits; shifting the shifted data right by the number of shifts to obtain the exploration value corresponding to the child node.

[0135] The process of determining the reward value corresponding to each node in the current search path based on the current search path and the target received signal includes: determining the current score corresponding to the current search path based on the partial Euclidean distance between the current search path and the target received signal; and updating the reward value corresponding to the target node based on the larger value of the current score and the current reward value corresponding to the target node, where the target node is any node in the current search path.

[0136] In step 606, if the current node is not fully expanded, the expansion value corresponding to each unexpanded child node is obtained based on the second part of the Euclidean distance between each unexpanded child node of the current node and the target received signal, and the unexpanded child node with the smallest median of the expansion values ​​is determined as the optimal child node of the current node.

[0137] Step 608: Add the optimal child node to the current search path corresponding to the current iterative search process, and use the optimal child node as the new current node.

[0138] Step 610: Determine whether the new current node is a leaf node of the Monte Carlo tree.

[0139] Step 612: If the new current node is a leaf node of the Monte Carlo tree, then the reward value corresponding to each node in the current search path is determined based on the current search path and the target received signal.

[0140] If the new current node is a leaf node of the Monte Carlo tree, steps 606 to 610 are executed again.

[0141] Step 614: determine whether a preset iteration end condition is met.

[0142] Step 616 : If the preset iteration end condition is met, then based on the reward value of each node in the Monte Carlo tree when the preset iteration end condition is met, determine the target transmission signal corresponding to the target reception signal.

[0143] If the preset iteration end condition is not met, steps 604 to 614 are executed again.

[0144] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0145] It should be understood that the term "based on" as used herein is used to describe one or more factors that influence a determination, and does not exclude other factors that may influence the determination. For example, the phrase "determine A based on B" means that the determination of A may be based entirely or at least partially on factor B. In other words, B is a factor that influences the determination of A, but does not exclude the determination of A being based on C.

[0146] Based on the same inventive concept, the present application also provides a signal detection device for implementing the aforementioned signal detection method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more signal detection device embodiments provided below can be found in the above-mentioned limitations on the signal detection method and will not be further elaborated here.

[0147] In an exemplary embodiment, Figure 7 As shown, a signal detection device is provided, including: an acquisition module 702, a search module 704 and a determination module 706, wherein:

[0148] The acquisition module 702 is configured to acquire a target received signal to be detected and a Monte Carlo tree corresponding to the transmitted signal.

[0149] The search module 704 is used to execute the current iterative search process according to the following search operations: when the current node in the Monte Carlo tree is fully expanded, for each child node of the current node, a shift processing is performed based on the number of visits to the child node, and the exploration value of the child node is determined. The optimal child node of the current node is determined based on the reward value and the exploration value of the child node, and the optimal child node is added to the current search path corresponding to the current iterative search process. The optimal child node is used as the new current node until the current search path is searched. The reward value corresponding to each node in the current search path is determined based on the current search path and the target received signal.

[0150] The determination module 706 is configured to determine the target transmit signal corresponding to the target receive signal based on the reward value of each node in the Monte Carlo tree when a preset iteration termination condition is satisfied.

[0151] In an exemplary embodiment, the search module 704 is used to obtain the first access count corresponding to the current node and the second access count corresponding to the child node; determine the moved data based on the first access count, and determine the number of shifted bits based on the second access count; shift the moved data right by the number of shifted bits to obtain the exploration value corresponding to the child node.

[0152] In an exemplary embodiment, the search module 704 is configured to determine a current score corresponding to the current search path based on a partial Euclidean distance between the current search path and a target received signal; and update a reward value corresponding to the target node based on the larger value of the current score and the reward value currently corresponding to the target node, where the target node is any node in the current search path.

[0153] In an exemplary embodiment, during the current iterative search process according to the following search operation, the search module 704 is further configured to, if the current node is not fully expanded, determine the expansion value corresponding to each unexpanded child node based on the partial Euclidean distance between each unexpanded child node of the current node and the target received signal; and determine the unexpanded child node with the smallest median of the expansion values ​​as the optimal child node of the current node. Accordingly, during the process of executing until the current search path is complete, the search module 704 specifically performs the following steps: until the new current node becomes a leaf node of the Monte Carlo tree.

[0154] In an exemplary embodiment, the search module 704 is used to mark unexpanded child nodes whose expansion values ​​are greater than or equal to a preset distance threshold as unavailable nodes, where unavailable nodes are nodes that are not visited during the iterative search process; and from the unexpanded child nodes whose expansion values ​​are less than the preset distance threshold, the unexpanded child node with the smallest expansion value is determined as the optimal child node of the current node.

[0155] In an exemplary embodiment, the search module 704 is configured to mark the current node as an unavailable node when all child nodes of the current node are unavailable nodes.

[0156] Each module in the above-mentioned signal detection device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0157] In an exemplary embodiment, a communication device is provided. The communication device may be a mobile terminal, and its internal structure may be as shown in FIG. Figure 8 As shown. The communication device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the communication device is used to provide computing and control capabilities. The memory of the communication device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the communication device is used to exchange information between the processor and external devices. The communication interface of the communication device is used to communicate with external terminals via wired or wireless means, and the wireless means can be achieved via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a signal detection method. The display unit of the communication device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the communication device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the communication device casing, or an external keyboard, touchpad or mouse.

[0158] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the communication device to which the scheme of the present application is applied. The specific communication device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0159] In one embodiment, a communication device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0160] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0161] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0162] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0163] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0164] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0165] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A signal detection method, characterized in that: The method comprises: Obtaining a target received signal to be detected and a Monte Carlo tree corresponding to the transmitted signal; The current iterative search process is performed according to the following search operation: when the current node in the Monte Carlo tree is fully expanded, for each child node of the current node, a shift process is performed based on the number of visits to the child node, and the exploration value of the child node is determined; the optimal child node of the current node is determined based on the reward value and the exploration value of the child node; the optimal child node is added to the current search path corresponding to the current iterative search process, and the optimal child node is used as the new current node until the current search path is searched. The reward value corresponding to each node in the current search path is determined based on the current search path and the target received signal; A target transmit signal corresponding to the target receive signal is determined based on a reward value of each node in the Monte Carlo tree when a preset iteration end condition is met.

2. The method according to claim 1, characterized in that The performing shift processing based on the number of visits to the child node to determine the exploration value of the child node includes: Obtain a first access count corresponding to the current node and a second access count corresponding to the child node; Determine the data to be moved according to the first access number, and determine the number of shifted bits according to the second access number; The shifted data is right-shifted by the number of shifting bits to obtain the exploration value corresponding to the child node.

3. The method according to claim 1, characterized in that The determining, based on the current search path and the target received signal, a reward value corresponding to each node in the current search path includes: determining a current score corresponding to the current search path according to a partial Euclidean distance between the current search path and the target received signal; Based on the larger value of the current score and the reward value currently corresponding to the target node, the reward value corresponding to the target node is updated, and the target node is any node in the current search path.

4. The method according to claim 1, wherein During the current iterative search process performed according to the following search operation, the method further includes: When the current node is not fully expanded, obtaining an expansion value corresponding to each non-expanded child node of the current node according to a partial Euclidean distance between each non-expanded child node and the target received signal; Determine the unexpanded child node with the smallest median of the expansion values ​​as the optimal child node of the current node; The process until the current search path is searched includes: Until the new current node is a leaf node of the Monte Carlo tree.

5. The method according to claim 4, characterized in that The step of determining the unexpanded child node with the smallest median of the expansion values ​​as the optimal child node of the current node includes: Marking the unexpanded child nodes whose expansion values ​​are greater than or equal to a preset distance threshold as unavailable nodes, where the unavailable nodes are nodes that are not visited during the iterative search process; From the unexpanded child nodes whose expansion values ​​are smaller than the preset distance threshold, the unexpanded child node with the smallest expansion value is determined as the optimal child node of the current node.

6. The method according to claim 5, characterized in that The method further comprises: In the case that all child nodes of the current node are unavailable nodes, the current node is marked as an unavailable node.

7. A signal detection device, characterized in that: The device comprises: An acquisition module, configured to acquire a target received signal to be detected and a Monte Carlo tree corresponding to the transmitted signal; A search module, configured to execute a current iterative search process according to the following search operation: when a current node in the Monte Carlo tree is fully expanded, performing a shift process on each child node of the current node based on the number of visits to the child node, determining an exploration value of the child node, determining an optimal child node of the current node based on the reward value and the exploration value of the child node, adding the optimal child node to a current search path corresponding to the current iterative search process, and using the optimal child node as a new current node until the current search path is searched, and determining a reward value corresponding to each node in the current search path based on the current search path and the target received signal; A determination module is used to determine a target transmit signal corresponding to the target receive signal based on a reward value of each node in the Monte Carlo tree when a preset iteration end condition is met.

8. A communication device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • MIMO detection method and device based on dynamic detection tree generation and intelligent fusion terminal

    CN121308789A