A signal detection method and system based on reinforcement learning
Through the signal detection method based on reinforcement learning, the SARSA algorithm is used to optimize the beam direction of the antenna array, and combined with the reward mechanism, the problems of low detection probability and high false alarm rate in traditional methods are solved, and fast and accurate object detection in complex electromagnetic environments are achieved.
Patent Information
- Application Number
- CN202510428447.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Traditional signal detection methods are difficult to determine the threshold value when the number of antenna array elements is small, resulting in low detection probability or high false alarm rate, slow mechanical scanning speed, complex electronic scanning control, and difficult to achieve fast and accurate target detection in complex electromagnetic environments.
Using a signal detection method based on reinforcement learning, the antenna array beam direction is optimized through the SARSA algorithm, combining positive and negative reward mechanisms to achieve closed-loop dynamic optimization, and adaptively adjust beam direction to improve the target detection probability.
In a complex electromagnetic environment, fast and accurate object detection is achieved, which improves the probability of object detection and reduces false alarm rate, avoids the speed limit of mechanical scanning and complex control of electronic scanning.
Smart Images

Figure CN119945587B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent electromagnetic signal processing, and particularly to a signal detection method and system based on reinforcement learning. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] The target detection probability depends on the precise adjustment of the beam direction of the antenna array. When the number of antenna elements is small, traditional signal detection methods usually rely on energy detection methods with fixed thresholds. It is not easy to determine the threshold value. If the threshold is too high, the signal cannot be detected; if the threshold is too low, the false alarm rate is too high, seriously affecting the signal detection probability. When the number of antenna elements is large, multiple beams with narrow beam widths can be formed, and target detection can be performed by controlling the beam direction. There are usually two methods: mechanical scanning and electronic scanning. The mechanical scanning speed is slow, which is not conducive to rapid target detection. Electronic scanning controls the transmitting or receiving direction of the antenna by changing the phase or amplitude of the radio frequency signal, but the beam scanning control is complex. Summary of the Invention
[0004] In order to solve the technical problems existing in the above background art, the present invention provides a signal detection method and system based on reinforcement learning. The present invention realizes closed-loop dynamic optimization of the antenna array beam direction by optimizing the target signal detection parameters based on reinforcement learning, combines the target detection probability, adapts to complex electromagnetic environment changes, and improves the target detection probability.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] The first aspect of the present invention provides a signal detection method based on reinforcement learning.
[0007] A signal detection method based on reinforcement learning includes:
[0008] Initialize the relevant parameters of the SARSA algorithm, use the beam angle region where a target may exist as the current action, obtain the current received perception data of the current action, and determine the threshold value;
[0009] Divide the current received perception data into L discrete angle grids, and at the current moment, judge one by one whether the signals in the L beam angle grids are higher than the threshold value, count the total number of angle grids where a target may exist at the current moment as the state at the current moment, and calculate the state at the next moment; calculate the target detection probability according to the proportion of all targets in the current received perception data, and then calculate the reward function at the next moment to update the greedy strategy;
[0010] According to the actions and states at the current moment, the updated greedy policy is used to select the action with the largest Q function as the action for the next moment. Combining the state of the next moment, the SARSA algorithm is iteratively looped. When the set conditions are met, the beam pointing is output.
[0011] Furthermore, at the current moment, it is determined one by one whether the signals in L beam angle grids are higher than the threshold value, and the total number of angle grids where targets may exist at the current moment is counted; the method includes: determining one by one whether the signals in L beam angle grids are higher than the threshold value at the current moment. If so, there may be a target; otherwise, there is no target, and the total number of angle grids where targets may exist at the current moment is counted.
[0012] Furthermore, the threshold value is expressed by the following formula:
[0013]
[0014] where represents the threshold value, represents the false alarm rate, represents the noise power, N represents the signal length, represents the complementary error inverse function.
[0015] Furthermore, the calculation of the target detection probability and then the calculation of the reward function for the next moment; are expressed by the following formula:
[0016]
[0017] where represents k+1 the reward function at the represents k state at the , respectively represent the target detection probabilities of the l , q th beam angle regions; represents the sum of the detection probabilities of the angle grids where targets may exist, that is, the target detection probability at the current moment; represents the detection probability of the remaining angle grids where targets may not exist.
[0018] Even further, the process of updating the greedy policy includes: using the target detection probability at the current moment as a positive reward, using the detection probabilities of the remaining angle grids where targets may not exist as negative rewards, and updating the greedy policy to achieve closed-loop feedback of the target detection probability.
[0019] Further, the total number of angular grids where all possible targets may exist at the current moment is statistically calculated and expressed by the following formula:
[0020]
[0021] wherein, represents the statistic, represents the steering vector, represents k all the data received within the l th beam angle region at the H moment, represents the covariance matrix of the noise,
[0022] The second aspect of the present invention provides a signal detection system based on reinforcement learning.
[0023] A signal detection system based on reinforcement learning includes:
[0024] A data acquisition module, which is configured to: initialize the relevant parameters of the SARSA algorithm, use the beam angle region where a target may exist as the current action, acquire the current received perception data of the current action, and determine the threshold value;
[0025] A parameter calculation module, which is configured to: divide the current received perception data into L discrete angular grids, judge one by one at the current moment whether the signals in the L beam angle grids are higher than the threshold value, statistically calculate the total number of angular grids where all possible targets may exist at the current moment as the state at the current moment, and calculate the state at the next moment; calculate the target detection probability according to the proportion of all targets in the current received perception data, and then calculate the reward function at the next moment to update the greedy policy;
[0026] An iterative output module, which is configured to: select, according to the action and state at the current moment, the action with the largest Q function value by using the updated greedy policy as the action at the next moment, combine with the state at the next moment, iteratively loop the SARSA algorithm, and output the beam pointing when the set conditions are met.
[0027] The third aspect of the present invention provides a computer-readable storage medium.
[0028] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the signal detection method based on reinforcement learning described in the first aspect above.
[0029] The fourth aspect of the present invention provides a computer device.
[0030] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the signal detection method based on reinforcement learning described in the first aspect above.
[0031] The fifth aspect of the present invention provides a computer program product or a computer program.
[0032] The present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the steps in the signal detection method based on reinforcement learning described in the first aspect above.
[0033] Compared with the prior art, the beneficial effects of the present invention are:
[0034] The present invention provides a signal detection method and system based on reinforcement learning. Relying on the powerful policy optimization ability of reinforcement learning, through the action execution and evaluation feedback mechanism, an optimal policy is formulated, and the detection probability of the signal is used as the evaluation criterion to provide closed-loop feedback for the model, realizing the autonomous dynamic optimization adjustment of the beam pointing of the antenna array and improving the target detection probability in a complex electromagnetic environment.
[0035] The present invention designs a reward mechanism that combines positive rewards and negative rewards, taking into account both the impact of the target detection probability and the false alarm rate. The negative reward is added to the reward function as a penalty term, and the evaluation result is more in line with the actual situation, realizing more accurate beam control.
[0036] The present invention automatically optimizes and adjusts the beam pointing through reinforcement learning, without the need for mechanical scanning or beam scanning control based on complex weight design, improving the speed and detection probability of target detection in a dynamic complex electromagnetic environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings forming a part of this invention are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0038] Figure 1 is a flowchart of the signal detection method based on reinforcement learning shown in the present invention;
[0039] Figure 2 is a flowchart of the closed-loop optimization of reinforcement learning shown in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0041] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which the present invention belongs.
[0042] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0043] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and systems according to various embodiments of the present disclosure. It should be noted that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Similarly, it should be noted that each block in the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, may be implemented using a dedicated hardware-based system for performing the specified functions or operations, or may be implemented using a combination of dedicated hardware and computer instructions.
[0044] Fast and accurate target signal detection in a complex electromagnetic environment is beneficial to perceiving the surrounding environment and further provides a basis for subsequent signal analysis and processing. Therefore, it is necessary to study fast and accurate detection for multiple targets in a complex environment. To this end, the present invention provides a signal detection method and system based on reinforcement learning, which will be described in detail below through several embodiments:
[0045] Embodiment 1
[0046] As Figure 1 、 Figure 2As shown in the figure, this embodiment provides a signal detection method based on reinforcement learning. In this embodiment, taking the application of this method to a server as an example, it can be understood that this method can also be applied to a terminal, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, web servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this. In this embodiment, the method includes the following steps:
[0047] Step 1: Model initialization: Initialize the parameters of the SARSA (state-action-reward-state-action) algorithm, including the Q matrix, state, action, learning rate, discount factor, gradient update step size, etc.
[0048] Step 2: Obtain the current antenna received data according to the current moment action.
[0049] Step 3: Divide the space where the antenna is located into beam angle regions. By comparing the received data within L beam angle regions at the current moment with the threshold value, determine whether there is a target within the beam angle region, and calculate the state at the next moment.
[0050] Step 4: Since the target detection situation is related to the size of the set threshold value, if the threshold value is too large, it is easy to miss detections, and if the threshold value is too small, it is easy to have false detections; therefore, the false alarm rate and the threshold value can be preset to detect the target in the antenna received data, calculate the target detection probability at the current moment according to the proportion of the target in all received data; calculate the reward function at the next moment according to the target detection probability at the current moment.
[0051] Step 5: According to the current action and state, adopt The greedy strategy selects the action with the largest Q function as the optimal action at the next moment, where the optimal action represents the beam angle region where a target may exist, so as to optimize and adjust the antenna array beam pointing. Update the Q function according to the current action, state, and the action, state, and reward function at the next moment, and enter the next round of loop.
[0052] Step 6: Gradually optimize the decision estimation of actions, enabling the antenna array to adaptively optimize and adjust the beam direction under changing environmental conditions, and improving the target detection probability.
[0053] Step 7: During the target detection process, adaptively adjust the beam direction according to different environmental states, and utilize the reward function feedback mechanism to cyclically update the strategy to improve the target detection probability in a complex electromagnetic environment.
[0054] Through designing a framework for reinforcement learning decision-making and target detection probability evaluation, the present invention realizes the closed-loop adaptive optimization adjustment of the beam direction of the antenna array for target detection, decision-making, evaluation, and beam direction optimization, and improves the adaptability of target detection to a complex dynamic electromagnetic environment.
[0055] The signal detection method using the SARSA algorithm in this embodiment is described in detail below, including:
[0056] Obtain parameter information affecting the target detection probability, including the currently received perception data, and respectively set the detection threshold, false alarm rate, etc.
[0057] Initialize the parameters of the SARSA algorithm, including the Q matrix, state, action, learning rate, discount factor, gradient update step size, etc.
[0058] Train and infer the beam direction decision of the antenna array to realize the parameter adjustment of the beam direction of the antenna array. The SARSA algorithm updates the Q function based on the following rules:
[0059]
[0060] Among them, represents the number of all beam angles where the target exists at time k; represents the beam angle where the target may exist at time k; represents the learning rate, ; represents the discount factor; represents the reward at time k + 1.
[0061] (1) State space:
[0062] First, review the target detection process. Divide the current perception area (the area where the receiving antenna data is located) into L discrete angular grids, and at the current moment, judge one by one whether the signal within the L beam angle ranges is higher than the current threshold value. The threshold value is related to the false alarm rate. If it is higher than the threshold value, there is a target; if it is lower than the threshold value, there is no target. The specific introduction is as follows:
[0063] For each angular grid, it can be divided into the following two situations:
[0064] ,
[0065] Among them, assume that represents only noise, that is ; represents k the noise within the l th beam angle region at time ; assume that represents the signal and noise, that is ; k represents l the amplitude of the signal within the th beam angle region at time
[0066] The relationship between the statistics of the received data and the threshold is as follows:
[0067] , .
[0068] Among them, represents the statistics of the received data containing the signal and noise, represents the statistics of the received data containing only noise, represents the threshold.
[0069] For each angular grid, the statistics can be calculated using the following formula:
[0070] .
[0071] Among them, represents the statistics, represents k all the data received within the l th beam angle region at time H represents the covariance matrix of the noise,
[0072]
[0073] .
[0074] Among them, represents the threshold, represents the false alarm rate, N represents the noise power, represents the signal length, represents the complementary error inverse function.
[0075] Define the total number of beam angle grids where targets may statistically exist. At the current time, if the If there is a target in a beam angle grid, it is set to 1; otherwise, it is set to 0. The state at the current moment is defined as the total number of beam angle grids where a target may exist. Therefore, assuming there are at most M targets in the sensing area (the total number of beam angle grids where a target may exist is M), the possible state set is .
[0076] (2)Action space:
[0077] Assuming there are at most M targets in the sensing area, the cardinality of the action set is M. Therefore, the action at the current moment can be defined as:
[0078] ,
[0079] where indicates that there is a target in the beam angle grid area, represents the i th beam angle grid area where there is a target. The action space is the set of all possible beam angle grids, denotes the action.
[0080] (3)Policy:
[0081] The algorithm adopts a greedy policy, representing the probability of randomly searching for a new action (beam pointing). The optimal action is selected by maximizing the Q function with a probability of . The action selection at the th moment is specifically expressed as follows:
[0082] .
[0083] where represents the action at time k + 1, represents the optimal action, represents the random action.
[0084] This indicates that if is set to 0, the antenna array beam pointing is not adjusted at all, and the optimal action is always selected. If is set to 1, the action remains random all the time, and the adjustment of the antenna array beam pointing does not utilize the previously learned information and does not save the Q function.
[0085] (4)Reward function:
[0086] The reward function defines the goal of the reinforcement learning problem. Therefore, the goal of the present invention is to maximize the reward function. The reward function can be divided into two parts, namely, the negative reward and the positive reward. If there is a false alarm, the negative reward can be used as a penalty term. The positive reward is the target detection probability within the beam angle grid where the target exists at the current moment, and the negative reward is the target detection probability within the remaining beam angle grids at the current moment. The reward function can be expressed as:
[0087]
[0088] wherein, represents k+1 the reward function at the represents k state at the and respectively represent the target detection probabilities of the l th q beam angle regions; the first term represents the sum of the detection probabilities of the angle grids where the target may exist, that is, the target detection probability at the current moment; the second term represents the detection probabilities of the remaining angle grids where the target may not exist.
[0089] In the present invention, the positive reward and the negative reward are considered in the reward function. The target detection probability within the beam angle grid where the target exists at the current moment is used as the positive reward, and the target detection probability within the remaining beam angle grids at the current moment is used as the negative reward. A penalty term is added to reduce the influence of false alarms on the decision-making, and the evaluation result is fed back to the reinforcement learning module to update the algorithm strategy, realizing the closed-loop feedback of the target detection probability.
[0090] The reinforcement learning of the present invention has a powerful policy optimization ability. By optimizing and adjusting the beam pointing through reinforcement learning and formulating the optimal policy by using the action execution and evaluation feedback mechanism, it can overcome the disadvantages of slow mechanical scanning speed and complex beam scanning control of electronic scanning at the same time, and realize fast and accurate target detection in a complex electromagnetic environment.
[0091] Embodiment 2
[0092] This embodiment provides a signal detection system based on reinforcement learning.
[0093] A signal detection system based on reinforcement learning includes:
[0094] A data acquisition module, which is configured to: initialize the relevant parameters of the SARSA algorithm, use the beam angle region where the target may exist as the current action, acquire the current received perception data of the current action, and determine the threshold value;
[0095] A parameter calculation module, which is configured to: divide the currently received sensing data into L discrete angular grids, determine whether the signals in the L beam angular grids are higher than the threshold value one by one at the current moment, count the total number of angular grids where targets may exist at the current moment as the state at the current moment, and calculate the state at the next moment; calculate the target detection probability according to the proportion of all targets in the currently received sensing data, and then calculate the reward function at the next moment, and update the greedy policy;
[0096] An iterative output module, which is configured to: according to the action and state at the current moment, select the action with the largest Q function using the updated greedy policy as the action at the next moment, combine the state at the next moment, iterate the SARSA algorithm in a loop, and output the beam pointing when the set conditions are met.
[0097] In some embodiments, the parameter calculation module is further configured to: determine whether the signals in the L beam angular grids are higher than the threshold value one by one at the current moment. If so, there may be a target; otherwise, there is no target, and count the total number of angular grids where targets may exist at the current moment.
[0098] In some embodiments, the threshold value is represented by the following formula:
[0099] .
[0100] Where represents the threshold value, represents the false alarm rate, represents the noise power, N represents the signal length, represents the complementary error inverse function.
[0101] In some embodiments, the calculation of the target detection probability and then the calculation of the reward function at the next moment are represented by the following formula:
[0102]
[0103] Where represents k+1 the reward function at the represents k state at the , respectively represent the target detection probabilities of the l , q th beam angular regions; represents the sum of the detection probabilities of the angular grids where targets may exist, that is, the target detection probability at the current moment; Represents the detection probability of the remaining angular grids where the target may not exist.
[0104] In some embodiments, the parameter calculation module is further configured to: use the target detection probability at the current moment as a positive reward, and use the detection probabilities of the remaining angular grids where the target may not exist as negative rewards to update the greedy strategy to achieve closed-loop feedback of the target detection probability.
[0105] In some embodiments, the total number of angular grids where the target may exist at the current moment is statistically calculated using the following formula:
[0106] .
[0107] Where represents the statistic, represents the steering vector, represents k at time l all the received data within the H th beam angle region, represents the covariance matrix of the noise,
[0108] By using the reinforcement learning algorithm to cyclically and real-time optimize and adjust the beam pointing of the antenna array, the present invention enables the target detection to dynamically adapt to the complex electromagnetic environment with multiple dynamic targets, and significantly improves the target detection probability.
[0109] Embodiment III
[0110] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the signal detection method based on reinforcement learning as described in Embodiment I above.
[0111] Embodiment IV
[0112] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the signal detection method based on reinforcement learning as described in Embodiment I above.
[0113] Embodiment V
[0114] This embodiment provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the signal detection method based on reinforcement learning described in the above-mentioned Embodiment 1.
[0115] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.
[0116] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0117] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0119] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0120] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A signal detection method based on reinforcement learning, characterized in that, Including: Initializing relevant parameters of the SARSA algorithm, taking the beam angle region where a target may exist as the current action, obtaining the current received perception data of the current action, and determining a threshold value; the threshold value is represented by the following formula: Among them, represents the threshold value, represents the false alarm rate, represents the noise power, N represents the signal length, represents the complementary error inverse function; Divide the currently received perception data into L discrete angular grids, and at the current moment, judge one by one whether the signals in the L beam angle grids are higher than the threshold value, count the total number of angle grids where targets may exist at the current moment as the state at the current moment, and calculate the state at the next moment; calculate the target detection probability according to the proportion of all targets in the currently received perception data, and then calculate the reward function at the next moment to update the greedy strategy; According to the actions and states at the current moment, the updated greedy policy is used to select the action with the largest Q function as the action for the next moment. Combining the state of the next moment, the SARSA algorithm is iteratively looped. When the set conditions are met, the beam pointing is output.
2. The signal detection method based on reinforcement learning according to claim 1, wherein Judging one by one whether the signals in L beam angle grids at the current moment are higher than the threshold value, and counting the total number of angle grids where a target may exist at the current moment; the method includes: judging one by one whether the signals in L beam angle grids at the current moment are higher than the threshold value, if so, a target may exist, otherwise, no target exists, and counting the total number of angle grids where a target may exist at the current moment.
3. The signal detection method based on reinforcement learning according to claim 1, wherein Calculating the target detection probability, and then calculating the reward function at the next moment; Represented by the following formula: Among them, represents k+1 the reward function at time represents k the state at time and respectively represent the target detection probabilities of the l and q th beam angle regions; represents the sum of the detection probabilities of the angle grids where a target may exist, that is, the target detection probability at the current time; represents the detection probabilities of the remaining angle grids where a target may not exist.
4. The signal detection method based on reinforcement learning according to claim 3, wherein The update The process of the greedy strategy includes: using the target detection probability at the current moment as a positive reward, using the detection probabilities of the remaining angular grids where the target may not exist as negative rewards, and updating the greedy strategy to achieve closed-loop feedback of the target detection probability.
5. The signal detection method based on reinforcement learning according to claim 1, characterized in that The total number of angle grids where a target may exist at the current moment is counted, and is represented by the following formula: Among them, represents the statistic, represents the steering vector, represents k all the data received within the l th beam angle region at time H represents the covariance matrix of the noise, represents the estimator of the covariance matrix of the noise.
6. A signal detection system based on reinforcement learning, characterized in that, Including: A data acquisition module, which is configured to: initialize relevant parameters of the SARSA algorithm, take the beam angle region where a target may exist as the current action, obtain the current received perception data of the current action, and determine a threshold value; the threshold value is represented by the following formula: Among them, represents the threshold value, represents the false alarm rate, represents the noise power, N represents the signal length, represents the complementary error inverse function; A parameter calculation module, which is configured to: divide the currently received perception data into L discrete angular grids, determine whether the signals in the L beam angular grids are higher than the threshold value one by one at the current moment, count the total number of angular grids where targets may exist at the current moment as the state at the current moment, and calculate the state at the next moment; calculate the target detection probability according to the proportion of all targets in the currently received perception data, and then calculate the reward function at the next moment to update the greedy strategy; An iterative output module, which is configured to: according to the actions and states at the current moment, adopt the updated greedy policy to select the action with the largest Q function as the action at the next moment, combine the state at the next moment, iterate the SARSA algorithm, and output the beam pointing when the set conditions are met.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the signal detection method based on reinforcement learning according to any one of claims 1-5.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the signal detection method based on reinforcement learning according to any one of claims 1-5.
9. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in the signal detection method based on reinforcement learning according to any one of claims 1-5.
Citation Information
Patent Citations
MIMO radar multi-target detection method and system based on strong target limitation
CN119165461A
Electronic device, method for controlling electronic device, and program
WO2024058225A1