Signal detection method and system based on reinforcement learning

Through the signal detection method based on reinforcement learning, the beam direction of antenna array is optimized, which solves the problem of difficult to determine the threshold value in the prior art, and improves the probability of target detection and the ability to adapt to complex electromagnetic environments.

CN119945587AActive Publication Date: 2025-05-06NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Patent Information

Application Number
CN202510428447.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of difficult to determine the threshold value in antenna array signal detection, resulting in low signal detection probability or excessive false alarm, especially in complex electromagnetic environments.

Method used

Using a signal detection method based on reinforcement learning, the target signal detection parameters are optimized through the SARSA algorithm, and the closed-loop dynamic optimization of antenna array beam direction is achieved in combination with the target detection probability, so as to adapt to changes in complex electromagnetic environments.

Benefits of technology

It improves the probability of target detection, reduces the false alarm rate, and realizes automatic dynamic optimization of antenna array beam direction, adapts to complex electromagnetic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945587A_ABST
    Figure CN119945587A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent electromagnetic signal processing, and provides a signal detection method and system based on reinforcement learning. The signal detection method based on reinforcement learning comprises the following steps: dividing current received sensing data into L discrete angle grids, judging whether signals in the L beam angle grids are higher than a threshold value one by one at the current moment, counting the total number of all angle grids in which a target possibly exists at the current moment as the state of the current moment, calculating the state of the next moment; calculating a target detection probability according to the proportion of all targets in the current received sensing data, further calculating a reward function at the next moment, and updating a # imgabs0 # greedy strategy; according to the action and the state of the current moment, the action with the maximum Q function is selected through the updated # imgabs1 # greedy strategy to serve as the action of the next moment, the state of the next moment and the iterative loop SARSA algorithm are combined, beam pointing is output when set conditions are met, and the target detection probability in the complex electromagnetic environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent electromagnetic signal processing, and in particular to a signal detection method and system based on reinforcement learning. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] The probability of target detection depends on the precise adjustment of the antenna array beam pointing. When the number of antenna elements is small, the traditional signal detection method is usually based on the energy detection method with a fixed threshold. The threshold value is not easy to determine. If the threshold is too high, the signal cannot be detected. If the threshold is too low, the false alarm is too large, which seriously affects the signal detection probability. When the number of antenna elements is large, multiple beams with narrow beam widths can be formed, and target detection can be performed by controlling the beam pointing. There are usually two methods: mechanical scanning and electronic scanning. Mechanical scanning is slow and not conducive to rapid detection of targets. Electronic scanning controls the transmission or reception direction of the antenna by changing the phase or amplitude of the RF signal, but the beam scanning control is complex. Summary of the invention

[0004] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a signal detection method and system based on reinforcement learning. The present invention optimizes the target signal detection parameters based on reinforcement learning and combines the target detection probability to realize closed-loop dynamic optimization of the antenna array beam pointing, adapt to the complex electromagnetic environment changes, and improve the target detection probability.

[0005] In order to achieve the above object, the present invention adopts the following technical solution: A first aspect of the present invention provides a signal detection method based on reinforcement learning.

[0006] A signal detection method based on reinforcement learning, comprising: Initialize the relevant parameters of the SARSA algorithm, take the beam angle area where the target may exist as the current action, obtain the current receiving perception data of the current action, and determine the threshold value; The currently received perception data is divided into L discrete angle grids. At the current moment, the signal in each of the L beam angle grids is determined to be higher than the threshold value. The total number of angle grids where all possible targets may exist is counted as the current state, and the state at the next moment is calculated. According to the proportion of all targets in the currently received perception data, the target detection probability is calculated, and then the reward function at the next moment is calculated to update Greedy strategy; According to the current action and status, the updated The greedy strategy selects the action with the largest Q function as the action at the next moment, combines the state at the next moment, iterates the SARSA algorithm, and outputs the beam pointing when the set conditions are met.

[0007] Furthermore, the method comprises: judging one by one whether the signals in the L beam angle grids are higher than the threshold value at the current moment, and counting the total number of angle grids where the target may exist at the current moment; the method comprises: judging one by one whether the signals in the L beam angle grids are higher than the threshold value at the current moment, if so, there may be a target, otherwise, there is no target, and counting the total number of angle grids where the target may exist at the current moment.

[0008] Furthermore, the threshold value is expressed by the following formula:

[0009] in, represents the threshold value, represents the false alarm rate, represents the noise power, N Indicates the signal length, Represents the complementary error inverse function.

[0010] Furthermore, the target detection probability is calculated, and then the reward function at the next moment is calculated; it is expressed by the following formula:

[0011] in, express k+1 The reward function at the moment, express k The state of the moment, , Respectively represent l , q Target detection probability in each beam angle region; The sum of the detection probabilities of the angle grids where the target may exist is the target detection probability at the current moment; Represents the detection probability of the remaining angle grids where the target may not exist.

[0012] Furthermore, the update The greedy strategy process includes: taking the target detection probability at the current moment as a positive reward, taking the detection probability of the remaining angle grids where the target may not exist as a negative reward, and updating Greedy strategy,realizes closed-loop feedback of target detection probability.

[0013] Furthermore, the total number of angle grids where all possible targets may exist at the current moment is counted, and is expressed by the following formula:

[0014] in, represents statistics, represents the steering vector, express k Moment l All the received data within the beam angle area, H represents the covariance matrix of the noise, Represents an estimator of the covariance matrix of the noise.

[0015] A second aspect of the present invention provides a signal detection system based on reinforcement learning.

[0016] A signal detection system based on reinforcement learning, comprising: A data acquisition module is configured to: initialize relevant parameters of the SARSA algorithm, take the beam angle area where the target may exist as the current action, obtain the current receiving perception data of the current action, and determine the threshold value; The parameter calculation module is configured to: divide the currently received perception data into L discrete angle grids, determine whether the signals in the L beam angle grids are higher than the threshold value at the current moment, count the total number of angle grids where all targets may exist at the current moment as the current state, and calculate the state at the next moment; calculate the target detection probability according to the proportion of all targets in the currently received perception data, and then calculate the reward function at the next moment, and update Greedy strategy; The iterative output module is configured to: use the updated The greedy strategy selects the action with the largest Q function as the action at the next moment, combines the state at the next moment, iterates the SARSA algorithm, and outputs the beam pointing when the set conditions are met.

[0017] A third aspect of the present invention provides a computer-readable storage medium.

[0018] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the signal detection method based on reinforcement learning as described in the first aspect above.

[0019] A fourth aspect of the present invention provides a computer device.

[0020] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the signal detection method based on reinforcement learning as described in the first aspect above are implemented.

[0021] A fifth aspect of the present invention provides a computer program product or a computer program.

[0022] The present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the signal detection method based on reinforcement learning as described in the first aspect above.

[0023] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a signal detection method and system based on reinforcement learning. Relying on the powerful strategy optimization capability of reinforcement learning, the optimal strategy is formulated through action execution and evaluation feedback mechanism, and the detection probability of the signal is used as an evaluation criterion to provide closed-loop feedback for the model, thereby realizing autonomous dynamic optimization adjustment of the antenna array beam pointing and improving the probability of target detection in complex electromagnetic environments.

[0024] The present invention takes into account the influence of target detection probability and false alarm rate by designing a reward mechanism that combines positive rewards with negative rewards. Negative rewards are added to the reward function as penalty items. The evaluation results are more in line with reality and more accurate beam control is achieved.

[0025] The present invention automatically optimizes and adjusts beam pointing through reinforcement learning, without the need for mechanical scanning or beam scanning control based on complex weight design, thereby improving the speed and detection probability of target detection in dynamic and complex electromagnetic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0027] Figure 1 is a flow chart of a signal detection method based on reinforcement learning shown in the present invention; Figure 2 It is a reinforcement learning closed-loop optimization flow chart shown in the present invention. DETAILED DESCRIPTION

[0028] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0029] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0030] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0031] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and systems according to various embodiments of the present disclosure. It should be noted that each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the flowchart and / or block diagram, and the combination of boxes in the flowchart and / or block diagram can be implemented using a dedicated hardware-based system that performs a specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.

[0032] Rapid and accurate target signal detection in complex electromagnetic environments is conducive to the perception of the surrounding environment, and further provides a basis for subsequent signal analysis and processing. Therefore, it is necessary to study the rapid and accurate detection of multiple targets in complex environments. To this end, the present invention provides a signal detection method and system based on reinforcement learning. The present invention is described in detail through several embodiments below: Embodiment 1 like Figure 1 , Figure 2As shown, this embodiment provides a signal detection method based on reinforcement learning. This embodiment uses the method applied to a server as an example. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application. In this embodiment, the method includes the following steps: Step 1: Model initialization: Initialize the parameters of the SARSA (state-action-reward-state-action) algorithm, including the Q matrix, state, action, learning rate, discount factor, gradient update step size, etc.

[0033] Step 2: Get the current antenna reception data according to the current moment action.

[0034] Step 3: Divide the space where the antenna is located into beam angles, and compare the received data in the L beam angle areas at the current moment with the threshold value to determine whether there is a target in the beam angle area, and calculate the state at the next moment.

[0035] Step 4: Since the target detection situation is related to the set threshold value, if the threshold value is too large, missed detection is likely to occur, while if the threshold value is too small, false detection is likely to occur; therefore, the false alarm rate and threshold value can be set in advance to detect the target in the antenna received data, and the target detection probability at the current moment is calculated according to the proportion of the target in all received data; based on the target detection probability at the current moment, the reward function at the next moment is calculated.

[0036] Step 5: Based on the current action and status, The greedy strategy selects the action with the largest Q function as the optimal action at the next moment, where the optimal action represents the beam angle area where the target may exist, thereby optimizing the antenna array beam pointing. According to the current action, state and the action, state and reward function at the next moment, the Q function is updated to enter the next cycle.

[0037] Step 6: Gradually optimize the decision estimation of the action so that the antenna array can adaptively optimize and adjust the beam pointing under changing environmental conditions to improve the probability of target detection.

[0038] Step 7: During the target detection process, the beam pointing is adaptively adjusted according to different environmental conditions, and the reward function feedback mechanism is used to cyclically update the strategy to improve the probability of target detection in complex electromagnetic environments.

[0039] The present invention realizes closed-loop adaptive optimization adjustment of antenna array beam pointing for target detection, decision-making, evaluation and beam pointing optimization by designing a framework of reinforcement learning decision-making and target detection probability evaluation, thereby improving the adaptability of target detection to complex dynamic electromagnetic environments.

[0040] The signal detection method using the SARSA algorithm in this embodiment is described in detail below, including: Obtain parameter information that affects the probability of target detection, including the currently received perception data, and set the detection threshold and false alarm rate, etc.

[0041] Initialize the parameters of the SARSA algorithm, including Q matrix, state, action, learning rate, discount factor, gradient update step size, etc.

[0042] Training and reasoning about antenna array beam pointing decisions are used to implement parameter adjustments for antenna array beam pointing. The SARSA algorithm updates the Q function based on the following rules:

[0043] in, Indicates the number of all beam angles where the target exists at time k; represents the beam angle of the possible target at time k; represents the learning rate, ; represents the impairment factor; represents the reward at time k+1.

[0044] (1) State space: First, let's review the target detection process. The current sensing area (the area where the antenna data is received) is divided into L discrete angle grids. At the current moment, we determine whether the signal within the L beam angle range is higher than the current threshold value. The threshold value is related to the false alarm rate. If it is higher than the threshold value, there is a target. If it is lower than the threshold value, there is no target. The specific introduction is as follows: For each angle grid, it can be divided into the following two cases: , Among them, assuming It contains only noise, that is ; express k Moment l The noise in the beam angle region; assuming It means that it contains signal and noise, that is ; express k Moment l The amplitude of the signal within the beam angle area; Represents the steering vector.

[0045] The relationship between the statistics of received data and the threshold value is as follows: , .

[0046] in, represents the statistics of the received data containing signal and noise, represents the statistics of the received data containing only noise, Indicates the threshold value.

[0047] For each angle grid, the statistics can be calculated using the following formula: .

[0048] in, represents statistics, express k Moment l All the received data within the beam angle area, H represents the covariance matrix of the noise, Represents an estimator of the covariance matrix of the noise.

[0049] The threshold value is related to the false alarm rate and can be calculated according to the following formula: .

[0050] in, represents the threshold value, represents the false alarm rate, represents the noise power, N Indicates the signal length, Represents the complementary error inverse function.

[0051] Define the total number of beam angle grids where there may be targets. If there is a target in the beam angle grid, it is set to 1, otherwise it is set to 0. The state at the current moment is defined as the total number of beam angle grids where there may be targets. Therefore, assuming that there are at most M targets in the perception area (the total number of beam angle grids where there may be targets is M), then the possible state set is .

[0052] (2) Action space: Assuming that there are at most M targets in the perception area, the cardinality of the action set is M. Therefore, the action at the current moment can be defined as: , in, ,express There is a target in the beam angle grid area. Indicates i There is a target in the beam angle grid area. Action space is the set of all possible beam angle grids, Indicates action.

[0053] (3) Strategy: Algorithm adopted Greedy strategy, represents the probability of randomly searching for a new action (beam pointing). The optimal action is The probability of is selected by maximizing the Q function. The specific action selection at the moment is as follows: .

[0054] in, represents the action at time k+1, represents the optimal action, Represents a random action.

[0055] This shows that if If set to 0, the antenna array beam pointing does not make any adjustments and always chooses the best action. If set to 1, the actions remain random, the antenna array beam pointing adjustment will not utilize previously learned information, and the Q function will not be saved.

[0056] (4) Reward function: The reward function defines the goal of the reinforcement learning problem, and therefore, the goal of the present invention is to maximize the reward function. The reward function can be divided into two parts, namely negative reward and positive reward. If there is a false alarm, the negative reward can be used as a penalty term. The positive reward is the probability of detecting the target within the beam angle grid range where the target exists at the current moment, and the negative reward is the probability of detecting the target within the remaining beam angle grid range at the current moment. The reward function can be expressed as:

[0057] in, express k+1 The reward function at the moment, express k The state of the moment, , Respectively representl , q The first term represents the sum of the detection probabilities of angle grids where targets may exist, that is, the target detection probability at the current moment; the second term represents the detection probabilities of the remaining angle grids where targets may not exist.

[0058] The present invention considers positive rewards and negative rewards in the reward function, takes the target detection probability within the beam angle grid range where the target exists at the current moment as the positive reward, takes the target detection probability within the remaining beam angle grid ranges at the current moment as the negative reward, adds a penalty term to reduce the impact of false alarms on decision-making, and feeds back the evaluation results to the reinforcement learning module, updates the algorithm strategy, and realizes closed-loop feedback of target detection probability.

[0059] The reinforcement learning of the present invention has a powerful strategy optimization capability. It optimizes and adjusts the beam pointing through reinforcement learning, and adopts the action execution and evaluation feedback mechanism to formulate the optimal strategy. It can simultaneously overcome the shortcomings of slow mechanical scanning speed and complex electronic scanning beam scanning control, and realize fast and accurate target detection in complex electromagnetic environments.

[0060] Embodiment 2 This embodiment provides a signal detection system based on reinforcement learning.

[0061] A signal detection system based on reinforcement learning, comprising: A data acquisition module is configured to: initialize relevant parameters of the SARSA algorithm, take the beam angle area where the target may exist as the current action, obtain the current receiving perception data of the current action, and determine the threshold value; The parameter calculation module is configured to: divide the currently received perception data into L discrete angle grids, determine whether the signals in the L beam angle grids are higher than the threshold value at the current moment, count the total number of angle grids where all targets may exist at the current moment as the current state, and calculate the state at the next moment; calculate the target detection probability according to the proportion of all targets in the currently received perception data, and then calculate the reward function at the next moment, and update Greedy strategy; The iterative output module is configured to: use the updated The greedy strategy selects the action with the largest Q function as the action at the next moment, combines the state at the next moment, iterates the SARSA algorithm, and outputs the beam pointing when the set conditions are met.

[0062] In some embodiments, the parameter calculation module is further configured to: determine one by one whether the signals in the L beam angle grids are higher than the threshold value at the current moment; if so, there may be a target; otherwise, there is no target, and count the total number of angle grids where there may be targets at the current moment.

[0063] In some embodiments, the threshold value is expressed by the following formula: .

[0064] in, represents the threshold value, represents the false alarm rate, represents the noise power, N Indicates the signal length, Represents the complementary error inverse function.

[0065] In some embodiments, the target detection probability is calculated, and then the reward function at the next moment is calculated; it is expressed by the following formula:

[0066] in, express k+1 The reward function at the moment, express k The state of the moment, , Respectively represent l , q Target detection probability in each beam angle region; The sum of the detection probabilities of the angle grids where the target may exist is the target detection probability at the current moment; Represents the detection probability of the remaining angle grids where the target may not exist.

[0067] In some embodiments, the parameter calculation module is further configured to: use the target detection probability at the current moment as a positive reward, and use the detection probability of the remaining angle grids where the target may not exist as a negative reward, and update Greedy strategy,realizes closed-loop feedback of target detection probability.

[0068] In some embodiments, the total number of angle grids where the target may exist at the current moment is counted, and is expressed by the following formula: .

[0069] in, represents statistics, represents the steering vector, express k Moment lAll the received data within the beam angle area, H represents the covariance matrix of the noise, Represents an estimator of the covariance matrix of the noise.

[0070] The present invention utilizes a reinforcement learning algorithm to perform cyclic real-time optimization and adjustment of the antenna array beam pointing, so that target detection can dynamically adapt to the complex electromagnetic environment where multiple dynamic targets exist, significantly improving the target detection probability.

[0071] Embodiment 3 This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the signal detection method based on reinforcement learning as described in the first embodiment are implemented.

[0072] Embodiment 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the signal detection method based on reinforcement learning as described in the first embodiment are implemented.

[0073] Embodiment 5 This embodiment provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of the signal detection method based on reinforcement learning described in the first embodiment.

[0074] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0075] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0076] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0078] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0079] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A signal detection method based on reinforcement learning, characterized in that: include: Initialize the relevant parameters of the SARSA algorithm, take the beam angle area where the target may exist as the current action, obtain the current receiving perception data of the current action, and determine the threshold value; The currently received perception data is divided into L discrete angle grids. At the current moment, the signal in each of the L beam angle grids is determined to be higher than the threshold value. The total number of angle grids where all possible targets may exist is counted as the current state, and the state at the next moment is calculated. According to the proportion of all targets in the currently received perception data, the target detection probability is calculated, and then the reward function at the next moment is calculated to update Greedy strategy; According to the current action and status, the updated The greedy strategy selects the action with the largest Q function as the action at the next moment, combines the state at the next moment, iterates the SARSA algorithm, and outputs the beam pointing when the set conditions are met.

2. The signal detection method based on reinforcement learning according to claim 1, characterized in that: The method comprises: judging one by one whether the signals in the L beam angle grids are higher than the threshold value at the current moment, and counting the total number of angle grids where the target may exist at the current moment; the method comprises: judging one by one whether the signals in the L beam angle grids are higher than the threshold value at the current moment, if so, there may be a target, otherwise, there is no target, and counting the total number of angle grids where the target may exist at the current moment.

3. The signal detection method based on reinforcement learning according to claim 1, characterized in that: The threshold value is expressed by the following formula: in, represents the threshold value, represents the false alarm rate, represents the noise power, N Indicates the signal length, Represents the complementary error inverse function.

4. The signal detection method based on reinforcement learning according to claim 1, characterized in that: The target detection probability is calculated, and then the reward function at the next moment is calculated; It is expressed by the following formula: in, express k+1 The reward function at the moment, express k The state of the moment, , Respectively represent l , q Target detection probability in each beam angle region; The sum of the detection probabilities of the angle grids where the target may exist is the target detection probability at the current moment; Represents the detection probability of the remaining angle grids where the target may not exist.

5. The signal detection method based on reinforcement learning according to claim 4, characterized in that: Said update The greedy strategy process includes: taking the target detection probability at the current moment as a positive reward, taking the detection probability of the remaining angle grids where the target may not exist as a negative reward, and updating Greedy strategy,realizes closed-loop feedback of target detection probability.

6. The signal detection method based on reinforcement learning according to claim 1, characterized in that: The total number of angle grids where all possible targets may exist at the current moment is calculated using the following formula: in, represents statistics, represents the steering vector, express k Moment l All the received data within the beam angle area, H represents the covariance matrix of the noise, Represents an estimator of the covariance matrix of the noise.

7. A signal detection system based on reinforcement learning, characterized in that: include: A data acquisition module is configured to: initialize relevant parameters of the SARSA algorithm, take the beam angle area where the target may exist as the current action, obtain the current receiving perception data of the current action, and determine the threshold value; The parameter calculation module is configured to: divide the currently received perception data into L discrete angle grids, determine whether the signals in the L beam angle grids are higher than the threshold value at the current moment, count the total number of angle grids where all targets may exist at the current moment as the current state, and calculate the state at the next moment; calculate the target detection probability according to the proportion of all targets in the currently received perception data, and then calculate the reward function at the next moment, and update Greedy strategy; The iterative output module is configured to: use the updated The greedy strategy selects the action with the largest Q function as the action at the next moment, combines the state at the next moment, iterates the SARSA algorithm, and outputs the beam pointing when the set conditions are met.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the signal detection method based on reinforcement learning as described in any one of claims 1 to 6 are implemented.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the signal detection method based on reinforcement learning as described in any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps in the signal detection method based on reinforcement learning as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • MIMO radar multi-target detection method and system based on strong target limitation

    CN119165461A

  • Passive detection resource parameter dynamic optimization system and method based on frequency spectrum monitoring

    CN119375836A

  • Apparatus, system and method of generating radar perception data

    US20200225317A1

  • Electronic device, method for controlling electronic device, and program

    WO2024058225A1

Cited By

  • Permanent magnet synchronous motor test method, device and system

    CN120629934A