Maritime antenna star finding method and device, and storage medium

By applying reinforcement learning model in maritime antennas and dynamically switching satellite tracking modules, the problem of unstable performance of automatic tracking satellite algorithms during environmental changes is solved, and the stability and reliability of signal transmission are improved.

CN120030890APending Publication Date: 2025-05-23广州肯赛特通信科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510106502.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

During the ship's navigation, the algorithm of automatically tracking satellites is unstable when the environment changes, resulting in inaccurate antenna beam direction and affecting signal transmission quality.

Method used

The reinforcement learning model is adopted to optimize the performance indicators of maritime antennas through the interaction between the agent and the environment, and dynamically switch different satellite tracking modules to adapt to the changes in ships and antennas.

Benefits of technology

It improves the performance stability of the automatic tracking satellite module, enhances the degree of adaptation with ships and their maritime antennas, gives full play to the advantages of different modules, and improves the reliability of signal transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030890A_ABST
    Figure CN120030890A_ABST
Patent Text Reader

Abstract

The invention discloses a satellite finding method and device for a maritime antenna and a storage medium. The method comprises the following steps: loading a plurality of satellite tracking modules configured for the maritime antenna; in the operation process of the maritime antenna, one satellite tracking module is selected as a maritime antenna tracking satellite to serve as a first target tracking module; constructing a reinforcement learning model; starting the reinforcement learning model, and selecting one satellite tracking module as a maritime antenna tracking satellite according to the ship to serve as a second target tracking module by taking optimization of the performance index of the maritime antenna as a target; converting the attitude output by the first target tracking module into an attitude suitable for being used by the second target tracking module, and taking the attitude as a target state; judging whether the target state meets a switching condition or not; and if yes, switching from the first target tracking module to the second target tracking module as a maritime antenna tracking satellite. According to the embodiment, the advantages of different automatic tracking satellite modules are exerted, so that the stability of the automatic tracking satellite performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of maritime antennas, and in particular to a satellite search method, equipment and storage medium for a maritime antenna. Background Art

[0002] During the voyage of the ship, the maritime antenna on the ship automatically tracks the satellite to ensure that the antenna beam points to the satellite, thereby realizing communication with other users through the satellite, obtaining meteorological information, navigation data, etc., and providing support for the normal operation of the ship.

[0003] At present, there are many automatic satellite tracking algorithms, each of which has its own advantages and disadvantages. Ships use one of the automatic satellite tracking algorithms.

[0004] However, the surrounding environment of a ship is constantly changing while it is sailing, which makes the automatic satellite tracking algorithm perform better in some cases and worse in some cases, resulting in poor performance stability. Summary of the invention

[0005] In view of this, the present invention provides a satellite-finding method, device and storage medium for a maritime antenna, so as to improve the stability of the automatic satellite tracking performance of the maritime antenna on a ship.

[0006] A first aspect of the present invention provides a satellite search method for a maritime antenna, which is applied to an antenna controller of a maritime antenna on a ship, and the method comprises:

[0007] Loading and configuring a plurality of satellite tracking modules for the maritime antenna;

[0008] During the operation of the maritime antenna, one of the satellite tracking modules is selected as the maritime antenna tracking satellite as a first target tracking module;

[0009] Constructing a reinforcement learning model; the intelligent agent in the reinforcement learning model is the antenna controller, the environment is the ship, and the action is selecting one of the satellite tracking modules as the maritime antenna tracking satellite;

[0010] Starting the reinforcement learning model, with the goal of optimizing the performance index of the maritime antenna, selecting one of the satellite tracking modules as the maritime antenna tracking satellite as the second target tracking module according to the ship;

[0011] Converting the posture output by the first target tracking module into a posture suitable for use by the second target tracking module as a target state;

[0012] Determine whether the target state satisfies a switching condition; if so, switch from the first target tracking module to the second target tracking module to track the satellite for the maritime antenna.

[0013] A second aspect of the present invention provides a satellite-finding device for a maritime antenna, which is applied to an antenna controller of a maritime antenna on a ship, and the device comprises:

[0014] A satellite tracking module loading module, used for loading various satellite tracking modules configured for the maritime antenna;

[0015] A first target tracking module selection module, used for selecting one of the satellite tracking modules as a tracking satellite for the maritime antenna as a first target tracking module during the operation of the maritime antenna;

[0016] A reinforcement learning model construction module is used to construct a reinforcement learning model; the intelligent agent in the reinforcement learning model is the antenna controller, the environment is the ship, and the action is to select one of the satellite tracking modules as the maritime antenna tracking satellite;

[0017] A second target tracking module selection module is used to start the reinforcement learning model, with the goal of optimizing the performance index of the maritime antenna, and select one of the satellite tracking modules as the maritime antenna tracking satellite according to the ship as the second target tracking module;

[0018] a target state conversion module, configured to convert the posture output by the first target tracking module into a posture suitable for use by the second target tracking module as a target state;

[0019] A switching condition judgment module is used to judge whether the target state meets the switching condition; if so, the tracking switching module is called;

[0020] A tracking switching module is used to switch from the first target tracking module to the second target tracking module to track satellites for the maritime antenna.

[0021] A third aspect of the present invention provides an electronic device, the electronic device comprising:

[0022] at least one processor; and

[0023] a memory communicatively connected to the at least one processor; wherein,

[0024] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the satellite search method for the maritime antenna as described in the first aspect above.

[0025] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the satellite search method for a maritime antenna as described in the first aspect above.

[0026] A fifth aspect of the present invention provides a computer program product, the computer program product comprising a computer program, and when the computer program is executed by a processor, the satellite search method for a maritime antenna as described in the first aspect above is implemented.

[0027] In this embodiment, multiple satellite tracking modules are loaded for configuring the maritime antenna; during the operation of the maritime antenna, one of the satellite tracking modules is selected as the maritime antenna tracking satellite as the first target tracking module; a reinforcement learning model is constructed; the agent in the reinforcement learning model is the antenna controller, the environment is the ship, and the action is to select one of the satellite tracking modules as the maritime antenna tracking satellite; the reinforcement learning model is started, with the goal of optimizing the performance index of the maritime antenna, and one of the satellite tracking modules is selected as the maritime antenna tracking satellite according to the ship as the second target tracking module; the posture output by the first target tracking module is converted into a posture suitable for use by the second target tracking module as the target state; it is determined whether the target state meets the switching condition; if so, the first target tracking module is switched to the second target tracking module as the maritime antenna tracking satellite. This embodiment adjusts the automatic tracking satellite module according to the situation of the ship and its maritime antenna, improves the degree of adaptation of the automatic tracking satellite module to the ship and its maritime antenna, and gives full play to the advantages of different automatic tracking satellite modules, thereby improving the stability of the automatic tracking satellite performance.

[0028] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0030] Figure 1 This is a flow chart of a satellite search method for a maritime antenna provided in Embodiment 1 of the present invention.

[0031] Figure 2 It is a structural diagram of a Q learning network provided in Example 1 of the present invention.

[0032] Figure 3It is a structural schematic diagram of a satellite-seeking device for a maritime antenna provided in Embodiment 2 of the present invention.

[0033] Figure 4 It is a structural schematic diagram of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can cover sequential implementations other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0036] Embodiment 1

[0037] See also Figure 1 , shows a flow chart of a satellite search method for a maritime antenna provided by Embodiment 1 of the present invention. The method can be executed by a satellite search device for a maritime antenna. The satellite search device for a maritime antenna can be implemented in the form of hardware and / or software. The satellite search device for a maritime antenna can be configured in an electronic device. The electronic device includes an antenna controller for a maritime antenna on a ship. In this case, the method can be applied to an antenna controller for a maritime antenna on a ship. Figure 1 As shown, the method includes:

[0038] Step 101: Load and configure various satellite tracking modules for the maritime antenna.

[0039] The antenna controller is configured with a satellite tracking algorithm library, which records a variety of satellite tracking modules. The satellite tracking module refers to the code compiled for the satellite tracking algorithm of the ship's maritime antenna, which is implemented in the form of plug-ins, library files, etc., for easy hot plugging.

[0040] In the specific implementation, one of the key technologies of various satellite tracking modules lies in the antenna tracking method and stable compensation, which usually includes the following types and their improved algorithms:

[0041] 1. Program tracking

[0042] According to the location information of the satellite navigation system, the attitude information of the maritime antenna and the target location information, the memory tracking is solved through attitude calculation to calculate the rotation angles of the azimuth and pitch motors, and then drive the execution motor to move within the predetermined range.

[0043] 2. Memory Tracking

[0044] Since the satellite's drift has a certain repeatability, it can be tracked based on historical record data in a short period of time.

[0045] 3. Orbit prediction and tracking

[0046] Within a certain period of time, the azimuth, pitch angle, time and the direction of change of azimuth and pitch output by comparing the previous tracking data at the point where the signal is maximum are recorded. Subsequently, the signal is directly transferred to the point where the historical memory has the maximum signal, and the signal level is sampled. Adjustment is made according to the original direction of movement, and the signal level is compared to control the maritime antenna to transfer to the point where the signal is strongest.

[0047] 4. Cone scanning tracking system

[0048] The feed horn is moved in a circle around the antenna's symmetry axis, or the sub-surface is tilted and rotated, so that the maritime antenna beam rotates in a conical shape.

[0049] 5. Single pulse tracking

[0050] The direction in which the antenna beam deviates from the satellite can be determined within a pulse interval, and the servo system can be driven to quickly align the maritime antenna with the satellite.

[0051] 6. Step tracking

[0052] At a certain time interval, the maritime antenna is rotated at a small angle in the azimuth or elevation plane, and the increase or decrease of the receiving level is judged within an appropriate integration time. If the receiving level increases, the antenna continues to rotate a small angle in the original direction. If the receiving level decreases, the antenna rotates in the opposite direction. The elevation direction and the azimuth direction are alternated in turn, so that the maritime antenna beam is gradually aligned with the satellite.

[0053] During the operation of the ship, the satellite tracking algorithm library can be loaded to facilitate switching between multiple satellite tracking modules.

[0054] Step 102: During the operation of the maritime antenna, one of the satellite tracking modules is selected as a maritime antenna tracking satellite as a first target tracking module.

[0055] During the operation of the maritime antenna, one of the satellite tracking modules can be selected from the satellite tracking algorithm library by using memory, performance, etc. as the maritime antenna tracking satellite. At this time, the satellite tracking module is recorded as the first target tracking module.

[0056] Among them, memory refers to selecting the satellite tracking module used last by the maritime antenna, and performance refers to selecting the satellite tracking module with the highest comprehensive performance.

[0057] Step 103: Build a reinforcement learning model.

[0058] In this embodiment, satellite tracking of maritime antennas on ships can be modeled to obtain a reinforcement learning model, wherein the reinforcement learning model is usually described using a Markov decision process (MDP), that is, the machine is in an environment, and each state is the machine's perception of the current environment; the machine affects the environment through actions, and when the machine performs an action, the environment will be transferred to another state with a certain probability; at the same time, the environment will feedback an incentive to the machine according to the potential incentive function.

[0059] In practical applications, the reinforcement learning model contains four basic elements: agent, environment, action and reward.

[0060] Among them, the intelligent agent can perceive the state of the environment, and according to the reward provided by the environment, select an appropriate action through learning to maximize the long-term reward. That is, if an action brings a positive reward from the environment, then this action will be strengthened, if an action brings a negative reward from the environment, then this action will be weakened.

[0061] The environment receives a series of actions performed by the agent, evaluates the quality of the actions, and converts them into a quantifiable reward to feed back to the agent. At the same time, the environment also provides the agent with its state.

[0062] Reward is a quantifiable scalar feedback signal provided by the environment to the agent, which is used to evaluate the quality of the action performed by the agent at a certain time. The goal of the agent to perform a series of actions is to maximize the future cumulative reward.

[0063] The state contains the information that the agent uses to select actions.

[0064] In this embodiment, the agent Agent in the reinforcement learning model is an antenna controller, the environment Environment is a ship (including a maritime antenna), and the action Agent is to select one of the satellite tracking modules as a maritime antenna tracking satellite.

[0065] Step 104: start the reinforcement learning model, with the goal of optimizing the performance indicators of the maritime antenna, and select one of the satellite tracking modules as the maritime antenna tracking satellite according to the ship as the second target tracking module.

[0066] In practical applications, the reinforcement learning model is started and run. The antenna controller acts as an intelligent agent and receives the operation information of the ship (including the maritime antenna) as the state. It executes the action of selecting one of the satellite tracking modules as the maritime antenna tracking satellite, and detects the performance index of the maritime antenna as the incentive Reward, so that the performance index of the maritime antenna is the optimal incentive Reward. At this time, the selected satellite tracking module is recorded as the second target tracking module.

[0067] In one embodiment of the present invention, step 104 may include the following steps:

[0068] Step 1041: At each time, record various meteorological information of the vessel's environment to obtain a meteorological sequence.

[0069] In this embodiment, a variety of meteorological information of the ship's environment at various times can be read from a meteorological detector on the ship to obtain a meteorological sequence.

[0070] Generally speaking, wind will cause the antenna to vibrate or even sway, causing the direction of the maritime antenna to shift, reducing the alignment accuracy between the maritime antenna and the satellite, deteriorating the signal transmission quality, and causing problems such as signal attenuation and increased bit error rate. Therefore, meteorological information can include wind force levels.

[0071] Rain can scatter and absorb radio signals, especially in high frequency bands, where signal attenuation is more obvious, affecting communication quality. In addition, rain can form a water film on the surface of the maritime antenna, changing the electrical properties of the maritime antenna, causing changes in the reflection coefficient and standing wave ratio of the maritime antenna, affecting the antenna's radiation efficiency and signal reception capability. Therefore, meteorological information can include rainfall levels.

[0072] Of course, the above-mentioned meteorological information is only used as an example. When implementing this embodiment, other meteorological information can be set according to actual conditions, such as fog level, snow level, lightning level, etc., and this embodiment does not limit this. In addition, in addition to the above-mentioned meteorological information, those skilled in the art can also use other meteorological information according to actual needs, and this embodiment does not limit this.

[0073] Step 1042: At each moment, record the attitude data of the ship to obtain an attitude sequence.

[0074] In this embodiment, when the ship rolls, the antenna will swing left and right in the horizontal direction, causing the angle between the antenna and the satellite to change and the antenna pointing to deviate. When the ship pitches, the distance between the antenna and the satellite and the signal propagation path are affected. Therefore, the attitude data of the ship recorded at each moment can be read from the attitude measurement system (APMS) on the ship to obtain an attitude sequence.

[0075] Step 1043: At each moment, record the change amplitude of the maritime antenna in attitude to obtain an amplitude sequence.

[0076] In this embodiment, the attitude of the maritime antenna at the current moment may be subtracted from the attitude of the previous moment to obtain the change amplitude of the attitude of the maritime antenna, and the change amplitudes may be accumulated at multiple moments to obtain an amplitude sequence.

[0077] Step 1044: In the dimension of uncertainty, the weather sequence, attitude sequence and amplitude sequence are input into the Q learning network in the reinforcement learning model to learn various satellite tracking modules as the first Q value of the maritime antenna tracking satellite.

[0078] In practical applications, the weather of the ship's environment at sea and the ship's own attitude have certain regularity in a long period of time, and there is a certain uncertainty in a short time. Tracking the maritime antenna is a short-term operation. Therefore, considering the short-term uncertainty, the weather sequence, attitude sequence and amplitude sequence are input into the Q-learning network (Q-Learning) in the reinforcement learning model to learn various satellite tracking modules as the first Q value of the maritime antenna tracking satellite.

[0079] In one embodiment of the present invention, Figure 2As shown, the Q-learning network (Q-Learning) in the reinforcement learning model is a deep learning model, especially a recurrent neural network (RNN), which includes a first long short-term memory network LSTM_1, a second long short-term memory network LSTM_2, a third long short-term memory network LSTM_3, a fourth long short-term memory network LSTM_4, a first fully connected layer FC_1, a second fully connected layer FC_2 and a third fully connected layer FC_3.

[0080] Initially, training the Q-learning network is divided into two stages:

[0081] The first stage is to train the first long short-term memory network LSTM_1, the second long short-term memory network LSTM_2, and the third long short-term memory network LSTM_3.

[0082] At this time, a fully connected layer and other structures are cascaded after the first long short-term memory network LSTM_1, the second long short-term memory network LSTM_2, and the third long short-term memory network LSTM_3 as a classifier, that is, the features output by the third long short-term memory network LSTM_3 are input into the classifier, and the classifier outputs uncertainty. The meteorological sequence and the posture sequence are used as samples, and the uncertainty is annotated as the label Label. The first long short-term memory network LSTM_1, the second long short-term memory network LSTM_2, the third long short-term memory network LSTM_3 and the classifier are supervised trained, and the features of the first long short-term memory network LSTM_1, the second long short-term memory network LSTM_2, the third long short-term memory network LSTM_3 and the classifier are updated, so that the features output by the first long short-term memory network LSTM_1, the second long short-term memory network LSTM_2, and the third long short-term memory network LSTM_3 can be used to identify uncertainty. When the training is completed, the classifier is discarded.

[0083] The second stage is to train the fourth long short-term memory network LSTM_4, the first fully connected layer FC_1, the second fully connected layer FC_2 and the third fully connected layer FC_3.

[0084] At this time, taking meteorological sequence, attitude sequence and amplitude sequence as samples, various satellite tracking modules are marked with Q values ​​of maritime antenna tracking satellites as label Labels, and supervised training is performed on the first long short-term memory network LSTM_1, the second long short-term memory network LSTM_2, the third long short-term memory network LSTM_3, the fourth long short-term memory network LSTM_4, the first fully connected layer FC_1, the second fully connected layer FC_2 and the third fully connected layer FC_3. The parameters of the first long short-term memory network LSTM_1, the second long short-term memory network LSTM_2 and the third long short-term memory network LSTM_3 are kept unchanged, and the parameters of the fourth long short-term memory network LSTM_4, the first fully connected layer FC_1, the second fully connected layer FC_2 and the third fully connected layer FC_3 are updated.

[0085] In this embodiment, the meteorological sequence can be input into the first long short-term memory network LSTM_1 to extract meteorological features representing uncertainty.

[0086] The posture sequence is input into the second long short-term memory network LSTM_2 to extract posture features representing uncertainty.

[0087] Use Concat and other functions to merge meteorological features and attitude features into the first uncertainty feature.

[0088] The first uncertainty feature is input into the third long short-term memory network LSTM_3 to extract the second uncertainty feature.

[0089] The amplitude sequence is input into the fourth long short-term memory network LSTM_4 to extract amplitude features.

[0090] Use functions such as Concat to merge the second uncertainty feature and the amplitude feature into the first multimodal feature.

[0091] The first multimodal features are input into the first fully connected layer FC_1 and mapped into the second multimodal features.

[0092] The second multimodal features are input into the second fully connected layer FC_2 and mapped into the third multimodal features.

[0093] The third multimodal features are input into the third fully connected layer FC_3 and mapped into the fourth multimodal features.

[0094] The fourth multimodal feature is activated using Sigmoid and other functions to obtain first Q values ​​of various satellite tracking modules for maritime antenna tracking satellites.

[0095] Step 1045: Select one of the satellite tracking modules according to the first Q value as the second target tracking module to optimize the performance index of the maritime antenna.

[0096] In this embodiment, the ∈-greedy method or the like can be used to select one of the satellite tracking modules according to the first Q value as the second target tracking module, thereby optimizing the performance index of the maritime antenna.

[0097] In the ∈-greedy method, there is a probability of ∈ to select the satellite tracking module with the largest first Q value, and there is a probability of (1-∈) to randomly select any satellite tracking module.

[0098] Step 1046: After switching from the first target tracking module to the second target tracking module, calculate excitation for the second target tracking module according to the amplitude sequence.

[0099] When switching from the first target tracking module to the second target tracking module later, the incentive Reward can be calculated for the second target tracking module according to the amplitude sequence at this time.

[0100] Exemplarily, the incentive is:

[0101]

[0102] Among them, R is motivation, Strength t is the signal strength of the maritime antenna at time t, Performance t is the performance parameter of the maritime antenna at time t (such as power consumption, resource utilization, etc.), Amplitude t is the attitude change amplitude of the maritime antenna at time t, T is the set of time, and α, β, γ and δ are all hyperparameters.

[0103] In this example, the signal strength is taken as the main factor, and the regularization terms are set from two aspects: performance parameters and attitude variation range. When the performance parameters and attitude variation range gradually increase, the regularization terms increase rapidly, thereby constraining the excitation.

[0104] Step 1047: In the dimension of uncertainty, the weather sequence, attitude sequence and amplitude sequence are input into the Q learning network in the reinforcement learning model to learn various satellite tracking modules as the second Q value of the maritime antenna tracking satellite.

[0105] When switching from the first target tracking module to the second target tracking module, the meteorological sequence, attitude sequence and amplitude sequence are input into the Q-learning network (Q-Learning) in the reinforcement learning model to learn various satellite tracking modules as the second Q value of the maritime antenna tracking satellite.

[0106] In one embodiment of the present invention, Figure 2As shown, the Q-learning network (Q-Learning) in the reinforcement learning model is a deep learning model, especially a recurrent neural network (RNN), which includes a first long short-term memory network LSTM_1, a second long short-term memory network LSTM_2, a third long short-term memory network LSTM_3, a fourth long short-term memory network LSTM_4, a first fully connected layer FC_1, a second fully connected layer FC_2 and a third fully connected layer FC_3.

[0107] In this embodiment, the meteorological sequence can be input into the first long short-term memory network LSTM_1 to extract meteorological features representing uncertainty.

[0108] The posture sequence is input into the second long short-term memory network LSTM_2 to extract posture features representing uncertainty.

[0109] Use Concat and other functions to merge meteorological features and attitude features into the first uncertainty feature.

[0110] The first uncertainty feature is input into the third long short-term memory network LSTM_3 to extract the second uncertainty feature.

[0111] The amplitude sequence is input into the fourth long short-term memory network LSTM_4 to extract amplitude features.

[0112] Use functions such as Concat to merge the second uncertainty feature and the amplitude feature into the first multimodal feature.

[0113] The first multimodal features are input into the first fully connected layer FC_1 and mapped into the second multimodal features.

[0114] The second multimodal features are input into the second fully connected layer FC_2 and mapped into the third multimodal features.

[0115] The third multimodal features are input into the third fully connected layer FC_3 and mapped into the fourth multimodal features.

[0116] The fourth multimodal feature is activated using Sigmoid and other functions to obtain second Q values ​​for various satellite tracking modules to track satellites for maritime antennas.

[0117] Step 1048: Update the Q learning network in the reinforcement learning model according to the incentive, the first Q value, and the second Q value.

[0118] In this embodiment, the incentive, the first Q value and the second Q value can be used to construct a loss function, and the loss function is used to continuously iteratively update the Q learning network in the reinforcement learning model (especially the fourth long short-term memory network LSTM_4, the first fully connected layer FC_1, the second fully connected layer FC_2 and the third fully connected layer FC_3).

[0119] Among them, the loss function is expressed as:

[0120] Loss=Qs t ,a t +α[r t +γmax a Q t+1 ,a t+1 -Q(s t ,a t )]

[0121] Among them, Loss is the loss value, Q(s t ,a t ) is the first Q value at time t, α is the learning rate, r is the decay coefficient, Q(s t+1 ,a t+1 ) is the second Q value at time t+1, r t is the excitation at time t, max a Indicates selecting the largest second Q value at the next moment.

[0122] Step 105: Convert the posture output by the first target tracking module into a posture suitable for use by the second target tracking module as the target state.

[0123] If the second target tracking module is different from the first target tracking module, considering the differences in algorithms such as gain among different satellite tracking modules, the posture output by the first target tracking module can be converted into a posture suitable for use by the second target tracking module as the target state, which facilitates the evaluation of the second target tracking module and thus measures the cost of switching.

[0124] In a specific implementation, the difference between the posture of the maritime antenna at the current moment and the posture of the maritime antenna output by the first target tracking module at the next moment can be calculated as the antenna rotation amplitude.

[0125] Query the gradient of the second target tracking module when updating the maritime antenna attitude.

[0126] The projection of the antenna rotation amplitude on the gradient is set as a posture suitable for use by the second target tracking module as a candidate state.

[0127] For example, since a vector can be decomposed into a weighted sum between two orthogonal projection vectors, the candidate state can be expressed as:

[0128]

[0129] Among them, Pose is the selected state, Turn is the antenna rotation amplitude, T is the transpose, G is the gradient, and M is the conversion matrix between the first target tracking module and the second target tracking module.

[0130] Use particle filtering, mean filtering and other algorithms to filter multiple candidate states, filter out noise and obtain the target state.

[0131] Step 106 , determine whether the target state meets the switching condition; if so, execute step 107 .

[0132] In this embodiment, the target state of the second target tracking module may be analyzed to evaluate whether the cost of switching the second target tracking module meets the switching condition, thereby reducing the fluctuation of the switching as much as possible.

[0133] In a specific implementation, a first antenna instance and a second antenna instance constructed for a maritime antenna using a master-slave relationship are queried; wherein the first antenna instance is a master instance and the second antenna instance is a slave instance, the master instance is an instance that actually controls the maritime antenna, and the slave instance is an instance that is a backup instance that controls the maritime antenna, and the master instance and the slave instance share the operating data of the maritime antenna.

[0134] The first antenna instance is connected to the first target tracking module, and the second antenna instance is in an idle state.

[0135] At this time, the second antenna instance is connected to the first target tracking module to use the second target tracking module to track satellites for the maritime antenna. Using the second target tracking module to track satellites for the maritime antenna is a simulation process, not an actual control process.

[0136] The second target tracking module uses the operation data of the maritime antenna to perform calculations and outputs the posture of the second antenna instance at the next moment. At this time, the posture of the second antenna instance output by the second target tracking module at the next moment can be queried in the specified memory area as a reference state.

[0137] The absolute value of the difference between the target state and the reference state is taken to obtain the state deviation.

[0138] If the state deviation is less than or equal to a preset threshold, it means that the cost of switching to the second target tracking module is small, and it is determined that the target state meets the switching condition.

[0139] If the state deviation is greater than a preset threshold, it means that the cost of switching to the second target tracking module is relatively high, and it is determined that the target state does not meet the switching condition.

[0140] Step 107: Switch from the first target tracking module to the second target tracking module to track the satellite for the maritime antenna.

[0141] When the target state meets the switching condition, it is possible to switch from the first target tracking module to the second target tracking module to continue tracking satellites for the maritime antenna.

[0142] In a specific implementation, the first antenna instance is switched from a master instance to a slave instance, and the first antenna instance is disconnected from the first target tracking module to stop using the first target tracking module to track satellites for the maritime antenna.

[0143] Switch the second antenna instance from slave to master to start tracking satellites for the maritime antenna using the second target tracking module.

[0144] The second antenna instance continues to simulate a maritime antenna tracking satellite and can quickly switch to the main instance when the cost is low, reducing the impact of the switch.

[0145] In this embodiment, multiple satellite tracking modules are loaded for configuring the maritime antenna; during the operation of the maritime antenna, one of the satellite tracking modules is selected as the maritime antenna tracking satellite as the first target tracking module; a reinforcement learning model is constructed; the agent in the reinforcement learning model is the antenna controller, the environment is the ship, and the action is to select one of the satellite tracking modules as the maritime antenna tracking satellite; the reinforcement learning model is started, with the goal of optimizing the performance index of the maritime antenna, and one of the satellite tracking modules is selected as the maritime antenna tracking satellite according to the ship as the second target tracking module; the posture output by the first target tracking module is converted into a posture suitable for use by the second target tracking module as the target state; it is determined whether the target state meets the switching condition; if so, the first target tracking module is switched to the second target tracking module as the maritime antenna tracking satellite. This embodiment adjusts the automatic tracking satellite module according to the situation of the ship and its maritime antenna, improves the degree of adaptation of the automatic tracking satellite module to the ship and its maritime antenna, and gives full play to the advantages of different automatic tracking satellite modules, thereby improving the stability of the automatic tracking satellite performance.

[0146] Embodiment 2

[0147] See also Figure 3 , showing a schematic diagram of the structure of a satellite search device for a maritime antenna provided in Embodiment 3 of the present invention. An antenna controller for a maritime antenna applied to a ship, such as Figure 3 As shown, the device comprises:

[0148] A satellite tracking module loading module 301 is used to load various satellite tracking modules configured for the maritime antenna;

[0149] A first target tracking module selection module 302 is used to select one of the satellite tracking modules as a first target tracking module during the operation of the maritime antenna as a tracking satellite for the maritime antenna;

[0150] A reinforcement learning model construction module 303 is used to construct a reinforcement learning model; the agent in the reinforcement learning model is the antenna controller, the environment is the ship, and the action is to select one of the satellite tracking modules as the maritime antenna tracking satellite;

[0151] A second target tracking module selection module 304 is used to start the reinforcement learning model, with the goal of optimizing the performance index of the maritime antenna, and select one of the satellite tracking modules as the maritime antenna tracking satellite according to the ship as the second target tracking module;

[0152] A target state conversion module 305, configured to convert the posture output by the first target tracking module into a posture suitable for use by the second target tracking module as a target state;

[0153] The switching condition judgment module 306 is used to judge whether the target state meets the switching condition; if so, the tracking switching module is called;

[0154] The tracking switching module 307 is used to switch from the first target tracking module to the second target tracking module to track satellites for the maritime antenna.

[0155] In one embodiment of the present invention, the second target tracking module selection module 304 includes:

[0156] A weather sequence recording module is used to record various weather information of the environment in which the ship is located at various times to obtain a weather sequence;

[0157] An attitude sequence recording module is used to record the attitude data of the ship at each moment to obtain an attitude sequence;

[0158] An amplitude sequence recording module is used to record the amplitude of the change in attitude of the maritime antenna at each moment to obtain an amplitude sequence;

[0159] A first Q value calculation module is used to input the meteorological sequence, the attitude sequence and the amplitude sequence into the Q learning network in the reinforcement learning model to learn various first Q values ​​of the satellite tracking module for the maritime antenna to track the satellite in the dimension of uncertainty;

[0160] A satellite tracking module selection module, used for selecting one of the satellite tracking modules as a second target tracking module according to the first Q value, so as to optimize the performance index of the maritime antenna;

[0161] an excitation calculation module, configured to calculate an excitation for the second target tracking module according to the amplitude sequence after switching from the first target tracking module to the second target tracking module;

[0162] A second Q value calculation module is used to input the meteorological sequence, the attitude sequence and the amplitude sequence into the Q learning network in the reinforcement learning model to learn various second Q values ​​of the satellite tracking module for the maritime antenna tracking satellite in the dimension of uncertainty;

[0163] A Q learning network updating module is used to update the Q learning network in the reinforcement learning model according to the incentive, the first Q value and the second Q value.

[0164] In one embodiment of the present invention, the Q learning network in the reinforcement learning model includes a first long short-term memory network, a second long short-term memory network, a third long short-term memory network, a fourth long short-term memory network, a first fully connected layer, a second fully connected layer and a third fully connected layer;

[0165] The first Q value calculation module is also used for:

[0166] Inputting the meteorological sequence into the first long short-term memory network to extract meteorological features representing uncertainty;

[0167] Inputting the posture sequence into the second long short-term memory network to extract posture features representing uncertainty;

[0168] fusing the meteorological feature with the attitude feature into a first uncertainty feature;

[0169] Inputting the first uncertainty feature into the third long short-term memory network to extract a second uncertainty feature;

[0170] Inputting the amplitude sequence into the fourth long short-term memory network to extract amplitude features;

[0171] fusing the second uncertainty feature and the amplitude feature into a first multimodal feature;

[0172] Inputting the first multimodal features into the first fully connected layer and mapping them into second multimodal features;

[0173] Inputting the second multimodal features into the second fully connected layer and mapping them into third multimodal features;

[0174] Inputting the third multimodal feature into the third fully connected layer and mapping it into a fourth multimodal feature;

[0175] Performing an activation operation on the fourth multimodal feature to obtain first Q values ​​for various satellite tracking modules to track satellites for the maritime antenna;

[0176] The second Q value calculation module is also used for:

[0177] Inputting the meteorological sequence into the first long short-term memory network to extract meteorological features representing uncertainty;

[0178] Inputting the posture sequence into the second long short-term memory network to extract posture features representing uncertainty;

[0179] fusing the meteorological feature with the attitude feature into a first uncertainty feature;

[0180] Inputting the first uncertainty feature into the third long short-term memory network to extract a second uncertainty feature;

[0181] Inputting the amplitude sequence into the fourth long short-term memory network to extract amplitude features;

[0182] fusing the second uncertainty feature and the amplitude feature into a first multimodal feature;

[0183] Inputting the first multimodal features into the first fully connected layer and mapping them into second multimodal features;

[0184] Inputting the second multimodal features into the second fully connected layer and mapping them into third multimodal features;

[0185] Inputting the third multimodal feature into the third fully connected layer and mapping it into a fourth multimodal feature;

[0186] An activation operation is performed on the fourth multimodal feature to obtain second Q values ​​of various satellite tracking modules for the maritime antenna to track satellites.

[0187] In one embodiment of the present invention, the incentive is:

[0188]

[0189] Where R is the incentive, Strength t is the signal strength of the maritime antenna at time t, Performance t is the performance parameter of the maritime antenna at time t, Amplitude t is the attitude change amplitude of the maritime antenna at time t, T is the set of time, and α, β, γ and δ are all hyperparameters.

[0190] In one embodiment of the present invention, the target state conversion module 305 includes:

[0191] An antenna rotation amplitude calculation module is used to calculate the difference between the posture of the maritime antenna at the current moment and the posture of the maritime antenna at the next moment output by the first target tracking module as the antenna rotation amplitude;

[0192] A gradient query module, used to query the gradient of the second target tracking module when updating the maritime antenna attitude;

[0193] A candidate state setting module, used for setting the projection of the antenna rotation amplitude on the gradient to a posture suitable for use by the second target tracking module as a candidate state;

[0194] The filtering processing module is used to perform filtering processing on the multiple candidate states to obtain a target state.

[0195] In one embodiment of the present invention, the candidate states are:

[0196]

[0197] Among them, Pose is the candidate state, Turn is the antenna rotation amplitude, T is transpose, G is the gradient, and M is the conversion matrix between the first target tracking module and the second target tracking module.

[0198] In one embodiment of the present invention, the switching condition determination module 306 includes:

[0199] An instance query module, used to query a first antenna instance and a second antenna instance constructed for the maritime antenna; the first antenna instance is a master instance, the second antenna instance is a slave instance, and the first antenna instance has been connected to the first target tracking module;

[0200] An instance access module, used to connect the second antenna instance to the first target tracking module, so as to track the satellite using the second target tracking module;

[0201] A reference state query module, used for querying the posture of the second antenna instance output by the second target tracking module at the next moment as a reference state;

[0202] A state deviation calculation module, used for taking an absolute value of a difference between the target state and the reference state to obtain a state deviation;

[0203] A first condition determination module, configured to determine that the target state satisfies a switching condition if the state deviation is less than or equal to a preset threshold;

[0204] The second condition determination module is used to determine that the target state does not meet the switching condition if the state deviation is greater than a preset threshold.

[0205] In one embodiment of the present invention, the tracking switching module 307 includes:

[0206] A first instance switching module, configured to switch the first antenna instance from the master instance to the slave instance, and release the first antenna instance from access to the first target tracking module, so as to stop using the first target tracking module to track satellites for the maritime antenna;

[0207] The second instance switching module is used to switch the second antenna instance from the slave instance to the master instance to start using the second target tracking module to track satellites for the maritime antenna.

[0208] The satellite-finding device for a maritime antenna provided in an embodiment of the present invention can execute the satellite-finding method for a maritime antenna provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to executing the satellite-finding method for a maritime antenna.

[0209] Embodiment 3

[0210] See also Figure 4 , shows a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0211] like Figure 4 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, ROM 12 and RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0212] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0213] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the satellite search method of the maritime antenna.

[0214] In some embodiments, the satellite search method for the maritime antenna may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the satellite search method for the maritime antenna described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the satellite search method for the maritime antenna in any other appropriate manner (e.g., by means of firmware).

[0215] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0216] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0217] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0218] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0219] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0220] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0221] Embodiment 4

[0222] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the satellite search method for a maritime antenna provided in any embodiment of the present invention is implemented.

[0223] In the process of implementation, the computer program product can be written in one or more programming languages ​​or a combination thereof to perform the computer program code of the present invention, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).

[0224] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0225] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A satellite search method for a maritime antenna, characterized in that: An antenna controller for a maritime antenna on a ship, the method comprising: Loading and configuring a plurality of satellite tracking modules for the maritime antenna; During the operation of the maritime antenna, one of the satellite tracking modules is selected as the maritime antenna tracking satellite as a first target tracking module; Constructing a reinforcement learning model; the intelligent agent in the reinforcement learning model is the antenna controller, the environment is the ship, and the action is selecting one of the satellite tracking modules as the maritime antenna tracking satellite; Starting the reinforcement learning model, with the goal of optimizing the performance index of the maritime antenna, selecting one of the satellite tracking modules as the maritime antenna tracking satellite according to the ship as the second target tracking module; Converting the posture output by the first target tracking module into a posture suitable for use by the second target tracking module as a target state; Determine whether the target state satisfies a switching condition; if so, switch from the first target tracking module to the second target tracking module to track the satellite for the maritime antenna.

2. The method according to claim 1, characterized in that The step of starting the reinforcement learning model with the goal of optimizing the performance index of the maritime antenna and selecting one of the satellite tracking modules as the maritime antenna tracking satellite as the second target tracking module according to the ship includes: At each moment, recording various meteorological information of the environment in which the ship is located to obtain a meteorological sequence; At each moment, recording the attitude data of the ship to obtain an attitude sequence; At each moment, recording the change amplitude of the attitude of the maritime antenna to obtain an amplitude sequence; In the dimension of uncertainty, the meteorological sequence, the attitude sequence and the amplitude sequence are input into the Q learning network in the reinforcement learning model to learn various first Q values ​​of the satellite tracking module for the maritime antenna to track the satellite; Selecting one of the satellite tracking modules according to the first Q value as a second target tracking module to optimize the performance index of the maritime antenna; After switching from the first target tracking module to the second target tracking module, calculating an excitation for the second target tracking module according to the amplitude sequence; In the dimension of uncertainty, the meteorological sequence, the attitude sequence and the amplitude sequence are input into the Q learning network in the reinforcement learning model to learn various second Q values ​​of the satellite tracking module for the maritime antenna to track the satellite; A Q learning network in the reinforcement learning model is updated according to the incentive, the first Q value, and the second Q value.

3. The method according to claim 2, characterized in that The Q learning network in the reinforcement learning model includes a first long short-term memory network, a second long short-term memory network, a third long short-term memory network, a fourth long short-term memory network, a first fully connected layer, a second fully connected layer and a third fully connected layer; In the dimension of uncertainty, the meteorological sequence, the attitude sequence and the amplitude sequence are input into the Q learning network in the reinforcement learning model to learn various first Q values ​​of the satellite tracking module for the maritime antenna to track the satellite, including: Inputting the meteorological sequence into the first long short-term memory network to extract meteorological features representing uncertainty; Inputting the posture sequence into the second long short-term memory network to extract posture features representing uncertainty; fusing the meteorological feature with the attitude feature into a first uncertainty feature; Inputting the first uncertainty feature into the third long short-term memory network to extract a second uncertainty feature; Inputting the amplitude sequence into the fourth long short-term memory network to extract amplitude features; fusing the second uncertainty feature and the amplitude feature into a first multimodal feature; Inputting the first multimodal features into the first fully connected layer and mapping them into second multimodal features; Inputting the second multimodal features into the second fully connected layer and mapping them into third multimodal features; Inputting the third multimodal feature into the third fully connected layer and mapping it into a fourth multimodal feature; Performing an activation operation on the fourth multimodal feature to obtain first Q values ​​for various satellite tracking modules to track satellites for the maritime antenna; In the dimension of uncertainty, the meteorological sequence, the attitude sequence and the amplitude sequence are input into the Q learning network in the reinforcement learning model to learn various second Q values ​​of the satellite tracking module for the maritime antenna to track the satellite, including: Inputting the meteorological sequence into the first long short-term memory network to extract meteorological features representing uncertainty; Inputting the posture sequence into the second long short-term memory network to extract posture features representing uncertainty; fusing the meteorological feature with the attitude feature into a first uncertainty feature; Inputting the first uncertainty feature into the third long short-term memory network to extract a second uncertainty feature; Inputting the amplitude sequence into the fourth long short-term memory network to extract amplitude features; fusing the second uncertainty feature and the amplitude feature into a first multimodal feature; Inputting the first multimodal features into the first fully connected layer and mapping them into second multimodal features; Inputting the second multimodal features into the second fully connected layer and mapping them into third multimodal features; Inputting the third multimodal feature into the third fully connected layer and mapping it into a fourth multimodal feature; An activation operation is performed on the fourth multimodal feature to obtain second Q values ​​of various satellite tracking modules for the maritime antenna to track satellites.

4. The method according to claim 2, characterized in that: The incentives are: Where R is the incentive, Strength t is the signal strength of the maritime antenna at time t, Performance t is the performance parameter of the maritime antenna at time t, Amplitude t is the attitude change amplitude of the maritime antenna at time t, T is the set of time, and α, β, γ and δ are all hyperparameters.

5. The method according to claim 1, characterized in that The step of converting the posture output by the first target tracking module into a posture suitable for use by the second target tracking module as the target state includes: Calculating the difference between the posture of the maritime antenna at the current moment and the posture of the maritime antenna at the next moment output by the first target tracking module as the antenna rotation amplitude; querying the gradient of the second target tracking module when updating the maritime antenna attitude; Setting the projection of the antenna rotation amplitude on the gradient to a posture suitable for use by the second target tracking module as a candidate state; Filtering is performed on the plurality of candidate states to obtain a target state.

6. The method according to claim 5, characterized in that The candidate states are: Among them, Pose is the candidate state, Turn is the antenna rotation amplitude, T is transpose, G is the gradient, and M is the conversion matrix between the first target tracking module and the second target tracking module.

7. The method according to any one of claims 1 to 6, characterized in that The determining whether the target state satisfies the switching condition includes: Querying a first antenna instance and a second antenna instance constructed for the maritime antenna; the first antenna instance is a master instance, the second antenna instance is a slave instance, and the first antenna instance has been connected to the first target tracking module; Connecting the second antenna instance to the first target tracking module to track a satellite using the second target tracking module; Querying the second target tracking module to output the posture of the second antenna instance at the next moment as a reference state; Taking an absolute value of a difference between the target state and the reference state to obtain a state deviation; If the state deviation is less than or equal to a preset threshold, determining that the target state meets the switching condition; If the state deviation is greater than a preset threshold, it is determined that the target state does not meet the switching condition.

8. The method according to claim 7, characterized in that The switching from the first target tracking module to the second target tracking module is for the maritime antenna to track the satellite, comprising: Switching the first antenna instance from the master instance to the slave instance, and disconnecting the first antenna instance from the first target tracking module, so as to stop using the first target tracking module to track satellites for the maritime antenna; The second antenna instance is switched from the slave instance to the master instance to begin tracking satellites for the maritime antenna using the second target tracking module.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the satellite search method for a maritime antenna according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the satellite search method for a maritime antenna according to any one of claims 1 to 8 is implemented.