Vision-based cross-domain unmanned platform cooperative inspection system

CN122506944APending Publication Date: 2026-08-04RONGQI INTELLIGENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RONGQI INTELLIGENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2026-05-07
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

岸基监视受视距、遮挡和海况影响明显,难以对动态目标进行持续确认;单一无人机虽然具备较强视场覆盖能力,但续航时间有限,且在发现目标后通常只能完成图像取证,难以继续实施水面跟踪和拦截;单一无人船虽可接近目标,但搜索范围受限,缺乏高位视觉引导时,面对远距、机动或短时失锁目标时,目标获取效率较低

Benefits of technology

1、空海协同链路完整:通过预警调度、空巡识别、坐标回传、水面协同、协同规划和拦截执行的顺序衔接,形成从预警接入到协同拦截的完整处理链。空中平台输出的目标表征序列、视觉锁稳定度、目标地理位置能够继续进入水面协同和路径规划过程,减少跨模块数据断裂;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122506944A_ABST
    Figure CN122506944A_ABST
Patent Text Reader

Abstract

This invention discloses a vision-based cross-domain unmanned platform collaborative inspection system, belonging to the field of cross-domain unmanned collaborative inspection technology. The system includes: an early warning scheduling module, an aerial patrol identification module, a coordinate feedback module, a surface collaboration module, a collaborative planning module, and an interception execution module. The early warning scheduling module generates inspection task sets and platform scheduling instructions; the aerial patrol identification module generates target representation sequences and visual lock stability; the coordinate feedback module generates target geographical locations and geographical mapping credibility; the surface collaboration module generates target coupling consistency and target prediction status; the collaborative planning module generates cross-domain collaboration consistency index, collaborative paths, and interception instructions; and the interception execution module generates interception status. By constructing a continuous processing chain of early warning scheduling, visual identification, coordinate feedback, surface collaboration, sector-loop planning, and collaborative interception, stable inspection and collaborative handling of dynamic targets by a cross-domain unmanned platform are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cross-domain unmanned collaborative inspection technology, specifically a vision-based cross-domain unmanned platform collaborative inspection system. Background Technology

[0002] Current maritime patrols primarily employ shore-based surveillance, single UAV patrols, or single unmanned surface vessel (USV) patrols. Shore-based surveillance is significantly affected by line-of-sight distance, obstructions, and sea conditions, making it difficult to continuously confirm dynamic targets. While single UAVs possess strong field-of-sight coverage capabilities, their endurance is limited, and after detecting a target, they can typically only complete image evidence collection, making it difficult to continue surface tracking and interception. Although single USVs can approach targets, their search range is limited, and without high-level visual guidance, their target acquisition efficiency is low when facing distant, maneuvering, or temporarily lost-locked targets.

[0003] Existing cross-platform collaborative solutions generally suffer from the following problems: First, there is a lack of a unified scheduling mechanism among early warning information, boundary constraints, and communication status, resulting in insufficient coordination between platform task allocation and timing. Second, it is difficult to reliably convert images of targets acquired by aerial platforms into geographical locations that can be directly accessed by surface platforms, leading to mapping errors between visual recognition results and spatial coordinates. Third, after receiving the target location, surface platforms lack an effective association mechanism with local candidate targets, making them prone to mistracking or target switching. Fourth, the collaborative planning process does not adequately utilize the reliability of visual recognition, the reliability of geographical mapping, and the consistency of cross-platform associations, making it difficult to simultaneously address boundary constraints, restricted area constraints, link status, and dynamic interception requirements. Fifth, during the interception execution phase, there is a lack of unified collaborative control criteria between UAVs and unmanned vessels, resulting in insufficient coordination between path execution, target acquisition, and status determination.

[0004] Therefore, a vision-based cross-domain unmanned platform collaborative inspection system is needed to achieve continuous connection between early warning scheduling, aerial patrol identification, coordinate transmission, surface coordination, collaborative planning, and interception execution. Summary of the Invention

[0005] Based on the shortcomings of the prior art described above, the purpose of this invention is to provide a vision-based cross-domain unmanned platform collaborative inspection system to solve the aforementioned technical problems.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a vision-based cross-domain unmanned platform collaborative inspection system, comprising: Early warning and dispatch module: Receives the boundary of the task sea area, the boundary of the restricted area, the status of the air and sea communication links and the early warning location, and performs task orchestration to generate inspection task sets and platform dispatch instructions; The aerial patrol identification module controls the UAV to collect image frames, pod attitude, body attitude and platform position and altitude based on the patrol task set. It performs structural topological coding and cross-frame displacement constraints on the image frames to generate target representation sequences and visual lock stability. Coordinate return module: Performs coordinate mapping on the target representation sequence, pod attitude, body attitude and platform position and altitude according to the platform scheduling instructions to generate the target geographical location and geographical mapping reliability; Surface Coordination Module: Based on the target's geographical location, the module controls the unmanned vessel to collect the status of candidate targets, correlates and extrapolates the target's geographical location and the candidate target status, and generates the target coupling consistency degree and the target predicted status. Collaborative planning module: Based on visual lock stability, geographic mapping credibility and target coupling consistency, a cross-domain collaborative consistency index is generated. The module performs fan-loop interception planning based on the target prediction status, mission sea area boundary, restricted area boundary and air-sea communication link status to generate collaborative paths and interception commands. Interception Execution Module: Controls UAVs and unmanned vessels to perform collaborative interception and generate interception status based on the collaborative path and interception command.

[0007] The present invention is further configured such that the execution of task orchestration to generate inspection task sets and platform scheduling instructions includes: Receive information on the boundaries of the mission sea area, the boundaries of restricted areas, the status of air and sea communication links, and early warning locations; Determine the schedulable area based on the boundaries of the mission sea area and the restricted area, and extract boundary constraint information; Link quality analysis is performed based on the status of air and sea communication links to generate link availability information; Constraints are adjusted on the warning location, schedulable area, boundary constraint information, and link availability information to generate the warning seed area and the core location of the task. Based on the early warning seed area, the core location of the mission, the current location of the UAV and the current location of the unmanned vessel, a set of inspection missions is generated through scheduling and matching. Based on the inspection task set, timing coordination and instruction encapsulation are performed to generate platform scheduling instructions, which include time-stamped windows and feedback cycles.

[0008] The present invention is further configured such that performing structural topological coding on the image frame to generate the target representation sequence includes: Based on the task-related search domain in the image frame defined by the inspection task set, candidate target regions are extracted within the task-related search domain. Based on the boundary lines of the candidate target region, perform contour segmentation and turn positioning to generate outer contour structure information; Based on the internal connectivity paths of the candidate target region, skeleton expansion and branch parsing are performed to generate local connectivity structure information; Generate regional distribution structure information based on the centroid location, main branch extension direction, and regional envelope relationship of the candidate target region; Topological association encoding is performed on external contour structure information, local connectivity structure information and regional distribution structure information to generate single-frame structural representations of candidate targets; The structural representations of a single frame are arranged according to the image acquisition time sequence to generate a target representation sequence.

[0009] The present invention is further configured such that the step of generating visual lock stability by performing cross-frame displacement constraints includes: Attitude compensation results are generated by performing attitude compensation based on the pod attitude and body attitude corresponding to adjacent image frames; The predicted position of the candidate target in the current frame is determined based on the target representation sequence and pose compensation results in adjacent image frames; Displacement continuity constraints are applied to the candidate target positions and predicted positions in the current frame to generate displacement consistency information; Scale consistency information is generated by constraining scale variation based on the distribution of outer contour span and main branch length in adjacent image frames. Orientation consistency information is generated by constraining the orientation change based on the main branch extension direction and the main direction of the outer contour in adjacent image frames. The displacement consistency information, scale consistency information, and orientation consistency information are correlated in a time sequence to generate continuous target correlation results; Visual lock stability is generated by extracting temporal preservation information based on the results of continuous target association.

[0010] The present invention is further configured such that the step of generating the target geographic location by performing coordinate mapping on the target representation sequence, pod attitude, aircraft attitude, and platform position altitude includes: Based on the time-stamped window and transmission cycle in the platform scheduling instructions, time-stamped registration is performed on the target representation sequence, pod attitude, airframe attitude, platform position and altitude to generate a mapped time series group; Based on the target representation sequence in the mapping time series, the pixel anchor coordinates are generated by time series stability screening of the outer contour inflection points, local connection nodes and regional distribution centers. The camera gaze vector in the camera coordinate system is generated by calculating the camera gaze vector based on the mapped time series group, pixel anchor point coordinates and preset camera intrinsic parameters. Based on the pod attitude, body attitude, and camera coordinate system line-of-sight vector in the mapping timing group, a progressive coordinate transformation is performed to generate the navigation coordinate system line-of-sight vector; The target position vector is generated by solving the landing point based on the platform position, altitude, and navigation coordinate system line-of-sight vector in the mapped time series group; The target location vector is transformed into geographic coordinates to generate the target geographic location.

[0011] The present invention further specifies that the credibility of the generated geographic mapping includes: The corrected target position vector is generated by performing time-series correction on the target position vectors at consecutive time points. The reprojection position is generated by back projection based on the corrected target position vector, pixel anchor point coordinates, pod attitude, body attitude, platform position height, and preset camera intrinsic parameters. The pixel anchor point coordinates and reprojection positions are compared for consistency to generate reprojection deviation information. Continuous deviation information is generated by performing temporal continuity analysis based on the corrected target position vector and the target position vector at consecutive time points; Attitude perturbation analysis is performed on the pod attitude, airframe attitude, and corrected target position vector to generate attitude-sensitive deviation information; Geographic mapping credibility is generated by consistent fusion of reprojection bias information, continuous bias information and attitude-sensitive bias information.

[0012] The present invention is further configured such that the step of associating and extrapolating the target geographic location and candidate target state to generate target coupling consistency and target predicted state includes: Generate a reference location by performing local coordinate mapping on the target geographic location; A local search domain is generated based on the reference position, and the unmanned surface vessel is controlled to collect the position, heading and speed of candidate targets within the local search domain to generate a candidate target state set; The target's geographical location and its previous state are used to perform motion completion to generate a reference target state; Spatial gating and motion gating are applied to the reference target state and the candidate target state set to generate an associative candidate set; The associated candidate set and the associated target state of the previous time step are temporally associated to generate the associated target state and the target coupling consistency degree; Target prediction states are generated by extrapolating the trajectory based on the associated target states.

[0013] The present invention is further configured such that the generation of the cross-domain collaborative consistency index based on visual lock stability, geographic mapping credibility, and target coupling consistency includes: Consistent center information is generated by centrally aggregating visual lock stability, geographic mapping credibility, and target coupling consistency. Discrete information for consistency is generated by performing discrete analysis on visual lock stability, geographic mapping credibility, and target coupling consistency with consistency center information. A cross-domain collaborative consistency index is generated by consistency fusion based on visual lock stability, geographic mapping credibility, target coupling consistency, and consistency discrete information.

[0014] The present invention is further configured such that the process of generating a coordinated path and interception instructions by performing sector-loop interception planning based on the target prediction status, mission sea area boundary, restricted area boundary, and air-sea communication link status includes: Based on the target prediction status, the mission sea area boundary, the restricted area boundary, and the air and sea communication link status, constraint analysis is performed to generate the planning feasible domain and constraint cost field. The target prediction state and the cross-domain coordination consistency index are collaboratively mapped to generate a fan-ring interception domain. Interception anchor points are generated by evaluating the acceptance point based on the sector-ring interception domain and the constraint cost field. A cooperative path is generated by solving the path between the current position of the unmanned vessel and the interception anchor point. Interception commands are generated by encapsulating patterns based on cross-domain collaboration consistency index, sector-ring interception domain, interception anchor point, target prediction state, and collaborative path.

[0015] The present invention is further configured such that controlling the UAV and unmanned surface vessel to perform cooperative interception and generate an interception state based on the cooperative path and interception command includes: The collaborative path, interception command, current state of the UAV and current state of the unmanned vessel are time-calibrated to generate an execution state group; The target prediction state, interception anchor point, and sector interception domain in the interception command are analyzed. Based on the target prediction state and interception anchor point, the UAV execution reference is generated. The current state of the UAV and the UAV execution reference are matched for deviation to generate UAV control commands. Based on the cooperative path and interception anchor point, an unmanned vessel execution reference is generated. The current state of the unmanned vessel and the unmanned vessel execution reference are matched for deviation to generate unmanned vessel control commands. The target prediction status, sector interception domain, UAV control commands and UAV control commands are collaboratively corrected to generate a collaborative interception control quantity, which is then sent to the UAV actuator and the UAV actuator. Based on the execution reference of the UAV, the execution reference of the UAV, the cooperative interception control quantity and the sector interception domain, the interception state is generated and sent to the upper-level task state machine.

[0016] This invention provides a vision-based cross-domain unmanned platform collaborative inspection system. The system comprises: an early warning and scheduling module that receives the boundary of the task sea area, the boundary of the restricted area, the status of the air-sea communication link, and the early warning location; and a task orchestration module that generates an inspection task set and platform scheduling instructions. An aerial inspection and identification module controls the UAV to collect image frames, pod attitude, aircraft attitude, and platform position and altitude based on the inspection task set; performs structural topology encoding and cross-frame displacement constraints on the image frames to generate a target representation sequence and visual lock stability. A coordinate return module performs coordinate mapping on the target representation sequence, pod attitude, aircraft attitude, and platform position and altitude according to the platform scheduling instructions to generate the target's geographical location and... Geographic mapping reliability; Surface coordination module: Controls unmanned surface vessels to collect candidate target states based on target geographic location, correlates and extrapolates target geographic location and candidate target states, and generates target coupling consistency and target prediction state; Cooperative planning module: Generates cross-domain cooperative consistency index based on visual lock stability, geographic mapping reliability, and target coupling consistency, and performs fan-loop interception planning based on target prediction state, mission sea area boundary, restricted area boundary, and air-sea communication link state to generate cooperative paths and interception commands; Interception execution module: Controls UAVs and unmanned surface vessels to perform cooperative interception based on cooperative paths and interception commands to generate interception state. The beneficial effects include: 1. Complete Air-Sea Collaborative Link: Through the sequential connection of early warning scheduling, air patrol identification, coordinate feedback, surface coordination, collaborative planning, and interception execution, a complete processing chain is formed from early warning access to collaborative interception. The target representation sequence, visual lock stability, and target geographical location output by the air platform can continue to enter the surface coordination and path planning process, reducing cross-module data fragmentation; 2. Clear target mapping relationship: Through structural topological coding, cross-frame displacement constraints, and coordinate mapping processing, a continuous correspondence is established between the target representation in the image domain and the target's geographical location in the spatial domain, and the mapping result is constrained by the geographical mapping credibility. This enables the visual recognition results to further support candidate target association, target prediction, and collaborative planning; 3. Sufficient constraints on interception planning: By unifying visual lock stability, geographic mapping credibility, and target coupling consistency into a cross-domain collaborative consistency index, and then combining the task sea area boundary, restricted area boundary, link status, and target prediction status for sector-ring interception planning, the collaborative path generation and interception command generation have a consistent data foundation, and the interception execution and status determination are more completely connected.

[0017] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 The flowchart illustrates a vision-based cross-domain unmanned platform collaborative inspection system as an exemplary embodiment of the present invention. Detailed Implementation

[0019] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0020] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0021] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0022] Example 1: Vision-based cross-domain unmanned platform collaborative inspection system, such as Figure 1 As shown, it includes: Early warning and dispatch module: Receives the boundary of the task sea area, the boundary of the restricted area, the status of the air and sea communication links and the early warning location, and performs task orchestration to generate inspection task sets and platform dispatch instructions; The aerial patrol identification module controls the UAV to collect image frames, pod attitude, body attitude and platform position and altitude based on the patrol task set. It performs structural topological coding and cross-frame displacement constraints on the image frames to generate target representation sequences and visual lock stability. Coordinate return module: Performs coordinate mapping on the target representation sequence, pod attitude, body attitude and platform position and altitude according to the platform scheduling instructions to generate the target geographical location and geographical mapping reliability; Surface Coordination Module: Based on the target's geographical location, the module controls the unmanned vessel to collect the status of candidate targets, correlates and extrapolates the target's geographical location and the candidate target status, and generates the target coupling consistency degree and the target predicted status. Collaborative planning module: Based on visual lock stability, geographic mapping credibility and target coupling consistency, a cross-domain collaborative consistency index is generated. The module performs fan-loop interception planning based on the target prediction status, mission sea area boundary, restricted area boundary and air-sea communication link status to generate collaborative paths and interception commands. Interception Execution Module: Controls UAVs and unmanned vessels to perform collaborative interception and generate interception status based on the collaborative path and interception command.

[0023] The present invention is further configured such that the execution of task orchestration to generate inspection task sets and platform scheduling instructions includes: The system receives the mission sea area boundary, restricted area boundary, air-sea communication link status, and early warning location. Specifically, it reads the mission sea area boundary and restricted area boundary from the mission configuration source. The mission sea area boundary is represented by a closed boundary point sequence, and the restricted area boundary is represented by a set of closed boundary point sequences. All boundary points have a unified coordinate reference. After receiving the data, a boundary validity check is performed. The check includes whether the boundary point sequence is continuous, whether the beginning and end are closed, whether the boundary lines intersect, and whether the restricted area extends beyond the mission sea area. If there are unclosed boundaries, self-intersecting boundaries, or inconsistent coordinate references, the data is directly marked as invalid input and the system enters the re-reading process. The air-sea communication link status is obtained from the UAV communication terminal, the UAV communication terminal, and the relay link monitoring. The recent sampling results from the testing end include at least four types of data: link strength, round-trip delay, packet loss ratio, and number of consecutive valid samples. Link strength reflects wireless connectivity, round-trip delay reflects control and transmission time consumption, packet loss ratio reflects link integrity, and the number of consecutive valid samples reflects link stability. The warning location is input using a unified coordinate format, and the warning timestamp is recorded. If the warning source provides an additional error range, that error range is received synchronously; otherwise, the system's preset error radius is used as the initial expansion range. After receiving the data, the boundary data, link data, and warning location are processed using a unified time scale and coordinate reference to generate a standardized input set that can be entered into the task orchestration process. The schedulable area is determined based on the mission sea area boundary and the restricted area boundary, and boundary constraint information is extracted. Specifically, the mission sea area boundary is considered as the outer allowed area, and the restricted area boundary is considered as the inner excluded area. Region clipping is performed on both to obtain the schedulable area. After region clipping, this area is not directly used as the scheduling result. Instead, regular sampling units are established within the schedulable area. The sampling units can use a regular grid or an equidistant scatter plot, and their spacing is determined by the minimum scheduling resolution preset by the system. This resolution is derived from the smaller of the minimum effective observation coverage scale of the UAV and the minimum discernible track adjustment scale of the UAV. Subsequently, the shortest distance to the mission sea area boundary and the shortest distance to the restricted area boundary are calculated for each sampling unit. To ensure sufficient safety margin for subsequent scheduling matching, the system presets two types of buffer distances: One type is the outer boundary buffer distance, used to limit the platform's proximity to the outer edge of the mission area; the other is the restricted area buffer distance, used to limit the platform's proximity to the restricted area boundary. These two types of buffer distances are not arbitrarily given, but are determined by the upper bound of the UAV's positioning error, the distance occupied by the UAV's turning, the displacement corresponding to command transmission delay, and the mission's allowable safety redundancy. During processing, each sampling unit is divided into a core safety zone, a transitional restricted zone, and a high-risk zone: when it is simultaneously far from both the mission area boundary and the restricted area boundary, it is classified as belonging to the core safety zone; when it is still within the schedulable area but falls into any buffer zone, it is classified as belonging to the transitional restricted zone; when it overlaps with the restricted area or exceeds the mission area boundary, it is directly judged as an unschedulable zone. The final output boundary constraint information is a boundary constraint table composed of area affiliation, outer boundary proximity level, and restricted area proximity level. Link quality analysis is performed based on the air-sea communication link status to generate link availability information. Specifically, the link status is analyzed in stages. First, the most recent continuous sampling sequence is extracted within a preset time window, and abnormal samples are removed for link strength, round-trip delay, packet loss ratio, and number of consecutive valid samples. The abnormal sample removal adopts the median deviation screening rule, that is, the median value of each type of link indicator within the time window is first calculated, and then sampling points that deviate from the median value by more than the preset tolerance range are removed to avoid short-term jitter and occasional interference directly affecting the availability judgment. Subsequently, sub-levels are generated for the four types of link indicators after removing abnormal samples. Link strength is divided into four levels: below the minimum access threshold, in the accessible range, in the stable access range, and in the priority access range. Round-trip delay is divided into four levels: timeout, high, normal, and preferred. Packet loss ratio is divided into four levels: high loss, medium loss, low loss, and near-no loss. The number of consecutive valid samples is... The data is categorized into four levels: occasional connection, intermittent connection, stable connection, and continuous connection. After the four sub-levels are generated, a joint evaluation method of first hard screening and then weighting is adopted: any link whose strength is lower than the minimum access threshold, round-trip delay exceeds the system's acceptable upper limit, or packet loss ratio exceeds the control loop's allowable upper limit is directly judged as unavailable. If the hard screening conditions are not triggered, the four sub-levels are then comprehensively scored according to preset weights, with link strength and round-trip delay having higher weights than the number of consecutive effective samples, in order to prioritize the real-time performance and stability of the control chain closed loop. The comprehensive score result is then mapped to the link availability level, which is divided into high availability, medium availability, low availability, and unavailable. If the task area is large, the link availability level of adjacent communication monitoring points is inherited for different spatial sampling units to form a link availability information distribution map. If the area is small, the current comprehensive link level is directly used as the link availability information for the entire task area. The system generates a seed area for early warning and a core location for the task by constraining and correcting the warning location, schedulable area, boundary constraint information, and link availability information. Specifically, an initial warning coverage area is generated centered on the warning location. If the warning input has an error range, this error range is used as the outer radius; if the warning input does not have an error range, the system's preset basic outer radius is used. After the initial warning coverage area is generated, it is first trimmed with the schedulable area, deleting portions that fall into the unschedulable area to form the first candidate area. Then, the generated boundary constraint table is called to filter the sampling units in the first candidate area by boundary level: unschedulable areas are directly deleted, high-risk areas do not enter the warning seed area, transitional restricted areas are retained but with reduced priority, and core safety areas are directly retained. After that, the generated link availability information is called to filter the retained units by link level: unavailable units are directly deleted, low-availability units are used as backup areas, and medium-availability units are used as backup areas. The unit serves as the schedulable area, and the highly available unit serves as the priority area. After completing the dual screening of boundaries and links, the system performs continuous region clustering on the remaining units, identifying spatially continuous unit clusters with both high boundary security levels and high link availability levels. The unit cluster closest to the warning location and whose continuous area meets the minimum task area requirement is selected as the warning seed area. The warning seed area refers to the area that subsequent inspection tasks will prioritize and that has an executable scheduling basis, rather than simply a circular area generated around the warning location. After the warning seed area is determined, priority is sorted within this area according to proximity to the warning location, lower boundary risk, and higher link level. The high-priority units are then weighted and their centers are calculated to generate the task core location. The task core location is not the geometric center, but a central location that integrates security and accessibility, making it more suitable as the organizational base for inspection tasks and subsequent platform support. Based on the early warning seed area, the core mission location, the current location of the UAV, and the current location of the UAV, a patrol task set is generated through scheduling and matching. Specifically, the current locations of the UAV and UAV are read, and a set of candidate patrol points for UAVs and a set of candidate standby points for UAVs are generated in combination with the early warning seed area and the core mission location. The UAV candidate patrol point set is deployed in the observation entrance zone corresponding to the area near the core mission location and above the early warning seed area. Its generation rule is to cover the main direction of the early warning seed area and avoid the high-risk boundary zone. The UAV candidate standby point set is deployed in the feasible access zone between the outer edge of the early warning seed area and the core mission location. Its generation rule is to facilitate subsequent interception while avoiding entering the restricted area buffer zone. For each UAV candidate patrol point, the system calculates four matching quantities: the distance from the current location of the UAV to the point, the boundary risk level corresponding to the point, the link availability level corresponding to the point, and the point's connection to the early warning. The observation coverage level of the seed area refers to the degree to which the seed area can be continuously covered by the UAV's electro-optical payload from the patrol point. First, the visible coverage ratio from the candidate patrol point to the seed area is calculated, and then the observation coverage level is divided according to the coverage ratio interval. For each UAV candidate standby point, the system calculates four matching quantities: the reach distance from the current position of the UAV to the point, the boundary risk level of the point, the link availability level of the point, and the accessibility level of the point to the mission core position. The accessibility level is calculated by first determining the number of path turns, the maximum turning angle, and the total heading adjustment from the candidate standby point to the mission core position, and then dividing the accessibility level according to the interval. Subsequently, the candidate UAV patrol points and the candidate UAV standby points are comprehensively ranked. During the ranking, points falling into high-risk boundary zones and low link level zones are eliminated first, and then points with shorter reach distances and higher coverage or accessibility levels are selected from the remaining points. After sorting, the drone patrol anchor points and unmanned surface vessel standby anchor points are determined, and together with the early warning seed area, the core mission location, the platform identity, and the mission mode, they are packaged into a patrol mission set; the patrol mission set shall include at least the mission area, the core mission location, the drone patrol anchor point, the unmanned surface vessel standby anchor point, the drone patrol mode, and the unmanned surface vessel standby mode. Based on the inspection task set, timing coordination and instruction encapsulation are performed to generate platform scheduling instructions. These instructions include a time-stamped window and a return cycle. Specifically, timing coordination is performed on the UAV and unmanned surface vessel (USV) based on the inspection task set. In the specific processing, the platform's preset motion parameters are first used to calculate the estimated arrival time of the UAV from its current position to the inspection anchor point, and then the estimated arrival time of the USV from its current position to the standby anchor point is calculated. The motion parameters are sourced from the UAV's cruise speed reference, the USV's average speed reference, and a path reduction factor. The path reduction factor is used to compensate for the additional time lost due to turning, acceleration / deceleration, and boundary detours during the platform's actual movement. If the USV's estimated arrival time is earlier than the UAV's, a waiting time is set for the USV to maintain link listening at the standby point. If the UAV's estimated arrival time is later, the USV's standby start time is postponed, and the UAV's inspection instruction is given a priority start flag. After timing coordination is completed, instruction encapsulation is performed. During encapsulation, not only is the instruction written... The anchor point location and task mode must also be written into the time-stamped window and the return cycle. The generation rule for the time-stamped window is: based on the expected start time, a preset preparation time is retained forward, and a minimum task processing cycle is covered backward, forming an effective time range that allows the same round of processing to be called. Subsequent air patrol identification, coordinate return, and surface coordination can only call data falling within this time range. The generation rule for the return cycle is: first read the basic return rhythm preset by the task, and then adjust it in combination with the link availability level. When the link level is high, a shorter cycle is used to maintain the state update density. When the link level is at a medium level, a medium cycle is used to balance the update frequency and link occupancy. When the link level is close to the lower limit but still available, a longer cycle is used to ensure that critical data can be transmitted stably. Finally, the platform scheduling instruction is output in the form of structured fields, including at least the platform identifier, task area, core task location, patrol anchor point or standby anchor point, start time, task mode, time-stamped window, and return cycle.

[0024] The present invention is further configured such that performing structural topological coding on the image frame to generate the target representation sequence includes: Based on the task-related search domain defined by the inspection task set in the image frame, candidate target regions are extracted within the task-related search domain. Specifically, the task area, inspection anchor point, and platform task mode in the inspection task set are read, and then combined with the current platform position, altitude, pod current pointing angle, and image frame timestamp of the UAV, the task-related search domain that should be prioritized in the current image frame is determined. During processing, the spatial area in the inspection task set is first projected onto the current image plane, and then the preset search width and preset search height are expanded to both sides with the main viewing direction corresponding to the inspection anchor point as the center, forming a candidate search window in the image domain. The preset search width refers to the horizontal expansion of the image with the projection position of the task core position in the current image frame as the center. The search width is calculated by covering the currently visible projection range of the warning seed area and adding compensation widths corresponding to the maximum possible lateral displacement of the target within the current processing cycle to the left and right sides. When this width exceeds the effective imaging lateral range of the current image frame, the effective imaging lateral range of the current image frame is used as the preset search width. The preset search height is calculated by centering on the projection position of the task core position in the current image frame, covering the currently visible projection range of the warning seed area vertically, and adding compensation heights corresponding to the maximum possible longitudinal displacement of the target within the current processing cycle to the upper and lower sides. When this height exceeds the effective imaging longitudinal range of the current image frame, the effective imaging longitudinal range of the current image frame is used as the preset search height. The search range is set as the preset search height; the compensation width corresponding to the maximum possible lateral displacement of the target within the current processing cycle and the compensation height corresponding to the maximum possible longitudinal displacement of the target within the current processing cycle are jointly determined by the target's maximum relative motion velocity, the current processing cycle length, and the current imaging scale; the effective imaging range is derived from the overlapping area of ​​the current pod's instantaneous field of view and the camera's effective image plane range; if the search window exceeds the image boundary, it is cropped according to the image boundary; after obtaining the task-related search domain, brightness equalization and local contrast tuning are first performed on the region to ensure that the boundaries and connectivity remain distinguishable under different imaging conditions; then, region separation is performed to remove obviously discontinuous segments or segments that are too similar to the background texture, retaining only those that are full-length. The regions defined by the lower limits of area, edge closure, and local contrast are used as candidate target regions. The lower limit of area is derived from the minimum resolvable size of the target at the current height and resolution. The lower limit of edge closure is defined as the maximum allowable interval threshold between the first and last ends of the boundary point column of the candidate target region. This threshold is set according to a certain proportion of the minimum resolvable side length of the target at the current image resolution and is corrected in combination with the average distance between adjacent boundary points. When the actual interval between the first and last ends is not greater than this threshold, the boundary closure is deemed to meet the requirements. The lower limit of local contrast is derived from the noise level of the image sensor. First, the background grayscale fluctuation amplitude in the neighborhood of the target region is extracted, and then a number of times this fluctuation amplitude is used as the lower limit of local contrast.By first limiting the task-related search domain and then extracting candidate target regions, background regions in the image that are irrelevant to the task will not enter the subsequent topology modeling process, thus enabling the structure encoding to be built on effective candidates under task constraints. Based on the boundary lines of the candidate target regions, contour segmentation and turning point localization are performed to generate outer contour structure information. Specifically, for each candidate target region, the closed boundary line is first extracted, and then a boundary point series is generated along the boundary line at a fixed sampling interval. Subsequently, contour segmentation is performed. The processing of contour segmentation is not based on simple length division, but on the continuity of directional changes and the curvature change trend of the boundary point series. The continuity of directional changes is used to determine whether adjacent boundary points belong to the same smooth contour segment, and the curvature change trend is used to determine whether there is a significant turn in the boundary. During processing, the local directional changes are calculated point by point along the boundary point series, and then boundary points that are continuously within the same directional change range are merged into a contour segment. When the directional change exceeds the preset turning point... When the threshold and the duration reach the minimum turning segment length, it is determined to be a valid turning position; the resulting outer contour structure information includes at least: the number of contour segments, the length level of each contour segment, the connection order between contour segments, the distribution density of turning positions, and the relative interval of turning positions; turning position localization comes from the geometric abrupt changes in the target shape on the boundary, rather than a set of manually defined corner points; contour segments represent continuous segments with relatively consistent shapes on the boundary; through contour segmentation and turning position localization, the outer boundary of the candidate target is no longer just a closed curve, but is transformed into a set of comparable, sortable, and traceable structural information, providing an external skeleton reference for the joint encoding of subsequent internal connection relationships and regional distribution relationships; Based on the internal connectivity paths of the candidate target region, skeleton expansion and branch parsing are performed to generate local connectivity structure information. Specifically, the candidate target region is first refined by compressing the original solid region into a central axis structure that maintains the connectivity relationship. The central axis structure is the skeleton. The skeleton is derived from the result of shrinking the target region layer by layer while preserving the connectivity topology, thus reflecting the basic connectivity pattern inside the target. After obtaining the skeleton, the endpoints, branching points and main paths on the skeleton are further identified. Endpoints represent the termination positions of local connection paths, bifurcation points represent the locations where branches appear in the connection paths, and the main path represents the skeleton path with the highest degree of connectivity and the longest extension distance. Subsequently, branch parsing is performed to record the connection relationship between each branch and the main path, the branch length level, the degree of deviation of the branch direction, and the branch distribution order, forming local connection structure information. The local connection structure is a structured description of which paths are connected, where they separate, and which path is dominant within the candidate target. The branch length level is derived from the ratio of the branch length to the main path length, the degree of deviation of the branch direction is derived from the angle between the branch and the main path, and the branch distribution order is derived from the arrangement order of the bifurcation points along the main path. Through skeleton expansion and branch parsing, the internal connectivity of the candidate target is clearly expressed, and even if different candidate regions are similar in outer contour size, they can be distinguished by their internal connectivity. The region distribution structure information is generated based on the centroid position, main branch extension direction, and region envelope relationship of the candidate target region. Specifically, the centroid position of the candidate target region is first calculated. The centroid position is the balance center of all effective pixels within the region on the image plane, used to represent the overall spatial offset of the target region. Then, the obtained main path is called, and the start and end directions and main extension directions of the main path are extracted as the main branch extension directions, used to represent the main unfolding direction of the candidate target. Next, the region envelope relationship is generated. Specifically, the candidate target region is surrounded by a minimum closed envelope surface, and the relative positional relationship of the outer contour structure, internal skeleton structure, and region centroid within this envelope surface is calculated. For example, whether the centroid is biased to one side, whether the main branch is close to the envelope boundary, and whether the turning points are concentrated in the same envelope direction. The region envelope relationship originates from the spatial occupancy status of the candidate target within the overall coverage area and describes how the internal structure corresponds to the external coverage boundary. The resulting regional distribution structure information includes at least: the centroid offset level, the main branch extension direction level, the proximity relationship between the main branch and the envelope boundary, and the distribution density of the outer contour turning points within the envelope plane; by adding the regional distribution structure information, the external boundary information and internal connection information formed in the first two steps are further placed into a unified spatial framework. A single-frame structural representation of candidate targets is generated by topological association encoding of external contour structure information, local connection structure information, and regional distribution structure information. Specifically, the external contour structure information, local connection structure information, and regional distribution structure information are first mapped to the same structural description framework. This framework is organized by structural nodes and structural connection relationships: the turning points and contour segment endpoints in the external contour are used as outer structural nodes, the endpoints and bifurcation points in the internal skeleton are used as inner structural nodes, and the centroid position and main branch direction are used as regional distribution nodes. Subsequently, connections are established according to the actual adjacency relationships between nodes: contour segments are connected through boundary connections, branches and main paths are connected through skeleton connections, and the centroid and main branches, and the centroid and envelope boundary are connected through distribution association relationships. After completing the organization of nodes and connection relationships, the node type, node order, connection direction, and relative position level are recorded in a unified order to form a single-frame structural representation. The single-frame structural representations are arranged according to the image acquisition time sequence to generate a target representation sequence. Specifically, the single-frame structural representations obtained in consecutive image frames are time-stamped and aligned. Time alignment refers to confirming whether the candidate target in the current frame belongs to the same target as the previous frame based on the structural correspondence between the structural representations of the previous frame and the current frame. During confirmation, the changes in the number of contour segments in the outer contour structure are compared first to see if they are within the allowable range, then the main path and branch relationships in the local connectivity structure are compared to see if they remain continuous, and finally the centroid shift and main branch direction in the regional distribution structure are compared to see if they are within the allowable range of change. Only when all three types of structural relationships meet the time continuity condition is the single-frame structural representation of the current frame continued to the previous frame. If any type of structural relationship is obviously broken, a new sequence is started or the representation is judged as discontinuous. The allowable intervals between adjacent image frames are as follows: The range of variation consists of position change threshold, scale change threshold, and orientation change threshold. The position change threshold is derived from the time interval between adjacent image frames, the upper limit of the platform's maneuvering speed, the upper limit of the target's relative displacement speed, and the current visual resolution. The scale change threshold is derived from the upper limit of the change in the relative distance between the target and the platform within the time interval between adjacent image frames. The orientation change threshold is derived from the upper limit of the platform's turning speed, the upper limit of the target's relative turning speed, and the upper limit of the orientation extraction error. After frame-by-frame concatenation, a target representation sequence is formed. This sequence not only records the single-frame structural representation at each moment but also preserves the structural continuity relationship. Therefore, subsequent cross-frame displacement constraints and visual lock stability generation can directly call this sequence for continuity judgment. Through this arrangement, the single-frame structural information is elevated into a target representation chain with temporal continuity, thus providing a clear data foundation for the subsequent stable locking of the aerial patrol recognition module.

[0025] The present invention is further configured such that the step of generating visual lock stability by performing cross-frame displacement constraints includes: Attitude compensation results are generated based on the pod attitude and aircraft attitude corresponding to adjacent image frames. Specifically, the pod attitude data and aircraft attitude data corresponding to the current image frame and the previous image frame are read first. The pod attitude data comes from the pod servo encoder or pod attitude calculation module and includes at least pitch, yaw, and roll angle information. The aircraft attitude data comes from the flight control inertial navigation module and includes at least aircraft attitude angle and heading change information. After reading, the two frames of attitude data are time-stamped. If the time difference exceeds the preset synchronization tolerance, interpolation is performed on the earlier frame attitude to make the two frames of attitude at the same comparison reference. Then, the two frames are... The changes in pod attitude and body attitude are converted into field of view offsets in the image plane. The field of view offsets are generated based on camera intrinsic parameters, pod installation geometry, and a fixed mapping relationship between the body coordinate system and the camera coordinate system. During processing, translational offsets, rotational offsets, and scale fine-tuning caused by changes in platform attitude are compensated item by item to obtain attitude compensation results. These results include at least the compensation offset direction, compensation offset magnitude, and compensated reference field of view position relative to the previous frame. Through this step, subsequent position predictions are based on an image reference whose platform attitude influence has been eliminated, and position continuity constraints can correspond to the relative changes of the target itself. The predicted position of the candidate target in the current frame is determined based on the target representation sequence and pose compensation results in adjacent image frames. Specifically, the structural representation of the confirmed candidate target in the previous frame is first read from the target representation sequence, which includes at least the outer contour center, local connectivity skeleton center, and region distribution center of the candidate target. Then, the generated pose compensation result is applied to the center position of the confirmed candidate target in the previous frame to obtain the theoretical migration position in the current frame. To improve the continuity of the predicted position, a set of displacement continuation is calculated by combining the position change trend in the target representation sequences of the previous two frames. The displacement continuation is derived from the center position difference between the previous two frames and the main branch direction. The changes and regional centroid offsets are calculated. Then, the theoretical migration position and displacement continuation are used together to correct the predicted position of the candidate target in the current frame. Two types of constraint parameters are set during the correction process: one is the maximum allowable displacement distance, which is derived from the joint estimation of the current frame rate, the upper limit of the platform's flight speed, and the target's maximum relative motion speed; the other is the orientation preservation tolerance, which is derived from the stable interval of the preceding main branch extension direction. If the corrected position exceeds the task-related search domain or deviates from the tolerance range, it will fall back to the theoretical migration position. The final output predicted position is the central reference point used for target matching in the current frame, and all subsequent displacement consistency judgments are performed around this reference point. The displacement continuity constraint between the candidate target position and the predicted position in the current frame is applied to generate displacement consistency information. Specifically, the center positions of all candidate targets are extracted in the current frame, and the spatial deviation between the center position and the predicted position of each candidate target is calculated. The spatial deviation is calculated using the distance in the current frame's image coordinate system, considering both the principal direction deviation and the secondary direction deviation. The principal direction is determined by the extension direction of the main branch in the previous frame, and the secondary direction is an auxiliary direction perpendicular to the principal direction. During processing, it is first determined whether the center position of the candidate target falls within the displacement threshold region around the predicted position. The width of the displacement threshold region is determined by three parts: the minimum detection unit corresponding to the image resolution, the platform... The attitude compensation residual error and the maximum allowable displacement amplitude of the target between two frames are considered. If a candidate target exceeds the threshold region, it is directly marked as a candidate with inconsistent displacement. If it falls within the threshold region, the deviation in the main direction is further compared to see if it is significantly greater than the deviation in the secondary direction. If the deviation direction is abnormal, the displacement consistency level of the candidate is reduced. After completing the comparison of all candidates, the candidate targets are sorted in ascending order of displacement deviation, and then three categories of displacement consistency information—high consistency, medium consistency, and low consistency—are assigned according to the threshold level. The displacement consistency information formed in this way not only indicates the degree of proximity but also reflects whether the offset direction is consistent with the previous motion trend, thus providing a more detailed displacement judgment basis for subsequent temporal correlation. Scale consistency information is generated by constraining scale variation based on the distribution of outer contour span and main branch length in adjacent image frames. Specifically, the outer contour span and main branch length distribution are extracted for the confirmed target in the previous frame and the candidate target in the current frame, respectively. The outer contour span is obtained by measuring the maximum coverage of the candidate target's outer boundary in the main and secondary directions, and the main branch length distribution is obtained by statistically analyzing the length levels of the main path and the subordinate branches. Subsequently, the confirmed target in the previous frame and the candidate target in the current frame are compared item by item in terms of the proportion of the outer contour main span, outer contour secondary span, main path length, and subordinate branch length. During processing, an upper limit and a lower limit for scale variation are set. The upper limit for scale variation is converted into the allowable variation range of the main span, secondary span, and main branch length based on the change in the relative distance between the target and the platform within adjacent time intervals. If the candidate target in the current frame falls within the allowable variation range for all four scale indicators, it is assigned a high-scale consistency level; if only a few indicators exceed the range but do not exceed the limit difference threshold, it is assigned a medium-scale consistency level; if the main span, main path length, and proportion of auxiliary branches are significantly unbalanced at the same time, it is assigned a low-scale consistency level. The scale consistency information output in this step indicates whether the structural size of the same candidate target remains continuous in adjacent frames, providing a scale-level discrimination basis for distinguishing different candidate targets as co-originating continuation targets or newly emerging targets. Orientation consistency information is generated by constraining orientation changes based on the main branch extension direction and the main outer contour direction in adjacent image frames. Specifically, the main branch extension direction and the main outer contour direction are first read from the single-frame structural representation of the confirmed target in the previous frame, and then similar orientation information is extracted from the candidate targets in the current frame. The main branch extension direction comes from the start and end connection directions of the main skeleton path, and the main outer contour direction comes from the longest extension direction of the candidate target's outer contour in the image plane. Subsequently, the change amplitude of the main branch extension direction and the change amplitude of the main outer contour direction between two frames are compared, and it is further determined whether these two types of directions maintain the same or nearly the same direction relationship in the current frame. During processing, orientation change tolerance and orientation coupling are set. The orientation consistency tolerance is determined by adding the target's maximum possible rotation angle and attitude compensation residual orientation deviation within adjacent time intervals; the orientation coupling tolerance is determined by the upper bound of the deviation distribution between the main branch direction and the main direction of the outer contour in historical samples; if both types of orientation changes fall within the tolerance and the two remain stably coupled in the current frame, a high orientation consistency level is assigned; if one type of orientation change is slightly larger but the coupling relationship still exists, a medium orientation consistency level is assigned; if both types of orientations change significantly at the same time and lose their coupling relationship, a low orientation consistency level is assigned; the orientation consistency information output in this step is used to characterize the continuity of the target structure in orientation and can supplement the rotational change information that cannot be covered by displacement consistency and scale consistency. The displacement consistency information, scale consistency information, and orientation consistency information are temporally correlated to generate continuous target association results. Specifically, the displacement consistency level, scale consistency level, and orientation consistency level corresponding to each candidate target in the current frame are first jointly organized to form a candidate association description. Then, temporal correlation is performed on the candidate association description. The processing logic of temporal correlation includes two layers: the first layer is intra-frame screening, which prioritizes retaining candidates with high displacement consistency level, high scale consistency level, and high orientation consistency level in the same frame; the second layer is cross-frame continuation, which continuously compares the candidates that have passed the screening in the current frame with the tail structure of the previous target representation sequence. The system determines whether the continuation conditions for the same target are met. The continuation conditions include at least the following: changes in center position remain within displacement tolerance, changes in scale remain within scale tolerance, changes in direction remain within orientation tolerance, and the number of consecutive missing frames does not exceed the preset number of unlocked buffer frames. The number of unlocked buffer frames is derived from the platform frame rate, link backhaul rhythm, and the target's short-term occlusion tolerance duration. If a candidate meets all the continuation conditions, it is connected to the original target representation chain. If it does not meet the conditions, an independent candidate chain is started or it is directly eliminated. After all continuations are completed, the continuous target association result is obtained. This result includes at least the association chain number, the current frame access position, the number of consecutively maintained frames, and the number of consecutively missing frames. Visual lock stability is generated by extracting temporal preservation information from continuous target association results. Specifically, temporal preservation analysis is performed on continuous target association results. The analysis includes at least the number of consecutively preserved frames, the proportion of consecutively preserved frames to the total number of frames in the current analysis window, the fluctuation range of position deviation, the fluctuation range of scale change, and the fluctuation range of orientation change. The number of consecutively preserved frames is derived from the number of frames in which the same candidate association chain remains unbroken within the current analysis window. The proportion of consecutively preserved frames is derived from the ratio of this number of frames to the total number of frames in the analysis window. The fluctuation ranges of position, scale, and orientation are derived from the level fluctuation ranges of displacement consistency information, scale consistency information, and orientation consistency information in the continuous target association results, respectively. Subsequently, visual lock stability is generated according to the following rules: duration of maintenance is prioritized, continuous proportion is secondary, and fluctuation amplitude is suppressed. If a certain association chain exists continuously within the analysis window and all three types of fluctuations remain in the stable range, it is judged as high visual lock stability. If there is a short-term break but the continuous proportion is still high, or only one of the three types of fluctuations is at a medium level of change, it is judged as medium visual lock stability. If the number of consecutive frames is insufficient or the three types of fluctuations are at a high level of change for a long time, it is judged as low visual lock stability. The length of the analysis window is derived from the common requirements of the platform's minimum stable observation time and the minimum effective mapping time of the coordinate backhaul module. The output visual lock stability is sent to the subsequent coordinate backhaul module along with the continuous target association results for subsequent target geographical location calculation and geographical mapping credibility generation. Through this step, the continuous target association results are further summarized into stability parameters that can directly drive the subsequent modules, thereby completing the complete conversion from cross-frame displacement constraints to visual lock stability.

[0026] The present invention is further configured such that the step of generating the target geographic location by performing coordinate mapping on the target representation sequence, pod attitude, aircraft attitude, and platform position altitude includes: Based on the time-stamped window and return cycle in the platform scheduling command, time-stamped registration is performed on the target representation sequence, pod attitude, airframe attitude, platform position, and altitude to generate a mapped time series. Specifically, the coordinate return module first receives the platform scheduling command and reads the time-stamped window and return cycle allowed for the current round of processing. Subsequently, it calls the target representation sequence output by the air roving identification module and simultaneously reads the pod attitude data, airframe attitude data, platform position data, and platform altitude data. The pod attitude data comes from the pod angle measurement component, the airframe attitude data comes from the flight control inertial calculation results, the position data comes from the navigation and positioning component, and the altitude data comes from the navigation altitude calculation results or the altimeter component fusion results. During processing, the return cycle is first used as the basis for... The process involves selecting a set of target representation, attitude, platform position, and altitude data closest to the current transmission time within a time-stamped window. If the timestamp of any data type falls outside the window, it is discarded, and the next set of valid data within the same period is searched. If a data type is missing in the current transmission period, the previous valid sample is searched forward along the time axis, and it is determined whether the sample still falls within the allowed range of the time-stamped window. Only samples that meet the condition are allowed to be added. After the filtering is completed, the five types of data are rearranged according to a unified time reference to form a mapped time series group. Through this step, the target representation, attitude, platform position, and altitude used subsequently are in the same processing rhythm, avoiding the accumulation of coordinate mapping errors caused by the inconsistency between image time and attitude time. Based on the target representation sequence in the mapping time series, temporal stability screening is performed on the outer contour inflection points, local connection nodes, and region distribution centers to generate pixel anchor point coordinates. Specifically, the current target representation sequence is read from the mapping time series, and three types of structural points are extracted: outer contour inflection points, local connection nodes, and region distribution centers. The outer contour inflection points are derived from the boundary turning positions identified in the previous structural topological encoding, the local connection nodes are derived from the key internal connection positions obtained from skeleton unfolding and branch parsing, and the region distribution centers are derived from the overall centroid distribution results of the candidate target in the image plane. Subsequently, temporal stability screening is performed on these three types of structural points. Specifically, several consecutive frames are traced back along the target representation sequence, and the positional drift amplitude, repetition frequency, and relative positional relationship with other structural points are statistically analyzed for each type of structural point in adjacent frames. Whether the positional drift is consistent; the positional drift amplitude is used to determine whether the point remains stable in the image plane; the number of repetitions is used to determine whether the point persists; the relative positional relationship is used to determine whether the point maintains a fixed role within the target structure; the system presets an upper limit for drift, a lower limit for continuous occurrence, and an upper limit for relative relationship changes. Structural points that simultaneously meet all three conditions are judged as highly stable structural points; if there are highly stable points among the outer contour turning points, the one located near the central axis of the main direction and with the highest number of consecutive occurrences is selected as the pixel anchor point; if the outer contour turning point does not meet the requirements, the same rules are applied to the local connection nodes; if the local connection nodes still do not meet the requirements, the regional distribution center is used as the backtracking anchor point; the pixel anchor point coordinates generated in this way have a clear source, a clear selection path, and complete backtracking rules, so that subsequent line-of-sight calculations can be established on the most temporally stable image reference point; The camera gaze vector in the camera coordinate system is generated by calculating the camera gaze vector based on the mapping time series, pixel anchor coordinates, and preset camera intrinsics. Specifically, the preset camera intrinsics are called. These intrinsics are derived from the camera calibration process. During calibration, the lens focal length, principal point position, and pixel scale ratio are solved using a calibration board with known scale or a calibration target with known geometric relationships. These intrinsics are then stored and fixed during the device manufacturing or system initialization phase. During processing, the pixel anchor coordinates are first used as input to calculate the pixel offset between them and the camera imaging center. Then, the pixel offset is converted into the actual directional offset on the imaging plane according to the preset camera intrinsics. After the directional offset conversion, normalization processing is also required. This process ensures that the direction represents only the viewing direction without carrying distance scale information. The normalized result is the viewing vector in the camera coordinate system. The camera coordinate system is a local coordinate system established with the camera optical center as the origin, the lens principal axis as the forward direction, and the horizontal and vertical directions of the imaging plane as the horizontal and vertical directions, respectively. The parameters used in this step include the focal length calibration value, principal point position, pixel scale ratio, and lens distortion correction table. The lens distortion correction table is derived from the distortion compensation relationship obtained during calibration. Before conversion, the pixel anchor point coordinates are first distorted and then the viewing direction is calculated to ensure that the viewing direction obtained from the pixel coordinates can accurately reflect the spatial orientation of the target in the camera coordinate system. Based on the pod attitude, body attitude, and camera coordinate system line-of-sight vector in the mapping time series, a progressive coordinate transformation is performed to generate the navigation coordinate system line-of-sight vector. Specifically, the pod attitude in the mapping time series is read first, followed by the body attitude at the corresponding moment. The pod attitude represents the rotation state of the camera relative to the body mounting reference, and the body attitude represents the rotation state of the UAV body relative to the navigation reference. During processing, the camera coordinate system line-of-sight vector is first transformed to the pod mounting reference direction according to the pod installation calibration relationship. Then, the first layer of transformation from the local camera direction to the body direction is completed using the current pod attitude. After completing the first layer of transformation, the current body attitude is used to complete the next transformation. The second layer of transformation is from the aircraft orientation to the navigation reference orientation. The progressive coordinate transformation refers to first processing the relative orientation relationship between the camera and the pod, and then processing the overall orientation relationship between the pod and the aircraft, along with the aircraft's attitude, relative to the navigation reference. The two layers of transformation are executed sequentially. The navigation coordinate system is usually established using a unified coordinate reference for the mission area. Its horizontal plane is used to describe changes in planar position, and its vertical direction is used to describe height relationships. After completing the progressive transformation, the line-of-sight vector of the navigation coordinate system is obtained. In this way, the local line of sight, which originally only represented the observation direction within the camera, is transformed into a spatial direction that can be jointly intersected with the platform's position and height, providing direct input for subsequent landing point calculation. The target position vector is generated by calculating the landing point based on the platform position, altitude, and navigation coordinate system line-of-sight vector in the mapping time sequence group. Specifically, the platform position and altitude in the mapping time sequence group are read first, and then the target reference plane is determined. The source of the target reference plane depends on the type of the task object: when the target is on the water surface, the reference plane is taken as the current water surface elevation of the task sea area; when the target is near the ground or close to the ground surface, the reference plane is taken as the preset benchmark elevation within the task area. Then, the landing point is calculated using the platform position as the starting point of the line of sight and the navigation coordinate system line-of-sight vector as the spatial extension direction. In specific processing, it is first determined whether the vertical component of the navigation coordinate system line-of-sight vector satisfies the intersection condition. If the vertical component is too small, it means that the current line of sight is approximately parallel to the reference plane and cannot form a stable landing point. At this time, the current frame is marked as an unreliable landing point frame and enters the next return cycle. If the intersection condition is satisfied, the intersection distance from the platform position along the line of sight to the reference plane is calculated, and then this distance is applied to the navigation coordinate system line-of-sight vector to obtain the spatial position of the target in the local navigation coordinate system. This location consists of a planar position component and a corresponding elevation component, forming a target position vector. The platform position used here comes from the navigation calculation result, the platform height comes from the elevation calculation result, and the target reference plane elevation comes from the mission configuration or marine baseline data. These three factors together define the intersection relationship between the line of sight and the reference plane. Through landing point calculation, pixel anchor points in the image domain are converted into target position vectors in space, thus completing the core mapping from image to space. The target location vector is transformed using geographic coordinates to generate the target geographic location. Specifically, the obtained target location vector undergoes geographic coordinate transformation. The reference used for geographic coordinate transformation comes from the unified geographic reference frame preset by the task system. This frame includes at least the geographic location corresponding to the origin of the navigation coordinate system, the correspondence between the local planar direction and the geographic direction, and the elevation reference datum. During processing, the planar position component in the target location vector is first converted to the unified geographic planar position, and then converted to latitude and longitude coordinates or other unified geographic coordinate expressions according to the correspondence between the navigation origin and direction. At the same time, the elevation component is converted to the target height under the unified elevation reference. If the system uses local projected coordinates as the internal reference of the task, the local coordinates are back-calculated to the unified geographic coordinate format before being output to subsequent modules to ensure that the surface collaboration module and the collaborative planning module can directly read and compare them. After the geographic coordinate transformation is completed, the target geographic location is generated. This result not only inherits the spatial location meaning of the target location vector, but also has the unified coordinate attributes required for cross-platform sharing. Therefore, it can be directly used as input for unmanned surface vessels to collect candidate target status and perform correlation extrapolation.

[0027] The present invention further specifies that the credibility of the generated geographic mapping includes: The target position vector is generated by time-series correction based on the continuous target position vectors. Specifically, the target position vectors within the current transmission cycle are read first, and the historical target position vectors saved in the previous several transmission cycles are also read to form a continuous time-series position sequence. The target position vectors are derived from the aforementioned coordinate mapping results and are already under a unified navigation coordinate reference. During processing, a continuous analysis window is established with the transmission cycle as the interval. The window length is jointly determined by the minimum continuous observation time of the target, the minimum stable flight time of the UAV, and the shortest trajectory length required for subsequent surface coordination. Subsequently, time-series correction is performed on the continuous time-series position sequence. The correction rules include three items: first, position jump screening, if the current target position is different from the previous position... If the displacement exceeds the upper limit determined by the target's maximum possible speed, the transmission cycle, and the platform's attitude residual error, it is marked as an abnormal displacement sample. Secondly, there is a direction continuity check: if the current position changes abruptly relative to the direction of motion formed in the previous two moments, and this change exceeds the preset steering tolerance, the participation level of that position in the correction is reduced. Thirdly, there is a neighborhood smoothing correction: for the position vectors that were not eliminated, they are weighted and integrated according to temporal proximity and motion continuity to generate a corrected target position vector. Here, the steering tolerance comes from the reasonable upper limit of the target's rotation angle within adjacent transmission cycles and the normal direction fluctuation range in historical observations. Through temporal correction, a consistent temporal continuity relationship is formed between the current target position vector and the historical trajectory. Based on the corrected target position vector, pixel anchor coordinates, pod attitude, aircraft attitude, platform position and height, and preset camera intrinsic parameters, a back projection is performed to generate the reprojected position. Specifically, the corrected target position vector is called, and the pixel anchor coordinates, pod attitude, aircraft attitude, platform position, height, and preset camera intrinsic parameters at the current moment are read simultaneously. The preset camera intrinsic parameters are derived from the camera calibration process and include at least the focal length, principal point position, pixel scale ratio, and lens distortion correction table. The pod attitude and aircraft attitude are derived from the platform attitude measurement results. The platform position and height are derived from the navigation positioning and height calculation results. Then, back projection is performed. In specific processing, the corrected target position vector is first converted from a unified navigation coordinate reference to a coordinate system based on the current platform position. Based on the referenced local spatial relationships, and following the reverse order of the forward coordinate mapping, the transformation from navigation direction to body direction, the transformation from body direction to camera direction, and the projection from camera direction to the imaging plane are completed sequentially. Finally, combined with the preset camera intrinsic parameters and distortion correction table, the theoretical landing point of the spatial target position on the current image plane is obtained. This theoretical landing point is the reprojection position. The processing chain here fully reuses the attitude relationships and imaging parameters involved in the aforementioned forward mapping, so the reprojection position can serve as a direct basis for geometric closure verification. The pixel anchor point coordinates are synchronously retained in this step as a reference for subsequent consistency comparison, thereby establishing a one-to-one correspondence between the landing point after the spatial position is remapped back to the image and the selected stable anchor point in the actual image. Reprojection deviation information is generated by comparing the pixel anchor coordinates and reprojection positions for consistency. Specifically, the pixel anchor coordinates and reprojection positions are compared for consistency under the same image coordinate reference. During processing, the offsets of the two in the principal and secondary directions of the image are first calculated, and then it is determined whether the offset falls within a preset consistency band. The sources of the preset consistency band include camera calibration residuals, upper limit of pixel anchor temporal drift, upper limit of attitude synchronization error, and the smallest resolvable offset unit under the image resolution. If the reprojection position falls within the consistency band, it is recorded as a high reprojection consistency level; if the offset is slightly higher than the consistency band but still within the extended tolerance, it is recorded as a medium reprojection consistency level; if the offset is significantly higher than the consistency band but still within the extended tolerance, it is recorded as a medium reprojection consistency level. If the offset exceeds the extended tolerance, it is recorded as a low reprojection consistency level. In addition to the offset magnitude, the offset direction and offset holding time are also recorded. The offset direction is used to distinguish between systematic offset and random offset. If the offset is along the same direction for several consecutive backhaul cycles, it indicates that there may be a fixed direction error in the mapping chain. The offset holding time is used to distinguish between occasional offset and continuous offset. If only a single-cycle anomaly occurs, it will have a lower weight in the subsequent consistency fusion. After the above processing, reprojection deviation information is formed. This information includes at least the offset level, offset direction, continuous offset duration, and offset exceeding limit flag. It will then enter the consistency fusion process together with continuous deviation information and attitude-sensitive deviation information. Continuous deviation information is generated through temporal continuity analysis based on the corrected target position vector and the target position vectors at consecutive time points. Specifically, the corrected target position vector is used as the current position reference, and then the sequence of target position vectors at consecutive time points is called for temporal continuity analysis. During processing, the position change, direction change, and position curvature change within several consecutive feedback cycles are extracted first. The position change reflects the smoothness of displacement between adjacent time points, the direction change reflects whether the direction of motion is continuous, and the position curvature change reflects whether the trajectory exhibits abnormal reversals or abrupt bends within a short period of time. The analysis window length is jointly determined by the minimum continuous locking duration and the shortest trajectory length required for collaborative planning. Subsequently, continuity judgment rules are set: if... If the current position maintains a smooth continuation relative to the historical trajectory, and the changes in position, direction, and curvature do not exceed the corresponding tolerances, it is judged as a high level of continuity and consistency. If only a single indicator approaches the upper limit of the tolerance, and the duration of the anomaly does not exceed the preset short-term disturbance duration, it is judged as a medium level of continuity and consistency. If multiple indicators exceed the limits simultaneously, or if the trajectory shows obvious jumps, reversals, or stagnation anomalies, it is judged as a low level of continuity and consistency. The tolerance here comes from the target's reasonable motion boundary, transmission cycle, platform navigation error, and residual error of timing correction. The final continuous deviation information includes at least the displacement continuity level, direction continuity level, trajectory curvature continuity level, and anomaly duration. This information is used to characterize the degree of fit between the current target position and the historical trajectory in the time dimension. Attitude perturbation analysis is performed on the pod attitude, aircraft attitude, and corrected target position vector to generate attitude sensitivity deviation information. Specifically, using the current pod attitude and aircraft attitude as a reference, small-range attitude perturbations are applied to both, and the sensitivity of the corrected target position vector to changes during remapping is observed. Attitude perturbations originate from equipment measurement errors, control jitter, and the allowable range of flight vibrations. Their magnitude is not arbitrarily set but is determined by the accuracy of pod angle measurement, the accuracy of aircraft inertial navigation calculation, and the residual amplitude of structural vibration. During processing, upper and lower limit perturbations are first applied to the pod attitude, then to the aircraft attitude, and finally a combined perturbation is applied to both. After each perturbation, the target position vector is recalculated along the same processing chain as the forward mapping. The corresponding changes are placed in the image plane, and the magnitude of the changes is compared with the position results under the current undisturbed state. If the position change after the disturbance remains within a small range, it indicates that the current geomapping is not sensitive to attitude fluctuations and is recorded as a high attitude stability level. If the disturbance causes a position change of a moderate magnitude, it is recorded as a medium attitude stability level. If a small disturbance causes a large position change, it is recorded as a low attitude stability level. To enhance the sufficiency of disclosure, the processing also records which type of disturbance dominates the position change, that is, distinguishing between pod attitude sensitivity, body attitude sensitivity, or combined attitude sensitivity. The final attitude sensitivity bias information includes at least the sensitivity level, the dominant sensitivity source, and the sensitivity fluctuation magnitude, which is used to measure the dependence of the current geomapping results on attitude measurement and control errors. Geographic mapping credibility is generated through consistent fusion of reprojection bias information, continuous bias information, and attitude-sensitive bias information. Specifically, the generated reprojection bias information, continuous bias information, and attitude-sensitive bias information are fused consistently. Before fusion, a hard threshold screening is performed: if the reprojection bias information shows that the reprojection position continuously exceeds the extended tolerance, or the continuous bias information shows that the trajectory continuity is broken for a long time, or the attitude-sensitive bias information shows that a small attitude disturbance causes significant mapping fluctuations, then the current sample is directly marked as a low-credibility sample. For samples that are not screened out, a rank fusion is then performed. The fusion rule adopts a hierarchical processing: first, the lowest-ranked item among the three types of bias information is compared. If the lowest-ranked item appears only once and the other two items remain at a high level, it is judged as medium-high credibility; if two items are at a medium level, it is judged as medium credibility; if any item is at a low level and continues for more than a preset duration, it is judged as low credibility. The preset duration is derived from the minimum effective return duration and the minimum stable planning duration. To facilitate subsequent collaborative planning calls, the final output geographic mapping credibility is expressed using an ordered level or normalized credibility value, which essentially reflects the comprehensive consistency of the current target geographic location in terms of geometric closure, temporal continuity, and attitude stability. After consistency fusion is completed, the geographic mapping credibility, together with the aforementioned target geographic location, enters the subsequent surface collaboration module and collaborative planning module as a key input for target association, fan-ring interception planning, and interception command generation.

[0028] The present invention is further configured such that the step of associating and extrapolating the target geographic location and candidate target state to generate target coupling consistency and target predicted state includes: The target's geographical location is mapped locally to generate a reference position. Specifically, the target's geographical location is first received from the coordinate feedback module. This target geographical location is expressed using a unified geographic reference datum, which is suitable for cross-platform transmission. However, the target status output by the UAV's local sensors is usually expressed using local planar coordinates. Therefore, local coordinate mapping is required first. During processing, the current navigation position of the UAV, the local coordinate origin of the UAV, and the direction reference of the current working area are read first. The local coordinate origin comes from the current navigation calculation result of the UAV or the preset working origin when the mission starts. The direction reference comes from the unified planar projection direction used in the mission sea area. Subsequently, the target's geographical location is converted into a planar position relative to the local origin of the unmanned surface vessel (USV), and the relative orientation of the target to the USV is preserved in the same plane. After mapping, a reference position is generated. The significance of this reference position is that it unifies the position results transmitted back from the aerial platform with the local perception results of the USV under the same planar coordinate framework, so that subsequent candidate target selection, gating matching, and trajectory extrapolation can all be carried out around the same spatial reference. The parameters used in the local coordinate mapping process include the navigation origin, the planar orientation reference, and the geographic-to-planar conversion ratio. These parameters are all derived from the mission coordinate reference configuration and the USV navigation system, and no additional external input is required. A local search domain is generated based on the reference position. The unmanned surface vessel (USV) is controlled to collect the position, heading, and speed of candidate targets within this domain, generating a candidate target state set. Specifically, a local search domain is constructed centered on the generated reference position. The extent of the local search domain is determined by three factors: first, the spatial uncertainty of the target's geographical location itself, derived from the credibility level corresponding to the geographic mapping credibility; second, the possible displacement range caused by air-sea transmission delays, derived from the current transmission cycle, link latency, and the target's maximum reasonable speed; and third, the effective detection range of the USV's local sensors, derived from the current operating parameters of radar, vision, or fusion sensing devices. During processing, a basic search range is first generated centered on the reference position, and then this range is scaled and corrected based on the geographic mapping credibility: when the credibility is high... The search domain is narrowed, and then appropriately widened when the confidence level is low. Subsequently, the unmanned surface vessel's sensing equipment is controlled to perform target extraction only within the local search domain, and a candidate target state is formed for each detected object. The candidate target state includes at least planar position, heading, and velocity. The planar position comes from the sensor positioning results, the heading comes from the continuous contour direction change of the target or the radar track direction estimation, and the velocity comes from the position change and track time difference between adjacent observation periods. To ensure the comparability of subsequent correlations, all candidate target states are uniformly time-scaled and coordinate-normalized before output, thus forming a candidate target state set. Through the constraint of the local search domain, the sensing range of the unmanned surface vessel is limited to the area directly related to the aerial transmission target, and the candidate state set can avoid introducing a large number of background targets that are irrelevant to the current task. The target's geographical location and the previously associated target state are used to perform motion completion to generate a reference target state. Specifically, after obtaining the reference position, the reference target state is further generated. Since the current aerial transmission results mainly provide the position, while subsequent gating and extrapolation require complete position, heading, and velocity states, motion completion is necessary. During processing, the previously confirmed associated target state is first read. The previously confirmed associated target state comes from the output of the previous round of surface coordination module and includes at least the target's position, heading, and velocity from the previous moment. If a previously confirmed associated target state exists, its heading and velocity are used as the initial motion reference for the current moment. Then, the position is updated by combining it with the reference position corresponding to the current target's geographical location. The current reference target state is established. If there is no associated target state from the previous moment, such as when entering the tracking process for the first time or when the previous association is interrupted, the preset initial speed and preset initial heading are called to form the initial motion reference. The preset initial speed comes from the median of the prior speed range of the task object, and the preset initial heading comes from the main direction of the local search domain or the statistical results of the most recent round of candidate trajectory directions of the UAV sensor. After motion completion, the reference target state includes three parts: current position, current reference heading, and current reference speed. This reference target state serves as a unified matching benchmark in subsequent spatial gating and motion gating, so that candidate targets are no longer compared only around spatial position, but are filtered around the complete motion state. Spatial gating and motion gating are applied to the reference target state and the candidate target state set to generate an associative candidate set. Specifically, the reference target state is compared with the candidate target state set one by one, and the associative candidate set is selected in the order of spatial gating first and then motion gating. The inputs for spatial gating include the reference position, the planar position of the candidate target, and the boundary of the local search domain. During processing, the planar offset distance from each candidate target to the reference position is calculated first, and then it is determined whether the offset falls within the preset spatial threshold range. The sources of the spatial threshold include the reliability of geographic mapping, the radius of the local search domain, and the upper limit of the local perception error of the unmanned surface vessel. After spatial gating, only candidates whose positions are close to the reference target are retained. For candidates that pass spatial gating, motion gating is further performed. The inputs for motion gating are... The input includes the heading and velocity of the reference target in its state, as well as the heading and velocity of the candidate target in its state. During processing, the deviations of the candidate target from the reference target in terms of heading difference and velocity difference are compared, and then a preset motion threshold is used to determine whether to retain it. The motion threshold is derived from the upper limit of the target's possible maneuverability, the change range of the target's state at the previous moment, and the accuracy of the UAV's local velocity estimation. If a candidate is close in position but deviates too much in heading or velocity, it is not included in the associative candidate set. If the position, heading, and velocity are all within the allowable range, it is included in the associative candidate set. After two levels of gating, the associative candidate set has eliminated objects with significant spatial deviations and significant mismatches in motion characteristics. Subsequent temporal association only needs to complete more refined target confirmation within this restricted set. The associated candidate set and the associated target state at the previous time step are temporally associated to generate the associated target state and the target coupling consistency; specifically, temporal association is performed on each candidate in the associated candidate set. The data used for temporal correlation includes: the current candidate's position, heading, and velocity; the position, heading, and velocity of the correlated target in the previous moment; and the current reference target's state. During processing, the system first compares the spatial continuity between the current candidate and the correlated target in the previous moment to determine if their displacement direction is consistent with the previous movement direction. Then, it compares the continuity of the current candidate and the correlated target in terms of heading and velocity changes to determine if their turning and acceleration / deceleration remain within a reasonable range. Subsequently, the current candidate and the current reference target's state are compared synchronously to verify whether they simultaneously meet the dual constraints of the airborne return position and historical trajectory. For each candidate, the system provides a spatial continuity level, heading continuity level, velocity continuity level, and reference position consistency level, and then generates a total correlation level according to preset joint rules. If the total correlation level reaches a high level, the candidate is confirmed as the correlated target state for the current moment. If several candidates simultaneously reach a medium level, the candidate with the higher spatial continuity level and reference position consistency level is selected first. If all candidates are below the minimum correlation requirements, no new correlated target state is generated for the current moment, and downstream waiting or reacquisition preparation logic is triggered. After the candidate selection is completed, the target coupling consistency is further formed based on the spatial continuity level, heading continuity level, speed continuity level and reference position consistency level; this consistency represents the comprehensive matching degree between the air-transmitted target and the surface-associated target at the current moment. Based on the associated target state, the trajectory extrapolation generates the target prediction state. Specifically, starting from the current associated target state, it calls the associated target state from the previous moment and the historical associated state from an even earlier moment to perform short-term trajectory extrapolation. Before extrapolation, the motion mode of the current associated target is determined. The motion mode is determined based on the speed change trend and heading change trend over several consecutive moments: if the speed change and heading change are small, it is judged as a uniform speed straight-ahead mode; if the speed is basically stable but the heading changes continuously, it is judged as a uniform speed turning mode; if the speed change is significant and the heading change is small, it is judged as a variable speed straight-ahead mode; if both speed and heading change significantly, it is judged as a compound maneuver mode. During processing, the corresponding extrapolation rule is selected according to the currently determined motion mode. For the uniform speed straight-ahead mode, it is extended forward along the current heading direction according to the current speed and the predicted time step; for the uniform speed turning mode, while maintaining the current speed, it is extended according to the most recent continuous trajectory. The direction of the next moment is corrected according to the rate of change, and then extrapolated along the corrected direction. For the variable speed straight-ahead mode, the speed is adjusted according to the recent continuous speed change trend and extrapolated while keeping the current heading unchanged. For the compound maneuver mode, the trends of the recent continuous speed change and heading change are used simultaneously, and a short extension is performed within the preset stable extrapolation time. The prediction time step here comes from the minimum planning cycle of the collaborative planning module, and the stable extrapolation time comes from the target state update frequency and the unmanned vessel path replanning rhythm. After the extrapolation is completed, the target prediction state is generated. The target prediction state includes at least the predicted position, predicted heading and predicted speed, and is sent to the collaborative planning module along with the target coupling consistency.

[0029] The present invention is further configured such that the generation of the cross-domain collaborative consistency index based on visual lock stability, geographic mapping credibility, and target coupling consistency includes: Consistent central information is generated by centrally aggregating visual lock stability, geographic mapping reliability, and target coupling consistency. Specifically, the collaborative planning module first receives visual lock stability, geographic mapping reliability, and target coupling consistency, and performs simultaneous verification and unified dimensional adjustment on the three inputs. Unified dimensional adjustment means converting all three inputs into comparable values ​​within the same evaluation interval to prevent subsequent aggregation distortion caused by different evaluation methods in previous modules. After adjustment, the central aggregation rule is invoked to generate consistent central information. The central aggregation rule adopts a bounded weighted aggregation method, first assigning aggregation contribution to each of the three inputs according to preset weights. The weight of visual lock stability is determined by the priority of the image recognition chain in the current task, the weight of geographic mapping reliability is determined by the importance of the spatial positioning chain, and the weight of target coupling consistency is determined by the importance of the surface collaboration chain. The weights are derived from the task initialization configuration, or initial values ​​can be obtained from historical operation statistics and then fixed in the system parameter table. Two types of constraints are set during the aggregation process: one is the minimum effective input threshold. If any of the three inputs is below the minimum effective threshold, it is marked as a weak support quantity and its participation level in the central aggregation is reduced. The other is the high-confidence locking threshold. If two of the three inputs reach the high-confidence level, the corresponding aggregation share is increased. After aggregation is completed, consistency center information is output. This information represents the common support level of the three inputs at the current moment, providing a reference benchmark for subsequent dispersion analysis, and enabling the collaborative planning stage to process data from the three chains of air patrol identification, coordinate backhaul, and surface coordination using the same evaluation center. Visual lock stability, geographic mapping reliability, and target coupling consistency are analyzed with consistency center information to generate consistent discrete information. Specifically, using the generated consistency center information as a reference, discrete analysis is performed on each of the three inputs. The discrete analysis includes deviation magnitude analysis, deviation direction analysis, and deviation persistence analysis. Deviation magnitude analysis is used to determine whether the deviation of each input relative to the consistency center is within the allowable range. The allowable range is derived from the acceptable fluctuation range of the three processing chains in historical samples and the current balance requirements of the task for aerial identification, surface association, and spatial positioning. Deviation direction analysis is used to determine whether the three inputs increase or decrease simultaneously, or whether there is an imbalance where one is significantly higher or lower than the other two. Deviation persistence analysis... The analysis is used to determine whether a deviation is a single-cycle fluctuation or a sustained state over several cycles. Specifically, the three inputs are first subtracted from the consistency center information, and then categorized into three levels based on their deviation range: slight deviation, moderate deviation, and significant deviation. The three deviation levels are then combined to determine whether the current state is equilibrium, a single-item imbalance, or a double-item imbalance. Next, the consistency center information from the most recent several cycles and the historical sequence of the three inputs are used to analyze whether the current deviation persists. If all three inputs are close to the center information, a low dispersion level is generated; if one input deviates significantly but for a short period, a moderate dispersion level is generated; if one or two inputs deviate from the center information for a long period with a stable direction, a high dispersion level is generated. After processing, consistency dispersion information is output. This information not only reflects the degree of dispersion among the three inputs but also retains the source and duration of the deviation, facilitating the subsequent consistency fusion stage in distinguishing between two different states: one with a high overall level but internal imbalance, and the other with a moderate overall level but internal coordination. A cross-domain collaborative consistency index is generated by fusion based on visual lock stability, geographic mapping credibility, target coupling consistency, and consistency discrete information. Specifically, visual lock stability, geographic mapping credibility, and target coupling consistency are used as basic support quantities, and consistency discrete information is used as a coordination correction quantity to perform consistency fusion processing. The fusion processing is completed using a hierarchical decision-making approach. First, a basic level determination is performed, and a basic consistency level is formed based on the high, medium, and low combinations of the three basic support quantities. The basic consistency level reflects the overall support strength of the three processing chains at the current moment. Subsequently, a discrete correction judgment is performed. Based on the discrete level, deviation source, and persistence status recorded in the consistent discrete information, an upper limit constraint or a reduction in order is applied to the basic consistency level. When the discrete level is low, it indicates that the three inputs are coordinated around the central information, and the basic consistency level can be retained. When the discrete level is medium, it indicates that there is a short-term imbalance in the local link, and the basic consistency level needs to be reduced to a limited level. When the discrete level is high and the duration exceeds the preset continuous period, it indicates that there is a significant incoordination relationship between the three chains of air patrol identification, coordinate backhaul, and surface coordination. At this time, a significant reduction in order is performed on the basic consistency level, and the result is marked as a coordination-sensitive state. The preset continuous period comes from the common requirements of the minimum effective planning window and the minimum stable interception control window, ensuring that the reduction correction is based on persistent imbalance rather than occasional fluctuations. Finally, the corrected result is mapped to a cross-domain coordination consistency index. The index can be expressed in an ordered hierarchical form or in a normalized interval value form. The specific form is determined by the interface agreement between the collaborative planning module and the subsequent fan-ring interception planning module. The output cross-domain collaborative consistency index will be directly entered into the fan-ring interception planning process to adjust the fan-ring domain scale, the assessment intensity of the receiving point, and the encapsulation level of the interception command. This ensures that the planning stage calls a comprehensive consistency quantity that simultaneously reflects the overall support level and the degree of internal coordination.

[0030] The present invention is further configured such that the process of generating a coordinated path and interception instructions by performing sector-loop interception planning based on the target prediction status, mission sea area boundary, restricted area boundary, and air-sea communication link status includes: Based on the target prediction status, mission sea area boundary, restricted area boundary, and air-sea communication link status, constraint analysis is performed to generate the planning feasible region and constraint cost field. Specifically, the target prediction status, mission sea area boundary, restricted area boundary, and air-sea communication link status are first read, while the current position, current course, and current speed of the unmanned vessel are simultaneously retrieved as path starting point information. The mission sea area boundary is represented by a closed boundary point sequence, and the restricted area boundary is represented by a set of closed boundary point sequences, both of which originate from the mission data output by the early warning and scheduling module. The air-sea communication link status is derived from the link quality monitoring results established during the scheduling phase, including at least the link availability level and link continuity level. The target prediction status is derived from the trajectory extrapolation results of the surface coordination module, including at least the predicted position, predicted course, and predicted speed. During processing, the mission sea area boundary and restricted area boundary are first subjected to region clipping, removing regional units falling within the restricted area and retaining regional units located within the mission sea area that do not overlap with the restricted area, forming the initial passable area. Subsequently, planning sampling units are established based on the current scale accuracy of the unmanned surface vessel (USV). For each sampling unit, its proximity to the boundary of the mission sea area and its proximity to the boundary of the restricted area are calculated. Then, the air-sea communication link status is mapped onto these sampling units to obtain the link availability level corresponding to each location. After that, constraint resolution is performed: units located outside the mission sea area or inside the restricted area are directly judged as impassable units; units located inside the mission sea area but close to the boundary buffer zone, the restricted area buffer zone, or the low-availability link zone are assigned a higher passage cost; units far from the boundary, far from the restricted area, and with stable links are assigned a lower passage cost. Here, the width of the boundary buffer zone is derived from the minimum turning distance occupied by the USV, the upper limit of navigation error, and the displacement redundancy corresponding to the command transmission lag. The low-availability link zone threshold is derived from the minimum communication conditions allowed by the control closed loop. After the above processing, a planning feasible region and a constraint cost field are obtained. The former is used to limit the path search range, and the latter is used to distinguish the passage priority of different locations within the same feasible region. The target prediction state and cross-domain coordination consistency index are used to collaboratively map and generate a fan-ring interception domain. Specifically, the predicted position and predicted heading in the target prediction state are used as the central information, and then combined with the cross-domain coordination consistency index to generate the fan-ring interception domain. In specific processing, the target's backward region is first determined based on the target's predicted heading, that is, the region in the opposite direction of the predicted heading is used as the priority direction for the unmanned surface vessel (USV) to intercept. Then, three types of fan-ring parameters are set: inner interception distance, outer interception distance, and angle. The inner interception distance comes from the minimum safe distance that the USV must maintain when approaching the target, the outer interception distance comes from the USV's current speed capability, the target's predicted speed, and the latest allowed interception distance, and the angle comes from the USV's lateral maneuverability and the target's possible short-term turning range. The cross-domain coordination consistency index is used here to adjust the fan-ring parameters: when the consistency index is high, it indicates that the three supporting relationships of visual lock stability, geographic mapping reliability, and target coupling consistency are relatively coordinated, and the fan-ring range can be appropriately converged to make the interception area more concentrated; when the consistency index is at a medium level, the fan-ring range is appropriately... The degree of expansion allows the receiving area to retain a certain degree of mobility redundancy; when the consistency index is low, the fan ring range is further widened and the outer receiving distance is extended to improve the fault tolerance space for subsequent receiving point evaluation; after parameter adjustment, the fan ring interception domain is intercepted within the planned feasible domain with the target prediction position as the center and the backward direction as the main axis; cooperative mapping refers to mapping the target motion prediction and cross-domain cooperative consistency into a geometric interception space at the same time. It does not depend solely on the target position or the consistency index, but rather transforms both into regional constraint results that can be called by path planning. Interception anchor points are generated based on the sector-ring interception domain and the constraint cost field. Specifically, candidate interception points are deployed within the sector-ring interception domain. The deployment density of candidate interception points is derived from the minimum trajectory adjustment scale of the unmanned surface vessel (USV) and the planned sampling resolution, ensuring that the interval between each candidate point is not less than the USV's control resolution capability. Subsequently, an interception point evaluation is performed on each candidate interception point. The acceptance point evaluation considers at least three factors: First, the reasonableness of the candidate point's backward acceptance relative to the predicted target position, i.e., whether the point is located in the priority area of ​​the target's main backward acceptance direction; second, the accessibility level of the candidate point in the constraint cost field, i.e., whether the point is adjacent to the boundary risk zone, the restricted area buffer zone, or the link degradation zone; third, the ease of approach for the unmanned vessel to reach the point from its current position, i.e., whether the required path turn is gentle and the path length is within the allowable range. During processing, candidate points falling in high-cost areas are first eliminated, and then the remaining candidate points are sorted in a first-level order according to the reasonableness of backward acceptance. The top-ranked candidate points are further sorted in a second-level order according to the ease of approach. Finally, the candidate point with the highest score is selected as the interception anchor point. The acceptance point evaluation is derived from the spatial acceptance relationship requirements of air-sea cooperative interception. The core is to allow the unmanned vessel to enter the target's backward interception area with an executable path, rather than directly selecting the point closest to the target. The interception anchor point generated by this step is both within the fan-ring interception domain and coordinated with the constraint cost field and the unmanned vessel's motion conditions, thus it can serve as the direct endpoint for cooperative path solving. A cooperative path is generated by solving for the current position of the unmanned surface vessel (USV) and the interception anchor point. Specifically, the current position of the USV is used as the starting point of the path, and the interception anchor point is used as the ending point of the path. The path is solved within the planned feasible region. The path solution adopts a step-by-step expansion method with constrained costs. In the specific process, several candidate propulsion directions are first generated around the current position of the USV, and then gradually expanded towards the interception anchor point. At each expansion step, it is checked whether the position is still within the planned feasible region, and the passage cost in the constraint cost field is read. If an expansion direction crosses an impassable area or enters an excessively high cost area, the expansion along that direction is stopped. If it is still within the allowable range, the extension continues and the cumulative path cost is recorded. The cumulative path cost consists of three parts: path length cost, used to constrain range consumption; proximity risk cost, used to constrain the path from approaching boundaries and restricted areas; and direction turning cost, used to constrain the path change to be gradual and avoid frequent large-angle maneuvers by the unmanned vessel. The direction turning cost is derived from the minimum turning radius and control response limit of the unmanned vessel, the path length cost is derived from the mission's allowed arrival time, and the proximity risk cost is derived from the constraint cost field generated in step one. After searching all candidate paths, the one with the minimum cumulative cost is selected as the cooperative path. The output cooperative path not only includes a continuous path point sequence from the current position to the interception anchor point, but also includes the direction sequence and turning sequence of the path segment. Interception commands are generated by pattern encapsulation based on the cross-domain collaboration consistency index, sector-ring interception domain, interception anchor point, target prediction state, and collaboration path. Specifically, pattern encapsulation is performed on the aforementioned results to generate interception commands. The inputs to pattern encapsulation include the cross-domain collaboration consistency index, sector-ring interception domain, interception anchor point, target prediction state, and collaboration path. During processing, the current interception mode is first determined based on the cross-domain coordination consistency index. Interception modes are categorized into at least three types: tightly coupled interception mode, conservative accompanying mode, and waiting-to-recapture mode. The tightly coupled interception mode is suitable for a consistently high consistency index, instructing the execution module to track the coordinated path with high convergence. The conservative accompanying mode is suitable for a medium consistency index, instructing the execution module to maintain significant redundancy within the sector-ring interception domain. The waiting-to-recapture mode is suitable for a low consistency index or continuously fluctuating consistency index, instructing the execution module to reduce approach intensity and wait for subsequent state updates. After the interception mode is determined, the mode identifier, target prediction status, sector-ring interception domain range, interception anchor point position, and coordinated path are encapsulated into an interception command. The interception command fields include at least: target prediction position, target prediction heading, target prediction speed, interception mode, interception anchor point position, sector-ring interception domain boundary description, coordinated path point list, and path execution order. After encapsulation, the interception command is sent to the interception execution module.

[0031] The present invention is further configured such that controlling the UAV and unmanned surface vessel to perform cooperative interception and generate an interception state based on the cooperative path and interception command includes: The collaborative path, interception command, current UAV status, and current UAV status are time-calibrated to generate an execution state group. Specifically, the interception execution module first receives the collaborative path and interception command output by the collaborative planning module, and simultaneously reads the current UAV status and UAV status in real time. The current UAV status includes at least the current position, current heading, current speed, and current altitude, while the current UAV status includes at least the current position, current heading, and current speed. The collaborative path contains a sequentially arranged list of path points, path segment directions, and path segment turning information. The interception command includes the target prediction status, interception anchor point, fan-ring interception domain, and interception mode. During processing, a control cycle is first established as a reference. A unified timeline is used to align the collaborative path, interception commands, and dual-platform status. If an input is earlier than the current control cycle, its time deviation is checked to see if it falls within the preset synchronization tolerance range. Data within the tolerance range is allowed to participate in the current control cycle, while data exceeding the tolerance range enters a waiting or update process. The synchronization tolerance is derived from the link transmission delay, control cycle length, and upper limit of the actuator response time. After time alignment, the collaborative path, interception commands, current UAV status, and current UAV status are organized into an execution status group. This execution status group is continuously invoked in subsequent steps to ensure that UAV control and UAV control are based on the planning results and platform status at the same moment. The interception command analyzes the target prediction state, interception anchor point, and fan-ring interception domain. Based on the target prediction state and interception anchor point, a UAV execution reference is generated. The current UAV state and the UAV execution reference are then matched to generate UAV control commands. Specifically, the target prediction state, interception anchor point, and fan-ring interception domain are analyzed from the execution state group, and the UAV execution reference is generated accordingly. During processing, the target's main observation direction is first determined based on the target's predicted position and predicted heading. Then, the UAV's priority observation side is determined by combining the interception anchor point. The priority observation side refers to the observation direction that can both cover the target's predicted position and provide overhead guidance for the UAV's journey to the interception anchor point. Subsequently, the UAV execution reference is generated based on the mission-preset observation radius, observation altitude, and field of view coverage angle. The observation radius is derived from the effective coverage range of the target and interception area by the electro-optical payload at the current altitude, and the observation altitude is derived from the UAV's... The minimum flight altitude, field-of-view resolution requirements, and link maintenance conditions are defined. The field-of-view coverage angle is derived from the pod's optical axis adjustment capability and the range of changes in the target's relative motion direction. After generating the UAV execution reference, the current state of the UAV is matched against this execution reference. The deviation matching includes at least position deviation, heading deviation, and altitude deviation. Position deviation represents the difference between the UAV's current horizontal position and the reference observation position; heading deviation represents the difference between the UAV's current orientation and the target's observation direction; and altitude deviation represents the difference between the UAV's current altitude and the reference observation altitude. The module converts these three types of deviations into UAV control commands according to preset control gains and platform motion constraints. The UAV control commands include at least speed adjustment, heading adjustment, and altitude adjustment. The control commands output in this step will serve as the basic input for subsequent collaborative corrections, while maintaining continuous coverage of the target and interception area by the UAV. The unmanned surface vessel (USV) execution reference is generated based on the cooperative path and interception anchor point. The current USV state and the execution reference are then matched to generate control commands. Specifically, the cooperative path and interception anchor point are read from the execution state group, and the USV execution reference is generated with the current USV position as the starting point and the interception anchor point as the target and ending point. During processing, the nearest path point corresponding to the USV's current position is first located on the cooperative path. Then, the next reference point within a preset forward look-ahead distance is searched along the path as the path tracking target for the current control cycle. The forward look-ahead distance is derived from the USV's current speed, turning radius, and control refresh cycle. Its purpose is to maintain path continuity while avoiding frequent large-angle corrections due to excessively short forward look-ahead distances. Subsequently, the USV reference heading is generated based on the directional relationship between the nearest path point and the forward look-ahead reference point, and the heading is determined based on the distance from the current position to the forward look-ahead reference point. The remaining distance is used to generate a reference thrust for the unmanned surface vessel (USV). After generating the execution reference, the current state of the USV is matched with the USV execution reference for deviation. The deviation matching includes at least lateral path deviation, longitudinal path thrust deviation, and heading deviation. The lateral path deviation represents the degree of deviation of the USV's current position from the centerline of the cooperative path. The longitudinal path thrust deviation represents the difference between the USV's current thrust progress and the reference thrust progress. The heading deviation represents the difference between the USV's current orientation and the reference heading. Based on the USV's current speed capability, rudder angle change limit, and minimum stability maneuvering requirements, the module converts these deviations into USV control commands. The USV control commands include at least thrust adjustment and steering adjustment. The control commands output in this step ensure that the USV continues to move towards the interception anchor point along the cooperative path and provides a surface execution baseline for subsequent cooperative corrections. The system collaboratively corrects the target prediction status, the fan-ring interception domain, UAV control commands, and UAV control commands to generate a collaborative interception control quantity, which is then sent to the UAV and UAV actuators. Specifically, after receiving individual UAV and UAV control commands, further collaborative corrections are performed. The purpose of collaborative correction is to maintain synchronization between the observation rhythm of the aerial platform, the approach rhythm of the surface platform, and the target prediction status. During processing, the system first determines the target's expected position and heading changes in the next control cycle based on the target prediction status. Then, using the fan-ring interception domain as a geometric constraint, it checks whether the UAV will still enter the fan-ring allowable area after advancing according to the current control commands. Simultaneously, it checks whether the UAV can still maintain simultaneous coverage of the target prediction area and the UAV approach area after adjusting according to the current control commands. If the UAV's expected entry path will cross the inner boundary of the fan-ring too early, the advance adjustment amount is reduced; if it is expected to enter the fan-ring domain late, the advance adjustment amount is increased; if the UAV's expected observation position will deviate from the target prediction area, the heading and position correction amounts are increased; if the UAV's expected observation angle will obscure the UAV's subsequent receiving direction, the reference observation side is adjusted. Next, the air-sea timing coordination relationship is calculated, specifically whether the time when the unmanned vessel is expected to arrive at the interception anchor point and the time when the target is expected to enter the interception zone are within the allowable difference range. The allowable difference range is derived from the unmanned vessel's minimum stable guidance time, the target status feedback cycle, and the UAV's supplementary observation margin. After completing the above corrections, a coordinated interception control quantity is formed. This control quantity includes the speed, heading, and altitude adjustment values ​​ultimately issued to the UAV's actuators, as well as the propulsion and steering adjustment values ​​ultimately issued to the unmanned vessel's actuators. Through this process, the control results of the UAV and the unmanned vessel are no longer two independent sets of instructions, but rather a unified execution quantity that coordinates and adjusts around the same target prediction state and the same sector interception domain. Interception status is generated based on UAV execution reference, UV execution reference, cooperative interception control variables, and the sector-ring interception domain. This interception status is then sent to the upper-level mission state machine. Specifically, at the end of each control cycle, a status determination is performed on the current interception process. The inputs for this status determination include UAV execution reference, UV execution reference, cooperative interception control variables, and the sector-ring interception domain. Simultaneously, the current actual UAV position, UV actual position, and target predicted position are used for auxiliary judgment. During processing, it is first determined whether the UAV execution status meets the observation-holding conditions, including whether the UAV is within the reference observation range, whether its heading covers the target's predicted direction, and whether its field of view covers the UV's receiving area. Then, it is determined whether the UV execution status meets the path-holding conditions, including whether the UV is within the cooperative path allowable deviation zone, whether its propulsion direction still points to the interception anchor point, and whether it is expected to remain within the sector-ring interception domain's allowable approach zone at the next moment. Finally, the sector-ring interception domain is considered in conjunction with the final determination. The spatial relationship between the current target and the unmanned surface vessel (USV) is determined to confirm whether the USV has entered the interception and reception phase, reached the stable companion phase, or experienced a mismatch due to deviation from the fan-ring domain. Based on the above determinations, the module generates an interception state. The interception state is divided into at least three types: cooperative approach state, cooperative interception state, and cooperative mismatch state. The cooperative approach state indicates that both the USV and the USV are approaching the target reception area according to the reference requirements. The cooperative interception state indicates that the USV has entered the fan-ring interception domain and the USV maintains effective observation coverage. The cooperative mismatch state indicates that there is a significant break in the USV's observation chain or the USV's reception chain. After generating the interception state, it is sent to the upper-level task state machine. The upper-level task state machine decides whether to maintain the current interception mode, switch to a conservative companion mode, or enter a waiting recapture and replanning process based on the state. Through this state determination process, the interception execution module outputs not only control results but also forms a phased feedback on the entire cooperative interception process, enabling the system to have closed-loop task management capabilities.

[0032] It should be noted that the specific methods by which each module and unit performs operations in the vision-based cross-domain unmanned platform collaborative inspection system provided in the above embodiments have been described in detail in the system embodiments and will not be repeated here. In practical applications, the vision-based cross-domain unmanned platform collaborative inspection system provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above, and this is not a limitation.

[0033] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A vision-based cross-domain unmanned platform collaborative inspection system, characterized in that, include: Early warning and dispatch module: Receives the boundary of the task sea area, the boundary of the restricted area, the status of the air and sea communication links and the early warning location, and performs task orchestration to generate inspection task sets and platform dispatch instructions; The aerial patrol identification module controls the UAV to collect image frames, pod attitude, body attitude and platform position and altitude based on the patrol task set. It performs structural topological coding and cross-frame displacement constraints on the image frames to generate target representation sequences and visual lock stability. Coordinate return module: Performs coordinate mapping on the target representation sequence, pod attitude, body attitude and platform position and altitude according to the platform scheduling instructions to generate the target geographical location and geographical mapping reliability; Surface Coordination Module: Based on the target's geographical location, the module controls the unmanned vessel to collect the status of candidate targets, correlates and extrapolates the target's geographical location and the candidate target status, and generates the target coupling consistency degree and the target predicted status. Collaborative planning module: Based on visual lock stability, geographic mapping credibility and target coupling consistency, a cross-domain collaborative consistency index is generated. The module performs fan-loop interception planning based on the target prediction status, mission sea area boundary, restricted area boundary and air-sea communication link status to generate collaborative paths and interception commands. Interception Execution Module: Controls UAVs and unmanned vessels to perform collaborative interception and generate interception status based on the collaborative path and interception command.

2. The vision-based cross-domain unmanned platform collaborative inspection system according to claim 1, characterized in that, The execution of task orchestration generates inspection task sets and platform scheduling instructions, including: Receive information on the boundaries of the mission sea area, the boundaries of restricted areas, the status of air and sea communication links, and early warning locations; Determine the schedulable area based on the boundaries of the mission sea area and the restricted area, and extract boundary constraint information; Link quality analysis is performed based on the status of air and sea communication links to generate link availability information; Constraints are adjusted on the warning location, schedulable area, boundary constraint information, and link availability information to generate the warning seed area and the core location of the task. Based on the early warning seed area, the core location of the mission, the current location of the UAV and the current location of the unmanned vessel, a set of inspection missions is generated through scheduling and matching. Based on the inspection task set, timing coordination and instruction encapsulation are performed to generate platform scheduling instructions, which include time-stamped windows and feedback cycles.

3. The vision-based cross-domain unmanned platform collaborative inspection system according to claim 1, characterized in that, Performing structural topological coding on image frames to generate target representation sequences includes: Based on the task-related search domain in the image frame defined by the inspection task set, candidate target regions are extracted within the task-related search domain. Based on the boundary lines of the candidate target region, perform contour segmentation and turn positioning to generate outer contour structure information; Based on the internal connectivity paths of the candidate target region, skeleton expansion and branch parsing are performed to generate local connectivity structure information; Generate regional distribution structure information based on the centroid location, main branch extension direction, and regional envelope relationship of the candidate target region; Topological association encoding is performed on external contour structure information, local connectivity structure information and regional distribution structure information to generate single-frame structural representations of candidate targets; The structural representations of a single frame are arranged according to the image acquisition time sequence to generate a target representation sequence.

4. The vision-based cross-domain unmanned platform collaborative inspection system according to claim 3, characterized in that, Executing cross-frame displacement constraints to generate visual lock stability includes: Attitude compensation results are generated by performing attitude compensation based on the pod attitude and body attitude corresponding to adjacent image frames; The predicted position of the candidate target in the current frame is determined based on the target representation sequence and pose compensation results in adjacent image frames; Displacement continuity constraints are applied to the candidate target positions and predicted positions in the current frame to generate displacement consistency information; Scale consistency information is generated by constraining scale variation based on the distribution of outer contour span and main branch length in adjacent image frames. Orientation consistency information is generated by constraining the orientation change based on the main branch extension direction and the main direction of the outer contour in adjacent image frames. The displacement consistency information, scale consistency information, and orientation consistency information are correlated in a time sequence to generate continuous target correlation results; Visual lock stability is generated by extracting temporal preservation information based on the results of continuous target association.

5. The vision-based cross-domain unmanned platform collaborative inspection system according to claim 1, characterized in that, Generate the target's geographic location by performing coordinate mapping on the target representation sequence, pod attitude, airframe attitude, and platform position and altitude, including: Based on the time-stamped window and transmission cycle in the platform scheduling instructions, time-stamped registration is performed on the target representation sequence, pod attitude, airframe attitude, platform position and altitude to generate a mapped time series group; Based on the target representation sequence in the mapping time series, the pixel anchor coordinates are generated by time series stability screening of the outer contour inflection points, local connection nodes and regional distribution centers. The camera gaze vector in the camera coordinate system is generated by calculating the camera gaze vector based on the mapped time series group, pixel anchor point coordinates and preset camera intrinsic parameters. Based on the pod attitude, body attitude, and camera coordinate system line-of-sight vector in the mapping timing group, a progressive coordinate transformation is performed to generate the navigation coordinate system line-of-sight vector; The target position vector is generated by solving the landing point based on the platform position, altitude, and navigation coordinate system line-of-sight vector in the mapped time series group; The target location vector is transformed into geographic coordinates to generate the target geographic location.

6. The vision-based cross-domain unmanned platform collaborative inspection system according to claim 5, characterized in that, The credibility of generated geomaps includes: The corrected target position vector is generated by performing time-series correction on the target position vectors at consecutive time points. The reprojection position is generated by back projection based on the corrected target position vector, pixel anchor point coordinates, pod attitude, body attitude, platform position height, and preset camera intrinsic parameters. The pixel anchor point coordinates and reprojection positions are compared for consistency to generate reprojection deviation information. Continuous deviation information is generated by performing temporal continuity analysis based on the corrected target position vector and the target position vector at consecutive time points; Attitude perturbation analysis is performed on the pod attitude, airframe attitude, and corrected target position vector to generate attitude-sensitive deviation information; Geographic mapping credibility is generated by consistent fusion of reprojection bias information, continuous bias information and attitude-sensitive bias information.

7. The vision-based cross-domain unmanned platform collaborative inspection system according to claim 1, characterized in that, The target's geographical location and candidate target status are correlated and extrapolated to generate target coupling consistency and target prediction status, including: Generate a reference location by performing local coordinate mapping on the target geographic location; A local search domain is generated based on the reference position, and the unmanned surface vessel is controlled to collect the position, heading and speed of candidate targets within the local search domain to generate a candidate target state set; The target's geographical location and its previous state are used to perform motion completion to generate a reference target state; Spatial gating and motion gating are applied to the reference target state and the candidate target state set to generate an associative candidate set; The associated candidate set and the associated target state of the previous time step are temporally associated to generate the associated target state and the target coupling consistency degree; Target prediction states are generated by extrapolating the trajectory based on the associated target states.

8. The vision-based cross-domain unmanned platform collaborative inspection system according to claim 1, characterized in that, The cross-domain collaboration consistency index is generated based on visual lock stability, geographic mapping credibility, and target coupling consistency, including: Consistent center information is generated by centrally aggregating visual lock stability, geographic mapping credibility, and target coupling consistency. Discrete information for consistency is generated by performing discrete analysis on visual lock stability, geographic mapping credibility, and target coupling consistency with consistency center information. A cross-domain collaborative consistency index is generated by consistency fusion based on visual lock stability, geographic mapping credibility, target coupling consistency, and consistency discrete information.

9. The vision-based cross-domain unmanned platform collaborative inspection system according to claim 8, characterized in that, The process of generating coordinated paths and interception commands for sector-loop interception based on the target's predicted status, mission sea area boundaries, restricted area boundaries, and air-sea communication link status includes: Based on the target prediction status, the mission sea area boundary, the restricted area boundary, and the air and sea communication link status, constraint analysis is performed to generate the planning feasible domain and constraint cost field. The target prediction state and the cross-domain coordination consistency index are collaboratively mapped to generate a fan-ring interception domain. Interception anchor points are generated by evaluating the acceptance point based on the sector-ring interception domain and the constraint cost field. A cooperative path is generated by solving the path between the current position of the unmanned vessel and the interception anchor point. Interception commands are generated by encapsulating patterns based on cross-domain collaboration consistency index, sector-ring interception domain, interception anchor point, target prediction state, and collaborative path.

10. The vision-based cross-domain unmanned platform collaborative inspection system according to claim 1, characterized in that, Based on the cooperative path and interception commands, the control of UAVs and unmanned surface vessels to perform cooperative interception and generate interception status includes: The collaborative path, interception command, current state of the UAV and current state of the unmanned vessel are time-calibrated to generate an execution state group; The target prediction state, interception anchor point, and sector interception domain in the interception command are analyzed. Based on the target prediction state and interception anchor point, the UAV execution reference is generated. The current state of the UAV and the UAV execution reference are matched for deviation to generate UAV control commands. Based on the cooperative path and interception anchor point, an unmanned vessel execution reference is generated. The current state of the unmanned vessel and the unmanned vessel execution reference are matched for deviation to generate unmanned vessel control commands. The target prediction status, sector interception domain, UAV control commands and UAV control commands are collaboratively corrected to generate a collaborative interception control quantity, which is then sent to the UAV actuator and the UAV actuator. Based on the execution reference of the UAV, the execution reference of the UAV, the cooperative interception control quantity and the sector interception domain, the interception state is generated and sent to the upper-level task state machine.