Method for identifying mother and baby Yangtze finless porpoises based on aerial image
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF AQUATIC LIFE ACAD SINICA
- Filing Date
- 2026-07-08
- Publication Date
- 2026-08-04
AI Technical Summary
在这种情况下,现有判别方式容易反复出现这样一种现象:两头江豚只是短时并游、交叉同行或者同向尾随,系统也会因为它们在画面中距离较近、方向相近、伴随时间较长而将其计作母幼,而真正处于贴附状态的幼豚又往往不能在每一时刻以完整独立个体呈现出来,结果导致普通伴游样本混入母幼样本,根源在于现有判断依据主要落在接近和伴随这些外部表现上,未能进一步辨别疑似幼豚的显现位置、显现先后及脱离时刻是否受母豚带动;
本方案通过对幼体候选段的出现时刻、保留时刻、体侧归属和跨轴状态进行联合判定,将母幼识别依据由接近伴随转为贴附约束关系判定,相对降低短时并游、交叉同行和同向尾随混入母幼样本的情况;
Smart Images

Figure CN122510941A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image recognition and aerial monitoring technology, and more specifically, to a method for identifying mother and calf Yangtze finless porpoises based on aerial images. Background Technology
[0002] In aerial monitoring work related to the identification of mothers and calves of Yangtze finless porpoises, existing technologies mainly focus on how to screen out targets that can reflect reproductive activities based on water surface images obtained by drones. In actual processing, individual porpoises are first extracted from continuous aerial images, and then the status of mothers and calves is judged by combining individual size differences, distance between them, direction of movement and duration of accompaniment. For example, when conducting routine patrols in the open waters of the middle and lower reaches of the Yangtze River, the flight process must maintain low interference with the activities of finless porpoises, while completing continuous observation over a large area within the limited altitude and patrol distance. At the same time, it is also affected by factors such as water surface reflection, obvious interference from wave patterns and wakes, the small size of juvenile porpoises and their brief exposure when attached to their mothers, and the short period of clear imaging. In this situation, the existing discrimination method is prone to the following phenomenon: two finless porpoises may only swim side by side, cross each other, or follow each other in the same direction for a short time. The system will also count them as mother and calf because they are close in the picture, have similar directions, and accompany each other for a long time. However, the calf that is actually attached to the mother often cannot be presented as a complete and independent individual at every moment. As a result, ordinary companion samples are mixed with mother and calf samples. The root cause is that the existing judgment criteria mainly focus on external manifestations such as proximity and accompaniment, and fail to further distinguish whether the appearance position, appearance order, and separation time of the suspected calf are driven by the mother. The technical problem this application aims to solve is: how to distinguish between the mother-and-feather attachment state of Yangtze finless porpoises and the short-term companionship state of non-mother-and-feather porpoises in aerial images. Summary of the Invention
[0003] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a method for identifying Yangtze finless porpoise mothers and calves based on aerial images. This method involves layering and extracting the process of the porpoise emerging from the water, the continuous manifestation of the mother body, and the attachment and manifestation process on both sides of the boundary in the aerial video. It also combines the appearance order, retention order, body continuity, and cross-axis status of the calves relative to the mother body segments to perform joint determination, thereby solving the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a method for identifying mother and calf Yangtze finless porpoises based on aerial imagery, comprising: S1. Acquire aerial video of the target water area, read each video frame in time sequence, perform water surface background alignment, difference region extraction and finless porpoise image recognition, group the finless porpoise out-of-water areas that appear continuously and are adjacent to each other into display segments, and output the display segment set. S2. Extract the front end position, rear end position and middle bulge position for each display segment. Determine the body axis based on the line connecting the front end position and the rear end position. Determine the body side based on the extension direction of the middle bulge position relative to the body axis. Connect the adjacent display segments with continuous body axis directions to form the parent segment and output the parent segment set. S3. For each parent segment, read the attached display area on both sides of the body axis that is connected to the parent segment frame by frame, perform attachment target image recognition on the attached display area, form attached records according to the first appearance time, the last retention time and the body side where it is located, and merge the attached records that are located on the same body side for multiple consecutive frames and are always attached to the boundary of the parent segment into juvenile candidate segments. S4. For each candidate segment of the juvenile body, compare the time of its first appearance with the time of the complete appearance of the corresponding parent segment, compare the time of its last retention with the time of its disappearance of the corresponding parent segment, and check the body side attribution in each video frame. If the candidate segment of the juvenile body appears later than the complete appearance of the parent segment, disappears earlier than the complete disappearance of the parent segment, and does not cross the body axis, the corresponding parent segment is determined as the parent-juvenile segment. S5. Merge adjacent mother-fetal segments that are consecutive to form a mother-fetal segment and generate the Yangtze finless porpoise mother-fetal identification result.
[0005] In a preferred embodiment, S1 includes: S1-1. Perform static water surface texture registration on adjacent video frames read in time sequence, solve the common water surface corresponding area between adjacent video frames, remove the edge areas that do not enter the common water surface corresponding area, and output the background aligned frame group. S1-2. Perform corresponding pixel difference calculation and connectivity merging on the background aligned frame group. Determine the difference regions in adjacent video frames that are continuously offset in position and have continuously connected ranges as candidate water exit areas. Perform finless porpoise image recognition on the candidate water exit areas and output the finless porpoise water exit areas. S1-3. Perform cross-frame continuation calculations on the order of appearance of the finless porpoise emergence areas, and merge the finless porpoise emergence areas with intersecting boundaries and continuous movement directions in adjacent video frames into the same display segment, and output the display segment set.
[0006] In a preferred embodiment, S2 includes: S2-1. For each display segment, extract the boundary point set of the display segment in the order of video frames, determine the end point group at the beginning of the time series and the end point group at the end of the time series in the boundary point set, and record the midpoint of the end point group as the front position and the back position respectively. Then, record the boundary point group where the distance between the front position and the back position is not zero as the middle raised position, and output the position record group. S2-2. Based on each position recording group, connect the front end position and the back end position to form a body axis, and use the body axis as the boundary to allocate the central bulge position to both sides of the body axis. Calculate the order of appearance of the side where the central bulge position is located in each video frame, and determine the side with the same appearance order as the body side. Output the body axis and body side.
[0007] In a preferred embodiment, S2 further includes: S2-3. For adjacent display segments in time, connect the rear end of the previous display segment with the front end of the next display segment to form a connecting line segment. Determine the turning direction of the connecting line segment with the body axis of the previous display segment and the body axis of the next display segment respectively. When the two turning directions are consistent and the body sides of the two display segments are consistent, the corresponding display segments are recorded as a connecting segment pair, and the connecting segment pair set is output. S2-4. Based on the set of continuation segments, perform concatenation in chronological order. Merge continuation segments in the previous continuation segment pair where the next displayed segment in the previous continuation segment pair is the same displayed segment as the previous displayed segment in the next continuation segment pair into the same connection chain. Merge the displayed segments in the same connection chain into the parent segment and output the parent segment set.
[0008] In a preferred embodiment, S3 includes: S3-1. Read the accompanying display areas on both sides of the boundary of each parent segment frame by frame, with the body axis as the boundary. Extract the boundary bonding position sequence, body axis projection interval and frame order position for each accompanying display area. Record the accompanying display areas with continuous boundary bonding position sequences, body axis projection intervals that do not cross the body axis and frame order positions that are consistent with the corresponding video frames of the parent segment as bonding units, and output the bonding unit set. S3-2. Establish a pairing relationship between the attaching unit sets according to the same parent segment, the same body side and adjacent frame sequence positions. Perform a continuation solution on the last attaching position of the previous attaching unit and the first attaching position of the next attaching unit. Establish a continuation edge when the two are located on the same side boundary of the parent segment and the corresponding body axis projection intervals overlap. Then delete the pairing relationship that does not satisfy the same side continuation or overlapping continuation and output the attaching continuation diagram.
[0009] In a preferred embodiment, S3 further includes: S3-3. Perform chain expansion using the attachment units in the attachment continuation diagram as nodes and the continuation edges as directed edges. Count the first frame position, the last frame position, the number of boundary attachment interruptions, and the number of body side rewrites for each expanded chain. Keep the expanded chains with zero boundary attachment interruptions and zero body side rewrites as accompanying record chains. At the same time, record the first frame position as the first occurrence time, the last frame position as the last retention time, and the body side where it is located as the record body side. Output the accompanying record set. S3-4. Establish record succession relationships for the accompanying record set according to the same parent segment and the same record body side. Perform frame sequence continuity check on the last retention time of the previous accompanying record and the first appearance time of the next accompanying record. Perform boundary coincidence check on the tail attachment position of the previous accompanying record and the first attachment position of the next accompanying record. When the frame sequence continuity check and the boundary coincidence check are successful, merge the corresponding accompanying records into the same young candidate segment and output the young candidate segment set.
[0010] In a preferred embodiment, S4 includes: S4-1. Read the front end position, back end position, and middle bulge position of each video frame according to the frame order position of the mother segment, and construct the front end sequence, back end sequence, and bulge sequence respectively. Perform linear interpolation on the front end sequence and back end sequence to fill in the missing frame positions. Perform three-point difference on the filled front end sequence and back end sequence to solve the volume axis change sequence. Then, record the first frame where the front end sequence, back end sequence, and bulge sequence all exist simultaneously as the complete appearance time, and the next frame after the last frame where they all exist simultaneously as the disappearance time. Record the first frame of the young body candidate segment as the first appearance time and the last frame as the last retention time. Output the time sequence record group. S4-2. Based on the time sequence record group, calculate the entry difference between the first appearance time and the complete appearance time, and the departure difference between the last retention time and the disappearance time, respectively. Construct the entry difference vector and the departure difference vector. Perform symbol decomposition on the entry difference vector and the departure difference vector. Retain the young candidate segments where all components of the entry difference vector are positive and all components of the departure difference vector are negative. Perform correlation backreading on the adjacent frames of the first frame and the last frame. When the correlation backreading result is consistent with the symbol decomposition result, output the time sequence establishment group.
[0011] In a preferred embodiment, S4 further includes: S4-3. Read the boundary bonding positions and body axis projection intervals of the juvenile candidate segments frame by frame according to the time sequence, and construct the body side sequence, bonding sequence, and cross-axis sequence. Perform sequence alignment on the body side sequence and record the positions where the body side codes of adjacent frames are inconsistent as overwrite bits. Perform interval intersection operation on the bonding sequence and record the positions where the bonding intervals of adjacent frames intersect empty as disconnection bits. Perform projection sign determination on the cross-axis sequence and record the positions where the body axis projection interval falls on both sides of the body axis as cross-axis positions. Then construct a conflict matrix composed of overwrite bits, disconnection bits, and cross-axis positions. Perform Gram product and eigenvalue decomposition on the conflict matrix and record the frame positions where the eigenvalue is not zero as conflict frames. Delete the juvenile candidate segments containing conflict frames and output the side-based group.
[0012] In a preferred embodiment, S4 further includes: S4-4. Using each candidate segment of the juvenile in the lateral group as a node, construct a decision graph with the entry difference vector, departure difference vector, body side sequence, attachment sequence, and cross-axis sequence as edge attributes. Perform adjacency matrix expansion and dynamic programming pathfinding on the decision graph. Accumulate the number of positive entry components, negative departure components, consecutive frames on the same side, and consecutive frames attached in each path. Delete paths containing cross-axis positions or conflicting frames. Then, sort the path in lexicographical order according to the order of the number of positive entry components, negative departure components, consecutive frames on the same side, and consecutive frames attached. Determine the parent segment corresponding to the first sorted path as the parent juvenile segment and output the set of parent juvenile segments.
[0013] In a preferred embodiment, S5 includes: S5-1. Construct a continuation relationship table according to the parent segment identifier, first frame position and last frame position of each parent segment. Perform parent segment continuation check and frame sequence continuation check on parent segments that are adjacent in time. When the last frame position of the previous parent segment and the first frame position of the next parent segment are adjacent in time and the corresponding parent segments belong to the same continuation chain, establish inter-segment connection and output parent-child connection diagram. S5-2. Using the mother-child segments in the mother-child connection graph as nodes and the inter-segment connections as directed edges, perform connection chain expansion. Record the first frame position of the first mother-child segment in the same connection chain as the fragment start point and the last frame position of the last mother-child segment as the fragment end point. Write the mother segment identifiers in the same connection chain into the fragment record in order and output the mother-child fragment set. S5-3. Generate the Yangtze finless porpoise mother-feather identification results according to the mother-feather segment set. Write the target water area identifier, segment start point, segment end point, mother segment identifier order and the number of mother-feather segments corresponding to each mother-feather segment into the identification result table, and output the Yangtze finless porpoise mother-feather identification results.
[0014] The technical effects and advantages of this invention are as follows: This scheme combines the occurrence time, retention time, body side attribution, and cross-axis status of candidate segments of offspring to determine the basis for mother-offspring identification from proximity and accompaniment to attachment constraint relationship, thereby relatively reducing the situation of short-term parallel swimming, cross-tracing, and same-direction tailing mixed into mother-offspring samples. By first aligning the water surface with the background of the aerial video, then extracting the difference areas and identifying the finless porpoise's emergence area, the interference of drone displacement, water surface reflection, and wave wakes on target extraction can be relatively reduced, which is conducive to the stable formation of subsequent display segment inputs. Extracting the front end position, rear end position, and middle bulge position from the display segment, further solving the body axis and body side, and connecting the front and rear continuous display segments into the mother segment can provide a continuous reference for reading the accompanying display area, thereby relatively improving the positioning consistency of the mother process; The accompanying display areas are extracted on both sides of the boundary of the mother segment, and the attachment units and larval candidate segments are formed according to the boundary attachment position sequence, body axis projection interval and frame sequence position. This can preserve attachment clues when the larvae are not fully and independently displayed, and relatively improve the recognition availability in local exposure scenarios. By performing continuous checks on the lateral sequence, splice sequence, and cross-axis sequence, and combining conflict matrix, path expansion, and sorting retention rules to screen out abnormal candidates, the system can constrain lateral rewriting, splicing, and cross-axis manifestation, thereby relatively improving the stability and consistency of maternal-fetal determination results. Attached Figure Description
[0015] Figure 1 This is a flowchart outlining the method steps of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Refer to the instruction manual appendix Figure 1 The method for identifying mother and calf Yangtze finless porpoises based on aerial imagery of the present invention includes: S1. Acquire aerial video of the target water area, read each video frame in time sequence, perform water surface background alignment, difference region extraction and finless porpoise image recognition, group the finless porpoise out-of-water areas that appear continuously and are adjacent to each other into display segments, and output the display segment set. In this embodiment, the purpose of S1 is to stably extract the finless porpoise display segments from the aerial video of the target water area for subsequent parent segment calculation. Its mechanism involves first establishing a correspondence between adjacent video frames consisting only of static water surface textures to eliminate interference from slight drone displacement and overall water surface perspective changes on difference detection. Then, it extracts out-of-water change regions with continuous spatial continuity from the background alignment results, and filters out target regions that meet the characteristics of finless porpoise images. Finally, these are grouped into display segments according to temporal sequence and spatial continuity. Specifically, the background alignment frame group output by S1 is used for candidate out-of-water area calculation, the finless porpoise out-of-water area output by S1 is used for cross-frame continuity calculation, and the display segment set output by S1 is used for extracting the front end, rear end, and central bulge positions in S2. This implementation process includes the following steps: The purpose of S1-1 is to preserve common and comparable water surface areas between adjacent video frames and eliminate false differences caused by edge drift. Its mechanism is to first solve the common water surface corresponding area between two frames through static water surface texture registration, and then retain effective pixels only in the common water surface corresponding area. The input is the previous video frame and the next video frame read in time sequence. The processing is as follows: first, extract the static water surface texture points with continuous gray-level changes and texture repetition in the local neighborhood in the two frames, remove the texture points located in the shoreline, hull, floating objects and identified motion areas, and then establish the texture correspondence between the previous video frame and the next video frame according to the gray-level arrangement of the texture points in the neighborhood. Use the texture correspondence to solve the coordinate transformation relationship between the frames, map the previous video frame to the coordinate system of the next video frame, find the overlapping area where there are effective pixels in both frames after mapping as the common water surface corresponding area, and set all edge areas that do not enter the common water surface corresponding area as blank areas to form a background aligned frame group. The output is the background aligned frame group and the corresponding boundary of the common water surface area. The boundary of the common water surface area is written into the current frame pair record for S1-2 to read. The abnormal or missing handling is as follows: when the number of available static water surface texture points in a certain video frame is insufficient to establish a texture pair, the coordinate transformation relationship solved in the previous frame pair is directly read as the alternative transformation relationship of the current frame pair. If the alternative transformation relationship cannot be solved in two consecutive adjacent video frames, the frame pair is marked as an invalid frame pair and writing to the background aligned frame group is stopped. The purpose of S1-2 is to extract real water-emergence change areas from the background-aligned frame group and filter out the finless porpoise water-emergence areas. Its working mechanism is to first perform difference calculation on the corresponding pixels in the corresponding area of the common water surface to locate the change area, and then form candidate water-emergence areas through connectivity merging. Subsequently, non-finless porpoise change areas are eliminated based on the finless porpoise image recognition field. The input is the background-aligned frame group and the boundary of the corresponding area of the common water surface. The processing action is: only in the corresponding area of the common water surface, perform the corresponding pixel difference calculation on the previous background-aligned frame and the next background-aligned frame, merge the pixels with non-zero difference and connected in the eight-code neighboring area into difference areas, extract the center position of the region, the area surrounding the boundary, the length and width arrangement of the region, the continuous state of the back arc, the separation state of the tail, and the displacement direction of the region along the time sequence for each difference area, and merge the difference areas in the previous video frame and the next video frame with continuous offset center position and continuous connection of the area surrounding the boundary into candidate water-emergence areas. Subsequently, image recognition of finless porpoises was performed on each candidate water exit area. The image recognition of finless porpoises used the outline arc shape, length and width arrangement, continuous back ridge shape, and independent water exit shape after separation from the tail as recognition fields. Candidate water exit areas that simultaneously meet the requirements of continuous back ridge, closed main body boundary, and tail area not merging into the main body boundary were identified as finless porpoise water exit areas. The output was the finless porpoise water exit area and its corresponding center position, boundary enclosing area, displacement direction, and frame sequence position. These fields were written into the water exit area record table for S1-3 to read. The abnormal or missing handling was as follows: when there was a one-to-many split in the difference area in adjacent video frames, the next difference area with a continuous boundary connection range with the previous difference area was retained, and the remaining difference areas were deleted. When there was a many-to-one merging, the merging result with the line connecting the center positions of the difference areas before merging in the same direction as the displacement direction of the merged area was taken as the candidate water exit area. If the consistency relationship did not exist, the merging result was not written into the water exit area record table. The purpose of S1-3 is to connect the finless porpoise emergence areas that are discretely distributed in adjacent video frames into the same display process. Its working mechanism is to establish cross-frame continuity based on frame order, boundary intersection, and continuous movement direction, and to merge the finless porpoise emergence areas belonging to the same continuity chain into display segments. The input is the finless porpoise emergence areas arranged in the order of appearance and the emergence area record table. The processing action is to establish cross-frame candidate pairs for the finless porpoise emergence areas in the adjacent video frames one by one. First, it is determined whether the boundary area of the previous finless porpoise emergence area overlaps with the boundary area of the next finless porpoise emergence area. Then, it is determined whether the displacement direction of the previous finless porpoise emergence area is the same as the displacement direction of the next finless porpoise emergence area. The cross-frame candidate pairs that simultaneously satisfy boundary intersection and continuous movement direction are established as continuity edges. Subsequently, an exit continuity graph is constructed using the finless porpoise exit area as a node and the connecting edge as a directed edge. The continuity chain is expanded in ascending order of frame sequence. The frame corresponding to the first finless porpoise exit area in the same continuity chain is recorded as the starting frame of the display segment, and the frame corresponding to the last finless porpoise exit area is recorded as the ending frame of the display segment. All finless porpoise exit areas in the same continuity chain are merged into the same display segment. The output is the display segment set and the starting frame, ending frame, and the order of the finless porpoise exit area numbers contained in each display segment. These fields are written into the display segment table for S2 to read. The abnormal or missing handling is as follows: when the same finless porpoise exit area establishes a connecting edge with multiple finless porpoise exit areas in the following video frame, the connecting edge with continuous boundary overlap length and the same change in movement direction as the previous finless porpoise exit area is retained, and the remaining connecting edges are deleted. When a finless porpoise exit area has no connecting edge in the following video frame, the finless porpoise exit area is written into the display segment table as the current display segment endpoint and no further continuity is performed. Through the above implementation process, the original aerial video can be stably converted into a set of display segments for subsequent mother segment processing, avoiding the generation or connection of display segments due to slight displacement of the UAV, overall water surface drift, wake interference, or splitting and merging of different regions. This ensures that the extraction of the front-end position, back-end position, and middle ridge position has a unified input basis. At the same time, this implementation process clearly defines the field composition, generation order, anomaly handling, and unique output of the corresponding area of the common water surface, the candidate water outflow area, the finless porpoise water outflow area, and the display segment, thereby eliminating the problems of unclear input sources, merging of different regions, and cross-frame continuation in subsequent implementations. In practical applications: For example, during drone patrols in an open water area in the middle and lower reaches of the Yangtze River, the continuously collected aerial videos show slight aircraft deflection and changes in water surface reflection. First, static texture points on the water surface are extracted from adjacent video frames, and the corresponding areas of the common water surface are solved. Then, the difference areas are extracted within the corresponding areas of the common water surface, and the finless porpoise emergence areas with the arc-shaped back and tail separation characteristics are screened out. Finally, the finless porpoise emergence areas with intersecting boundaries and continuous movement directions in adjacent video frames are merged into the same display segment, thereby obtaining a set of display segments that can be directly read for subsequent calculation of body axis, body side, and parent body segments.
[0018] S2. Extract the front end position, rear end position and middle bulge position for each display segment. Determine the body axis based on the line connecting the front end position and the rear end position. Determine the body side based on the extension direction of the middle bulge position relative to the body axis. Connect the adjacent display segments with continuous body axis directions to form the parent segment and output the parent segment set. In this embodiment, the purpose of S2 is to extract the body axis and body sides that can characterize a single continuous manifestation process of the finless porpoise based on the manifestation segment set output by S1, and further connect the preceding and following manifestation segments belonging to the same continuous manifestation process into a parent segment. Its working mechanism is to first extract stable end and ridge information from each manifestation segment to form the front end position, rear end position, and middle ridge position; then, use the front end position and rear end position to determine the body axis, and use the stable side of the middle ridge position relative to the body axis to determine the body sides; subsequently, perform continuity calculations on temporally adjacent manifestation segments, and finally connect manifestation segments that meet the continuity conditions into the same link and merge them into a parent segment; wherein, the position record group output by S2 is used for body axis and body side calculation reading, the body axis and body sides output by S2 are used for manifestation segment continuity judgment reading, and the parent segment set output by S2 is used for reading the accompanying manifestation area in S3; this implementation process includes the following steps: The purpose of S2-1 is to extract stable position fields that can be used for body axis and body side calculations from each display segment. Its working mechanism is to determine the end point groups at both ends through the boundary point set of the display segment, and extract the middle bulge position after removing the end influence. The input is the display segment set and the boundary record of the finless porpoise's water exit area corresponding to each display segment. The processing action is as follows: read the boundary point set of each display segment in each video frame in the order of video frames, first perform sequential arrangement of the boundary points along the extension direction of the display segment boundary, then classify the consecutive adjacent boundary points at the first side of the arrangement sequence into the first side end point group, and classify the consecutive adjacent boundary points at the last side of the arrangement sequence into the last side end point group. The number of points contained in the end point group is given by the preset configuration, which is determined based on the aerial video resolution and the average number of boundary points in the finless porpoise's water exit area per frame. Subsequently, the average coordinates of the first and last end point groups are calculated separately. The two average coordinates are recorded as the front end position and the back end position, respectively. The front end position and the back end position are then connected to form a front-back line. Boundary points that are not located at the front-back position are projected onto the front-back line. Boundary points whose projection positions are located within the middle section of the front-back line are taken as middle boundary points. The vertical distance from the middle boundary points to the front-back line is calculated point by point. The middle boundary points that are always located on the same side and appear consecutively in the frame order in multiple consecutive frames are merged into the middle bulge point group. The average coordinates of the middle bulge point group are recorded as the middle bulge position. The output is a position record group, which includes at least the display segment identifier, video frame number, front end position, back end position, and central bulge position. The position record group is written into the position record table for S2-2 to read. The abnormal or missing handling is as follows: when the first end point group or the last end point group is missing in a video frame, the front end position or back end position is not written in that video frame; when there are multiple central bulge point groups in a video frame, only the central bulge point groups that have corresponding points on the same side in adjacent video frames are retained, and the other central bulge point groups are deleted. The purpose of S2-2 is to solve the body axis and body sides corresponding to each display segment based on the position recording group. Its working mechanism is to form the body axis using the front and rear positions, and then determine the body sides based on the lateral maintenance relationship of the central bulge position in consecutive video frames. The input is the position recording group. The processing action is as follows: read the front and rear positions corresponding to the same display segment in the position recording group one by one, connect the front and rear positions to form the body axis, and use the body axis as the dividing line to allocate the central bulge position to both sides of the body axis. Then, generate the lateral code sequence of the central bulge position in the order of video frame number, where the left side of the body axis is recorded as the first lateral code, the right side of the body axis is recorded as the second lateral code, and the corresponding video frame that has not written the central bulge position is recorded as an empty code. Next, perform continuous statistics on the side code sequence. The side codes that appear frequently and remain unchanged from the first to the last frame of the display segment are recorded as body sides. If two side codes form a continuous sequence in the same display segment, only the side code corresponding to the middle video frame of the display segment is retained as the body side. The output is the body axis and body side corresponding to each display segment. The display segment identifier, video frame number, body axis endpoint coordinates, and body side are written into the body axis and body side table for S2-3 to read. The abnormal or missing handling is as follows: when a display segment does not form a central bulge position in all video frames, the display segment will not enter the continuation judgment. When an empty code appears in the first or last frame of the side code sequence, the empty code will not participate in the body side determination. When the empty code is located in the middle video frame, the empty code is filled with a side code that is consistent before and after the empty code. If the side codes before and after the empty code are inconsistent, the video frame is recorded as a side code interruption frame and will not be used as the basis for body side retention in subsequent continuation judgment. The purpose of S2-3 is to filter out consecutive segments with a continuous display relationship from adjacent display segments in time. Its working mechanism is to use the turning relationship between the connecting line segment between the display segments and the front and rear body axes, as well as whether the front and rear body sides are consistent, to determine whether two display segments belong to the same continuous display process. The input quantities are the display segment set, the body axis and body side table, and the position record table. The processing action is as follows: the display segment set is sorted by time according to the start frame and end frame of the display segment. For the preceding and following display segments that are adjacent in time after sorting, the rear end position corresponding to the last frame of the preceding display segment and the front end position corresponding to the first frame of the following display segment are read, and the two positions are connected to form a connecting line segment. Next, calculate the turning sign from the body axis of the previous display segment to the connecting line segment and the turning sign from the connecting line segment to the body axis of the next display segment. One turning is from the direction of the body axis from the start to the end point to the direction of the connecting line segment, and another turning is from the direction of the connecting line segment to the direction of the body axis of the next display segment. Clockwise is recorded as the first turning sign, and counterclockwise as the second turning sign. When the two turning signs are consistent and the body sides of the previous and next display segments are consistent, the preceding and following display segments are recorded as a pair of connecting segments. The output is a set of connecting segment pairs, which includes at least the preceding... The system identifies the first display segment, the next display segment, the connecting segment, the preceding turning symbol, the following turning symbol, and the body side consistency mark. It also writes the set of connecting segment pairs into the connecting relationship table for S2-4 to read. The abnormal or missing handling is as follows: when the first frame of the next display segment is missing the front end position or the last frame of the previous display segment is missing the back end position, no connecting segment is established. When the same previous display segment can form a connecting segment pair with multiple subsequent display segments, only the connecting segment pair with body side consistency and whose two turning symbols are consistent with the previous display segment's preceding connecting direction is retained, and the rest of the connecting segment pairs are deleted. The purpose of S2-4 is to connect the sets of consecutive segments into a link and merge them into a parent segment. Its working mechanism is to establish a link based on the shared relationship between the first and last visible segments of the consecutive segments, and then merge the visible segments in the same link into the same parent segment. The input is the set of consecutive segments and the set of visible segments. The processing is as follows: First, read the set of consecutive segments in chronological order, establish a link between the next visible segment in the previous consecutive segment pair and the previous visible segment in the next consecutive segment pair that is the same visible segment, and then construct a visible segment link graph with the consecutive segments as edges and visible segments as nodes. Then, perform chain expansion according to the visible segment link graph, mark the visible segments that do not have a predecessor consecutive segment pair as the starting visible segment of the link, and write the visible segments that can be continuously reached along the time increment direction into the same link in sequence until there are no subsequent consecutive segments, and stop the expansion, and merge all visible segments in the same link into the same parent segment. The output is a set of parent segments, which includes at least a parent segment identifier, the order of the display segments identifier, the starting display segment of the connection chain, the ending display segment of the connection chain, and the body side. The parent segment set is written into the parent segment table for S3 to read. The handling of anomalies or missing segments is as follows: when the same display segment enters two connection chains at the same time, only the connection chain that was formed first in time sequence is retained, and the other connection chain deletes the display segment and continues to expand. When two connection chains re-merge into the same display segment in subsequent video frames, the connection chain with more display segments before the same display segment is retained. If the number of display segments is the same, the connection chain with the smaller starting frame number of the starting display segment is retained. Through the above implementation process, the front end position, rear end position, and middle ridge position can be stably solved from the display segment set output by S1, further forming the body axis, body side, and parent body segment set. This provides a unified, continuous, and uniquely assigned input basis for subsequent reading of the accompanying display area and solving of candidate segments of the young body. At the same time, this implementation process clearly defines the value selection method, generation order, anomaly handling, and unique retention rules for end point groups, middle ridge positions, body side determination, turning symbols, connecting segment pairs, and connecting chains, thereby eliminating problems such as unclear body axis, body side drift, incorrect connection of display segments, and multiple assignments of parent body segments. In practical applications: For example, in continuous aerial video of the target water area, the same process of a finless porpoise emerging from the water is divided into multiple temporally adjacent display segments by S1. First, the boundary point set is extracted frame by frame from each display segment to generate the front position, rear position and central bulge position. Then, the body side is solved by the side where the central bulge position appears continuously. Subsequently, the turning signs of the connecting line segments and the front and rear body axes are calculated for temporally adjacent display segments, and the connecting segment pairs with the same turning signs and the same body side are screened out. Finally, the connecting segment pairs sharing the display segment are expanded into the same connecting chain and merged into the parent segment, thereby obtaining the parent segment set that can be read frame by frame by S3 for the accompanying display area.
[0019] S3. For each parent segment, read the attached display area on both sides of the body axis that is connected to the parent segment frame by frame, perform attachment target image recognition on the attached display area, form attached records according to the first appearance time, the last retention time and the body side where it is located, and merge the attached records that are located on the same body side for multiple consecutive frames and are always attached to the boundary of the parent segment into juvenile candidate segments. In this embodiment, the purpose of S3 is to identify the accompanying display process that occurs when the mother segment adheres to the mother segment on both sides of the boundary, and to group the discrete attachment display information into candidate segments of the offspring that can be read for subsequent mother-offspring determination. Its working mechanism is to first read the accompanying display area frame by frame based on the body axis and boundary of the mother segment, extract the boundary attachment position sequence, body axis projection interval and frame sequence position to form attachment units that meet the attachment conditions, and then establish the connection relationship between attachment units under the same mother segment, the same body side and adjacent frame sequence positions to form attachment continuity. The diagram is then processed by chain expansion of the attachment continuation diagram, filtering out the accompanying record chains with uninterrupted boundary attachment and no alteration of the body side. Finally, accompanying records that are continuous in time and boundary with the same parent segment and the same recording body side are merged into the same juvenile candidate segment. The attachment unit set output by S3 is used for attachment continuation diagram construction, the accompanying record set output by S3 is used for juvenile candidate segment merging, and the juvenile candidate segment set output by S3 is used for the first occurrence time, last retention time, and body side attribution verification in S4. This implementation process includes the following steps: The purpose of S3-1 is to filter out the accompanying display areas that maintain a close fit with the boundary of the parent segment from the video frames corresponding to each parent segment, and to write the accompanying display areas that meet the fitting conditions into the fitting unit set. Its working mechanism is to limit the reading range with the body axis as the boundary, and then use three sets of conditions—boundary fitting, same-side projection of the body axis, and frame sequence correspondence—to jointly constrain the establishment of the accompanying display areas. The input quantities are the parent segment table, the body axis and body side table, the display segment boundary record, and the finless porpoise emergence area boundary record in each video frame. The processing action is to read the body axis, body side, and parent segment in the corresponding video frame frame by frame. For each body segment boundary, search for the outer connected region that directly contacts the boundary of the parent body segment according to the pixel adjacency relationship. Record the outer connected region as the candidate region of the accompanying display area. Extract the boundary contact points that are connected to the boundary of the parent segment for each candidate region of the accompanying display area, and arrange the boundary contact points in the direction of the boundary of the parent segment to form a boundary contact position sequence. Then, map all the boundary points of the candidate regions of the accompanying display area to the body axis with perpendicular feet, and take the start position and end position of all perpendicular feet on the body axis to form the body axis projection interval. At the same time, record the current video frame number as the frame sequence position. Subsequently, a continuity judgment is performed on the boundary bonding position sequence. The continuity judgment is performed based on whether there is an interruption in the numbering of the boundary bonding points on the boundary of the parent segment. When the number of interruption points is zero, the boundary bonding position sequence is considered continuous. A cross-axis judgment is performed on the volume axis projection interval. The cross-axis judgment is performed based on whether the perpendicular feet of the boundary points of the candidate area of the attached display area are simultaneously distributed on both sides of the volume axis. When the perpendicular feet only fall on one side of the volume axis, the volume axis projection interval is considered not to cross the volume axis. A correspondence judgment is performed on the frame sequence position. The correspondence judgment is performed based on whether the video frame in which the candidate area of the attached display area is located belongs to the frame sequence range of the corresponding parent segment. When the boundary bonding position sequence is continuous, the body axis projection interval does not cross the body axis, and the frame order position is consistent with the corresponding video frame of the parent segment, the candidate area of the accompanying display area is recorded as a bonding unit and written into the bonding unit table; the output is a bonding unit set, which includes at least the parent segment identifier, video frame number, body side, boundary bonding position sequence, body axis projection interval, first bonding position, and last bonding position, and the bonding unit set is written into the bonding unit table for S3-2 to read; the abnormal or missing handling is as follows: when there are multiple candidate areas of the accompanying display area on a certain body side in the same video frame, only the candidate areas of the accompanying display area with a continuous number of bonding points with the parent segment boundary and whose body axis projection interval is located within the middle section of the parent segment are retained, and the remaining candidate areas of the accompanying display area are deleted; when the candidate area of the accompanying display area only contacts the parent segment boundary with a single point, it is not written into the bonding unit table; The purpose of S3-2 is to establish a sequential relationship between attachment units and form an attachment continuity diagram. Its working mechanism is to use the boundary position continuity relationship and body axis projection overlap relationship between attachment units under the same parent segment, the same body side, and adjacent frame sequence position to filter out the continuity edges belonging to the same attachment display process. The input quantities are the attachment unit table and the parent segment table. The processing actions are as follows: sort the attachment unit set according to the parent segment identifier, body side, and video frame number. Under the same parent segment and the same body side, establish a sequential pairing relationship between the attachment units in the previous video frame and the attachment units in the next video frame. For each sequential pairing relationship, read the last attachment position of the previous attachment unit and the first attachment position of the next attachment unit. If the two are the same number or adjacent numbers on the boundary number of the parent segment, they are recorded as continuous first and last attachments. Then read the body axis projection interval of the previous attachment unit and the body axis projection interval of the next attachment unit, and perform interval intersection operation. When the intersection of the two intervals is not empty, it is recorded as body axis projection interval overlap. Subsequently, it is simultaneously checked whether the previous and subsequent attachment units are located on the same side boundary of the parent segment. If the body side of the previous attachment unit is consistent with that of the subsequent attachment unit, and the beginning and end are continuously attached and the body axis projection intervals overlap, then a connection edge is established between the previous and subsequent attachment units, and the corresponding pairing relationship is written into the attachment connection table. If the pairing relationship does not satisfy the same-side connection or does not satisfy the body axis projection interval overlap, then the pairing relationship is deleted, and no connection edge is established. The output is an attachment connection diagram, which includes at least the starting attachment unit number. The attachment unit number, parent segment identifier, body side, first and last attachment continuity mark, and body axis projection interval overlap mark are terminated, and the attachment continuity diagram is written into the continuity diagram for S3-3 to read; the abnormal or missing handling is as follows: when the same previous attachment unit can establish a continuity edge with multiple attachment units in the next video frame, only the continuity edge with a small difference between the first attachment position and the last attachment position number of the previous attachment unit and a long body axis projection interval intersection length is retained, and the rest of the continuity edges are deleted; when there is no attachment unit on the same body side in the next video frame, the previous attachment unit will no longer establish a continuity edge to the next frame; The purpose of S3-3 is to extract stably consistent trailing record chains from the attachment continuity graph and form a trailing record set. Its mechanism involves expanding the attachment unit chain using attachment units as nodes and connecting edges as directed edges, then filtering out unstable expanded chains using the number of boundary attachment interruptions and body-side rewrites. The inputs are the continuity graph, the attachment unit table, and the parent segment table. The processing steps are: constructing a directed graph using attachment units as nodes and connecting edges as directed edges; searching for attachment units without preceding connecting edges in ascending video frame order as the chain start node; and performing chain expansion along the directed edge direction. The process continues until the current attached unit has no subsequent connecting edge, forming an unfolded chain. For each unfolded chain, the first frame position, the last frame position, the number of boundary attachment interruptions, and the number of body-side rewrites are counted. The first frame position is the frame sequence position corresponding to the chain start node, and the last frame position is the frame sequence position corresponding to the chain end node. The number of boundary attachment interruptions is counted based on whether there are any gaps between adjacent attached units where no connecting edge has been established, and the number of gaps is recorded as the number of boundary attachment interruptions. The number of body-side rewrites is counted based on whether there are any side-specific changes in the body-side code sequence of all attached units in the unfolded chain, and the number of side-specific changes is recorded as the number of body-side rewrites. When the number of boundary bonding interruptions is zero and the number of body-side rewrites is zero, the unfolded chain is retained as an accompanying record chain. At the same time, the first frame position is recorded as the first occurrence time, the last frame position is recorded as the last retention time, and the body side where the unfolded chain is located is recorded as the recording body side. The output is an accompanying record set, which includes at least the parent segment identifier, the accompanying record chain number, the first occurrence time, the last retention time, the recording body side, the first bonding position, and the last bonding position. The accompanying record set is written into the accompanying record table for S3-4 to read. The abnormal or missing handling is as follows: when the same bonding unit can enter multiple unfolded chains at the same time, only the unfolded chain with consecutive frame positions and a large number of nodes is retained. When the number of nodes is the same, the unfolded chain with the earlier first frame position is retained. When the body side code of a certain video frame in the unfolded chain is empty, the empty code is filled back with the body side code that is consistent with the previous and next video frames. If the previous and next body side codes are inconsistent, the number of body-side rewrites of the unfolded chain is incremented by one. The purpose of S3-4 is to group accompanying records belonging to the same parent segment and maintaining the same record body side into candidate segments for young bodies. Its mechanism involves determining whether adjacent accompanying records can be continuously displayed as the same attachment target through frame sequence continuity checks and boundary coincidence checks. The inputs are the accompanying record table, the attachment unit table, and the parent segment table. The processing actions are as follows: The accompanying record set is grouped according to the parent segment identifier and record body side; within the same parent segment and the same record body side, a record-interval relationship is established between the preceding and following accompanying records; for each record-interval relationship, the last retention time of the preceding accompanying record and the first appearance time of the following accompanying record are read, and frame sequence continuation is performed. Continued verification: Frame sequence continuous verification is performed based on whether the first occurrence time of the subsequent accompanying record is the first or second frame after the last retention time of the preceding accompanying record. The allowed interval frame number is given by the preset configuration, which is determined based on the aerial video frame rate and the statistical results of the number of frames of short-term occlusion of the Yangtze finless porpoise. Then, the tail attachment position of the preceding accompanying record and the first attachment position of the subsequent accompanying record are read, and boundary coincidence verification is performed. Boundary coincidence verification is performed based on whether the difference between the tail attachment position and the first attachment position on the boundary number of the parent segment does not exceed the preset boundary displacement. The preset boundary displacement is given by statistical estimation, and the statistical object is the difference in the attachment positions of adjacent frames of continuously attached units within the same parent segment. When both frame sequence continuity and boundary overlap verification are successful, the corresponding preceding and following records are merged into the same juvenile candidate segment and written into the juvenile candidate segment table in chronological order. The output is a juvenile candidate segment set, which includes at least the juvenile candidate segment number, parent segment identifier, record body side, first appearance time, last retention time, accompanying record chain number order, and boundary attachment range. The juvenile candidate segment set is written into the juvenile candidate segment table for S4 to read. The handling of abnormalities or missing records is as follows: when the same preceding accompanying record can satisfy both frame sequence continuity and boundary overlap verification with multiple following accompanying records, only the following accompanying record with the smaller difference between the first attachment position and the last attachment position of the preceding accompanying record is retained, and the other following accompanying records are not merged. When two juvenile candidate segments re-merge into the same accompanying record in subsequent video frames, only the juvenile candidate segment with the earlier first appearance time is retained, and the other juvenile candidate segment is terminated from being written before the merging frame. Through the above implementation process, the accompanying display area can be stably extracted from both sides of the boundary of the parent segment, and the attachment unit, attachment continuation diagram, accompanying record set and juvenile candidate segment set can be formed in sequence. This organizes the local, short-term, and discrete accompanying display information into continuous objects with clear time range, clear body side attribution, and clear boundary attachment relationship, providing a unified input for the first appearance time, last retention time, and body side attribution verification in the subsequent parent-juvenile segment determination. At the same time, this implementation process clearly defines the reading boundary of the accompanying display area, the composition of the boundary attachment position sequence, the formation of the body axis projection interval, the establishment and deletion of the continuation edge, the retention of the unfolded chain, the verification of the continuation between records, and the unique output of the juvenile candidate segment, thereby eliminating the problems of unclear attachment target boundary, multiple attribution of record chains, and repeated merging of juvenile candidate segments. In practical applications: For example, in consecutive video frames of the same parent segment, multiple small-scale appendage display areas that are connected to the parent segment boundary appear consecutively near the right boundary of the parent segment. First, the boundary attachment position sequence and body axis projection interval of each appendage display area are extracted frame by frame, and the attachment units that are always located on the same side of the body axis are screened out. Then, the attachment units in adjacent video frames are connected to form an attachment continuity diagram. Subsequently, the appendage record chain with uninterrupted boundary attachment and no body side modification is obtained. Finally, the appendage records that are continuous in time and whose boundary attachment positions are connected are merged into the same juvenile candidate segment, thereby obtaining a set of juvenile candidate segments that can be directly read by S4.
[0020] S4. For each candidate segment of the juvenile body, compare the time of its first appearance with the time of the complete appearance of the corresponding parent segment, compare the time of its last retention with the time of its disappearance of the corresponding parent segment, and check the body side attribution in each video frame. If the candidate segment of the juvenile body appears later than the complete appearance of the parent segment, disappears earlier than the complete disappearance of the parent segment, and does not cross the body axis, the corresponding parent segment is determined as the parent-juvenile segment. In this embodiment, the purpose of S4 is to perform a joint determination of temporal constraints, body-side constraints, and adhesion constraints on the candidate segments of the young organism, and to screen out the mother-young organism segments that satisfy the mother-young attachment relationship from the candidate segments of the young organism segments. Its working mechanism is to first solve the complete appearance time, disappearance time, first appearance time, and last retention time based on the temporal position of the mother segment and the candidate segments of the young organism segments, and then use the entry difference and departure difference to screen out the candidate segments of the young organisms that satisfy the mother-young and young organism attachment relationships. Subsequently, the body-side continuity, boundary adhesion continuity, and other constraints are checked frame by frame for the retained candidate segments of the young organisms. If a candidate segment crosses the body axis, delete any candidate segments that exhibit lateral rewriting, attachment interruption, or cross-axis manifestation. Finally, construct a decision graph from the remaining candidate segments and perform path expansion and path preservation to determine the parent segment corresponding to the candidate segment that satisfies all constraints as the parent-child segment. The timing record group output by S4 is used for constructing and reading the entry and exit difference vectors; the timing establishment group output by S4 is used for verifying the body-side sequence and attachment sequence; and the lateral establishment group output by S4 is used for constructing the decision graph and reading the parent-child segment output. This implementation process includes the following steps: The purpose of S4-1 is to unify the timing reference of the parent segment and the candidate segments for the young segment and generate the timing record group required for subsequent judgment. Its working mechanism is to use the completeness of the appearance of the front end position, the back end position and the middle bulge position in the parent segment to determine the complete appearance time and disappearance time of the parent segment, and then use the start and end frame positions of the candidate segments for the young segment to determine the first appearance time and the last retention time. The input quantities are the parent segment table, the position record table and the candidate segment table for the young segment. The processing actions are as follows: read the front end position, the back end position and the middle bulge position in all video frames covered by the parent segment according to the parent segment identifier, and write them into the front end sequence, the back end sequence and the bulge sequence in the order of video frames respectively. Perform linear interpolation on the missing frame positions in the front end sequence and the back end sequence. Linear interpolation is only performed on missing frames located between known front and back frames, and the interpolation position is generated equally according to the coordinate difference of the known front and back positions by the number of frames. No interpolation is performed on missing positions located before the first frame or after the last frame. Then, a three-point difference is performed on the completed front-end sequence and the back-end sequence. The three-point difference calculates the direction of change of the current position based on the position coordinates of the current frame and the adjacent frames before and after it, so as to obtain the body axis change sequence of the mother segment between adjacent frames. Then, the first video frame in which the front-end sequence, the back-end sequence and the bulge sequence exist at the same time is read and recorded as the complete display time. The last video frame in which the three exist at the same time is read and the first video frame after the last video frame is recorded as the disappearance time. When there is no next video frame after the last video frame, the last video frame itself is written into the disappearance flag bit as the disappearance time. Then, the first and last frame positions of the juvenile candidate segments are read according to their numbers, and recorded as the first appearance time and the last retention time, respectively. The output is a timing record group, which includes at least the parent segment identifier, juvenile candidate segment number, complete appearance time, disappearance time, first appearance time, last retention time, and body axis change sequence. The timing record group is written into the timing record table for S4-2 to read. The abnormal or missing handling is as follows: when the bulge sequence is missing in all video frames of the parent segment, the parent segment will not enter the subsequent judgment of S4; when either end of the interpolation interval is missing, the current missing frame will remain null and a missing mark will be written into the timing record table. The purpose of S4-2 is to screen out candidate segments of young organisms that satisfy the sequence of mother-to-young and young-to-retreat from a temporal perspective. Its mechanism involves constructing an entry difference using the complete appearance time and the first appearance time, and a departure difference using the disappearance time and the last retention time. The temporal relationship is then verified through symbolic decomposition and adjacent frame correlation readback. The inputs are a temporal record table, a mother segment table, and a candidate young organism table. The processing involves reading each temporal record group sequentially, calculating the frame sequence difference between the first appearance time and the complete appearance time for each candidate young organism segment, and recording this difference as the entry difference. Difference; calculate the frame sequence difference between the last retained time and the disappearance time, and record it as the departure difference; for all young candidate segments within the same parent segment, the entry difference is formed into an entry difference vector according to the young candidate segment numbering order, and the departure difference is formed into a departure difference vector according to the same order; then perform symbol decomposition on the entry difference vector and the departure difference vector, and the symbol decomposition is recorded as positive if the component is greater than zero, zero if equal to zero, and negative if less than zero. Young candidate segments whose entry difference vector has all positive components and whose departure difference vector has all negative components are regarded as time-passing objects; Then, for each time series, a related readback is performed through the object. The related readback reads the boundary bonding position and the volume axis projection range in the video frame before the first occurrence time, the video frame after the first occurrence time, the video frame before the last retention time, and the video frame after the last retention time. It then determines whether there is no boundary bonding record corresponding to the candidate segment of the juvenile body in the video frame before the first occurrence time, whether there is a continuous bonding record in the video frame after the first occurrence time, whether there is a continuous bonding record in the video frame before the last retention time, and whether there is no continuous bonding record in the video frame after the last retention time. When the relevant readback result is consistent with the sign decomposition result of the positive sign of the entry difference vector and the negative sign of the exit difference vector, the candidate segment of the young body is retained; the output is a timing establishment group, which includes at least the parent segment identifier, the candidate segment number of the young body, the entry difference, the exit difference, the entry difference vector, the exit difference vector, and the relevant readback consistency flag, and the timing establishment group is written into the timing establishment table for S4-3 to read; the abnormal or missing handling is as follows: when the first occurrence time or the last retention time is located in the first frame or the last frame of the parent segment, resulting in the absence of the previous video frame or the next video frame, the corresponding missing adjacent frame is recorded as an empty frame and the relevant readback is performed only with the existing adjacent frames; when the entry difference is zero or the exit difference is zero, the candidate segment of the young body is not written into the timing establishment table; The purpose of S4-3 is to remove young candidate segments from the time-series establishment group that have body-side rewriting, splicing breaks, or cross-axis appearance. Its working mechanism is to construct body-side sequences, splicing sequences, and cross-axis sequences by frame-by-frame boundary splicing positions and body axis projection intervals, and then use a conflict matrix to centrally express the three types of conflicts and locate conflict frames through matrix decomposition. The input quantities are the time-series establishment table, young candidate segment table, splicing unit table, and parent segment table. The processing action is as follows: read the boundary splicing positions and body axis projection intervals of each video frame corresponding to the young candidate segment according to the time-series establishment group, generate a body-side code for each video frame according to whether the boundary splicing position is on the left or right side of the body axis, write the first body-side code on the left side, write the second body-side code on the right side, and write an empty code if missing. The body-side codes of all video frames are used to form a body-side sequence according to the frame order. A splicing sequence is generated based on the boundary splicing position intervals of adjacent video frames. Each component in the splicing sequence records whether there is an intersection between the splicing intervals of two adjacent frames. If the intersection is not empty, it is recorded as a continuous splicing code; if the intersection is empty, it is recorded as a disjoint splicing code. Projection sign determination is performed on the volume axis projection intervals in each video frame. The projection sign determination is performed based on whether the start and end points of the projection interval are located on one side of the volume axis or span both sides of the volume axis. If the start and end points of the projection interval are both on the same side of the volume axis, it is recorded as a non-cross-axis code; if they fall on both sides of the volume axis, it is recorded as a cross-axis code. The determination results of all video frames are used to form a cross-axis sequence in frame order. Then, sequence alignment is performed on the volume side sequence. The positions where the volume side codes of adjacent frames are inconsistent are recorded as overwrite bits. Interval intersection is performed on the splicing sequence. The positions where the intersection of the splicing intervals of adjacent frames is empty are recorded as disjoint bits. The position where the cross-axis code is read from the cross-axis sequence is recorded as the cross-axis bit. A conflict matrix is constructed by arranging video frame positions in rows and by arranging three types of conflicts: overwritten bits, disconnected bits, and cross-axis bits. When a certain type of conflict exists in a corresponding video frame, the first conflict code is written to the corresponding row and column; otherwise, a zero code is written. A Gram product is performed on the conflict matrix to obtain a conflict correlation matrix. Then, eigenvalue decomposition is performed on the conflict correlation matrix, and the positions of video frames with non-zero eigenvalues are recorded as conflict frames. Young candidate segments containing any conflict frame are deleted, and only young candidate segments without conflict frames are retained. The output is a side-success group, where side-success is defined. Each group includes at least the parent segment identifier, the candidate segment number of the young segment, the body side sequence, the bonding sequence, the cross-axis sequence, the conflict matrix, and the conflict frame position. The side-identification established group is written into the side-identification established table for S4-4 to read. The abnormal or missing handling is as follows: when a blank code appears in the middle of the body side sequence and the body side codes of the two frames before and after the blank code are consistent, the blank code is backfilled with the consistent body side codes of the two frames before and after; when the body side codes before and after the blank code are inconsistent, the position of the blank code is directly recorded as the rewrite position; when the bonding interval is missing, causing the interval intersection operation to be unable to be performed, the position is recorded as the disconnection position. The purpose of S4-4 is to determine the final parent segment from candidate segments that satisfy temporal and lateral constraints. Its mechanism involves constructing a decision graph using candidate segments as nodes and temporal and lateral attributes as edge attributes. Then, through adjacency matrix expansion, dynamic programming pathfinding, and lexicographical sorting, the parent segment corresponding to the path that satisfies all constraints and is ranked first is retained. The inputs are the lateral constraint validity table, the temporal validity table, and the parent segment table. The processing involves: using each candidate segment in the lateral constraint validity group as a node, constructing a decision graph with the temporal sequential relationship between any two candidate segments under the same parent segment as edges, and writing the entry difference vector, departure difference vector, body-lateral sequence, adjacency sequence, and cross-axis sequence into the corresponding edge attributes. Then, the adjacency matrix is expanded on the decision graph to generate a set of reachable paths from the start node to the end node. Dynamic programming is then performed on each path to accumulate the number of positive entry components, the number of negative exit components, the number of consecutive frames on the same side, and the number of consecutive frames attached to each path. The number of positive entry components is the sum of the number of positive components in the entry difference vector of all juvenile candidate segments on the path. The number of negative exit components is the sum of the number of negative components in the exit difference vector of all juvenile candidate segments on the path. The number of consecutive frames on the same side is the total number of video frames corresponding to consecutive codes on the same side in all body-side sequences on the path. The number of consecutive frames attached to each other is the total number of video frames corresponding to consecutive codes attached to each attached sequence on the path. For each path, simultaneously check whether it contains cross-axis positions or conflicting frames. If it does, delete the path and do not participate in the sorting. Sort the remaining paths lexicographically in the order of the number of positive-sign entering components, negative-sign leaving components, number of consecutive frames on the same side, and number of consecutive frames attached. When the cumulative values of these four items are exactly the same, prioritize retaining the path with the smaller starting frame number of the corresponding parent segment. The parent segment corresponding to the path that ranks first in the sorting is determined as the parent-child segment, and the parent-child segment is written into the parent-child segment table. The output is the set of parent-child segments. The set of parent and child segments must include at least the parent and child segment number, the parent segment identifier, the order of the child candidate segment numbers, the cumulative path count, and the path sorting position. The set of parent and child segments is written into the parent and child segment table for S5 to read. The handling of anomalies or missing segments is as follows: when there is no path under the same parent segment that passes all constraints, the parent segment is not written into the parent and child segment table; when the same child candidate segment enters multiple paths with higher sorting positions at the same time, only the path with higher sorting position is retained, and the remaining paths delete the child candidate segment and re-execute the path accumulation and path sorting. Through the above implementation process, candidate segments of juveniles can be screened layer by layer from four levels: temporal relationship, lateral relationship, boundary attachment relationship, and cross-axis relationship. Finally, the parent segment corresponding to the candidate segment of the juvenile that truly satisfies the parent-juvenile attachment relationship is determined as the parent-juvenile segment, thereby avoiding misjudging short-term companionship, lateral jump, or discontinuous attachment as parent-juvenile. At the same time, this implementation process clearly defines the generation method, calculation method, anomaly handling, and unique output rules for complete manifestation time, disappearance time, entry difference vector, departure difference vector, lateral sequence, attachment sequence, cross-axis sequence, conflict matrix, conflict frame, path accumulation, and sorting position, thereby eliminating the problems of unclear temporal sequence, unclear lateral, unclear matrix, and unclear path preservation rules in S4. In practical applications: For example, when multiple candidate segments of juveniles are formed on the right side of the same parent segment, the complete appearance time and disappearance time are first solved based on the front end position, the back end position, and the middle bulge position. Then, the entry difference vector and the exit difference vector are constructed by combining the first appearance time and the last retention time of the candidate segments of juveniles. The candidate segments of juveniles that appear before the parent and retreat after the juveniles are screened out. Then, the candidate segments of juveniles are checked frame by frame to see if they are always located on the same side of the parent body, whether the boundary is continuous, and whether they cross the body axis. The conflict matrix is decomposed and the candidate segments of juveniles corresponding to the conflict frames are deleted. Finally, a decision graph is constructed for the remaining candidate segments of juveniles and the path is expanded and sorted. The parent segment corresponding to the first sorted path is output as the parent-juvenile segment.
[0021] S5. Merge adjacent and continuous mother-fetus segments into a mother-fetus fragment and generate the Yangtze finless porpoise mother-fetus identification result. In this embodiment, the purpose of S5 is to further group the established mother-fetal segments into mother-fetal fragments based on time and maternal continuity, and output the Yangtze finless porpoise mother-fetal identification results for retrieval, statistics, and verification. Its mechanism involves first establishing inter-segment continuity relationships based on the maternal segment identifier, first frame position, and last frame position of the mother-fetal segment; filtering out consecutive mother-fetal segments belonging to the same maternal continuous display process; then performing link expansion on the inter-segment continuity relationships to form mother-fetal fragments with a unique segment start point, a unique segment end point, and a unique maternal segment identifier order; finally, writing the mother-fetal fragments into the identification result table. Specifically, the mother-fetal connection diagram output by S5 is used for mother-fetal fragment expansion and reading; the mother-fetal fragment set output by S5 is used for identification result table generation and reading; and the identification result table output by S5 is used for subsequent statistical analysis and manual verification. This implementation process includes the following steps: The purpose of S5-1 is to establish inter-segment connections between mother-child segments that reflect the continuous appearance relationship of the same mother. Its mechanism is to use the continuation relationship of the mother segment and the frame order adjacency relationship to jointly constrain whether consecutive mother-child segments can be classified into the same mother-child segment. The input quantities are the mother-child segment table, the mother segment table, and the continuation relationship table. The processing actions are: reading the mother segment identifier, first frame position, and last frame position corresponding to each mother-child segment according to the mother-child segment table, and sorting according to the mother segment identifier and first frame position to construct the continuation relationship table. Each record includes at least the previous mother-child segment number, the next mother-child segment number, the previous mother segment identifier, the next mother segment identifier, the previous last frame position, and the... The position of the next frame; then, for the mother-juvenile segments that are adjacent in time, the mother segment continuation check and frame sequence continuation check are performed. The mother segment continuation check is performed according to whether the mother segment corresponding to the previous mother-juvenile segment and the mother segment corresponding to the next mother-juvenile segment belong to the same connection chain written by S2. Specifically, it is determined by reading the connection chain identifier in the mother segment table. When the connection chain identifier is consistent, it is recorded that the mother segment continuation is successful. The frame sequence continuation check is performed according to whether the position of the first frame of the next mother-juvenile segment is the first or second frame after the position of the last frame of the previous mother-juvenile segment. The allowed number of interval frames is given by the preset configuration, which is determined based on the frame rate of the aerial video and the statistical results of the number of frames of short-term occlusion of the Yangtze finless porpoise. When both the parent segment continuation check and the frame sequence continuation check are successful, an inter-segment connection is established between the previous parent segment and the next parent segment, and the corresponding connection is written into the parent-child connection table. The output is a parent-child connection graph, which includes at least the starting parent segment number, the ending parent segment number, the connection chain identifier, the frame sequence continuation mark, and the parent segment continuation mark. The parent-child connection graph is written into the parent-child connection graph for S5-2 to read. The handling of abnormalities or missing information is as follows: when the same previous parent segment can satisfy both types of checks with multiple subsequent parent segments, only the subsequent parent segment with the smaller difference between the first frame position and the last frame position of the previous parent segment is retained, and the remaining inter-segment connections are deleted. When the parent segment identifier or the connection chain identifier is missing, no inter-segment connection is established. The purpose of S5-2 is to expand the mother-child segments in the mother-child connection graph into mother-child fragments and form fragment records. Its working mechanism is to perform connection chain expansion with mother-child segments as nodes and inter-segment connections as directed edges, merging all mother-child segments in the same connection chain into the same mother-child fragment. The inputs are the mother-child connection graph, the mother-child segment table, and the mother segment table. The processing actions are as follows: construct a directed connection graph with mother-child segments as nodes and inter-segment connections as directed edges; find the mother-child segment without predecessor inter-segment connections in chronological order as the first mother-child segment, and perform connection chain expansion along the inter-segment connection direction until the current mother-child segment has no successor inter-segment connections, thus forming a mother-child connection chain; for each mother-child connection chain, read the first frame position of the first mother-child segment and record it as the fragment start point, read the last frame position of the last mother-child segment and record it as the fragment end point, and then read the corresponding mother segment identifiers in the order in which the mother-child segments appear in the connection chain, and write the mother segment identifiers sequentially into the fragment records. Then, all mother-child segments in the same mother-child connection chain are merged into the same mother-child segment, and a mother-child segment number is generated. The output is a mother-child segment set, which includes at least the mother-child segment number, segment start point, segment end point, mother segment identifier order, mother-child segment number order, and connection chain length. The mother-child segment set is written into the mother-child segment table for S5-3 to read. The abnormal or missing handling is as follows: when the same mother-child segment enters two mother-child connection chains at the same time, only the mother-child connection chain with the longer connection chain length is retained. When the connection chain lengths are the same, the mother-child connection chain with the earlier segment start point is retained. When two mother-child connection chains merge into the same mother-child segment in subsequent video frames, only the mother-child connection chain with more mother-child segments written before the merger is retained, and the other mother-child connection chain terminates at the mother-child segment before the merger. The purpose of S5-3 is to generate structured and searchable identification results of Yangtze finless porpoise mothers and calves based on the set of mother-calves segments. Its working mechanism is to write the target water area information, time range, and mother segment order corresponding to each mother-calves segment into the identification result table to form the final output. The inputs are the mother-calves segment table, the target water area record table, and the mother-calves segment table. The processing actions are: reading the mother-calves segment number, segment start point, segment end point, and mother segment identifier order one by one according to the set of mother-calves segments, reading the target water area identifier corresponding to each mother-calves segment in combination with the target water area record table, and counting the number of mother-calves segments contained in each mother-calves segment. Subsequently, an identification result table is generated. The identification result table includes at least the result number, target water area identifier, mother-fetal segment number, segment start point, segment end point, mother segment identifier order, and number of mother-fetal segments. Each field is written into the identification result table in the order of mother-fetal segment number. The output is the identification result of Yangtze finless porpoise mother-fetal segment, i.e., the completed identification result table, which is written to the result storage area for subsequent statistical analysis and manual review. The handling of anomalies or missing information is as follows: when the target water area identifier is missing for the same mother-fetal segment, the target water area identifier of the corresponding aerial video is directly read and filled in; when the segment start point is later than the segment end point, it is not written into the identification result table and the mother-fetal segment is recorded as an invalid segment; when there are duplicate entries in the mother segment identifier order, only the first entry is retained and subsequent duplicate entries are deleted. Through the above implementation process, the mother-fetus segments obtained from discrete determination can be further organized into mother-fetus segments with complete time range and clear maternal continuity, forming a Yangtze finless porpoise mother-fetus identification result with fixed fields, clear order, and direct callability. This avoids the repeated counting of multiple adjacent mother-fetus segments or the split output of the same maternal continuity process. At the same time, this implementation process clearly defines the rules for establishing inter-segment connections, the order of connection chain expansion, the method for determining the segment start and end points, the rules for writing the maternal segment identifier in order, and the composition of the identification result table fields and the handling of anomalies. This eliminates the problems of unclear continuity, unclear segment boundaries, and unclear entry points for writing the result table in S5. In practical applications: For example, in an aerial video of the same target water area, multiple mother-juvenile segments that are temporally adjacent and belong to the same connection chain have been output. First, inter-segment connections are established between these mother-juvenile segments based on the continuity relationship of the mother segments and the continuity relationship of the frame order. Then, the mother-juvenile segments with sequential connection relationships are expanded into the same mother-juvenile connection chain. The first frame position of the first mother-juvenile segment is taken as the segment start point, and the last frame position of the last mother-juvenile segment is taken as the segment end point. Subsequently, the mother segment identification order and the number of mother-juvenile segments in the connection chain are written into the recognition result table. Thus, the Yangtze finless porpoise mother-juvenile recognition results can be directly used for patrol statistics and manual review.
[0022] The working principle of this scheme is as follows: First, video frames are read sequentially from aerial videos of the target water area. Water surface background alignment is performed on adjacent video frames to remove interference caused by slight drone displacement and overall water surface changes. Then, continuously occurring difference areas are extracted, and the finless porpoise emergence areas are identified from these. Adjacent and consecutive finless porpoise emergence areas are grouped into display segments. Next, the front, rear, and central bulge positions are extracted from each display segment, and the body axis and sides are determined. Display segments that can be continuously connected are then joined into a parent segment. Then, on both sides of the parent segment boundary, adjacent display areas that are attached to the parent body are searched frame by frame and merged to form juvenile candidate segments. Finally, the results are combined... The candidate segments of the juvenile finless porpoise are screened layer by layer based on the order of appearance and disappearance, whether the sides of the segments are consistent, whether they are continuously attached, and whether they cross the body axis. This process determines the segments that truly satisfy the mother-juvenile attachment relationship. Finally, adjacent segments that belong to the same mother in a continuous process are merged into a single segment to generate the final identification result of the Yangtze finless porpoise mother-juvenile relationship. In other words, this method does not simply look for two finless porpoises in the picture, but rather narrows down the judgment range step by step by looking for the process of the finless porpoise emerging from the water, then for the continuous appearance of the mother, then for the attachment, and finally for whether the mother-juvenile relationship is satisfied. This reduces the possibility of misidentifying ordinary swimming finless porpoises as mother and finless porpoises. For example, when conducting drone patrols in the open sections of the middle and lower reaches of the Yangtze River, the footage often shows reflections, ripples, contrails, and finless porpoises briefly surfacing and then quickly diving back down. Young porpoises often remain close to their mothers, making it difficult for them to remain fully visible like independent targets. In such cases, this approach first separates the areas in the aerial video that truly belong to the finless porpoises emerging from the water surface changes. Then, it connects multiple emergence areas belonging to the continuous emergence process of the same porpoise into a "mother segment." Subsequently, it only searches for small, attached areas near the boundary of the mother segment that remain in close contact with the mother porpoise, and examines these attached areas. The criteria for identifying a finless porpoise are as follows: whether the visible area is always on the same side as the mother, whether it appears later than the mother, whether it disappears earlier than the mother, and whether it crosses the body axis to the other side. If these conditions are met consecutively, the corresponding mother segment is identified as a mother-fetus segment, and adjacent mother-fetus segments are merged into a complete mother-fetus segment. In this way, even if there are two finless porpoises swimming side by side, crossing each other, or following each other in the same direction in the picture, they will not be misidentified as mother and fetus just because they are close, have similar directions, or have been together for a long time. Instead, they will only be identified if the continuous performance of the attachment relationship is truly met.
[0023] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for identifying mother and calf Yangtze finless porpoises based on aerial imagery, characterized in that, include: S1. Acquire aerial video of the target water area, read each video frame in time sequence, perform water surface background alignment, difference region extraction and finless porpoise image recognition, group the finless porpoise out-of-water areas that appear continuously and are adjacent to each other into display segments, and output the display segment set. S2. Extract the front end position, rear end position and middle bulge position for each display segment. Determine the body axis based on the line connecting the front end position and the rear end position. Determine the body side based on the extension direction of the middle bulge position relative to the body axis. Connect the adjacent display segments with continuous body axis directions to form the parent segment and output the parent segment set. S3. For each parent segment, read the attached display area on both sides of the body axis that is connected to the parent segment frame by frame, perform attachment target image recognition on the attached display area, form attached records according to the first appearance time, the last retention time and the body side where it is located, and merge the attached records that are located on the same body side for multiple consecutive frames and are always attached to the boundary of the parent segment into juvenile candidate segments. S4. For each candidate segment of the juvenile body, compare the time of its first appearance with the time of the complete appearance of the corresponding parent segment, compare the time of its last retention with the time of its disappearance of the corresponding parent segment, and check the body side attribution in each video frame. If the candidate segment of the juvenile body appears later than the complete appearance of the parent segment, disappears earlier than the complete disappearance of the parent segment, and does not cross the body axis, the corresponding parent segment is determined as the parent-juvenile segment. S5. Merge adjacent mother-fetal segments that are consecutive to form a mother-fetal segment and generate the Yangtze finless porpoise mother-fetal identification result.
2. The method for identifying Yangtze finless porpoise mothers and calves based on aerial imagery according to claim 1, characterized in that: S1 includes: S1-1. Perform static water surface texture registration on adjacent video frames read in time sequence, solve the common water surface corresponding area between adjacent video frames, remove the edge areas that do not enter the common water surface corresponding area, and output the background aligned frame group. S1-2. Perform corresponding pixel difference calculation and connectivity merging on the background aligned frame group. Determine the difference regions in adjacent video frames that are continuously offset in position and have continuously connected ranges as candidate water exit areas. Perform finless porpoise image recognition on the candidate water exit areas and output the finless porpoise water exit areas. S1-3. Perform cross-frame continuation calculations on the order of appearance of the finless porpoise emergence areas, and merge the finless porpoise emergence areas with intersecting boundaries and continuous movement directions in adjacent video frames into the same display segment, and output the display segment set.
3. The method for identifying Yangtze finless porpoise mothers and calves based on aerial imagery according to claim 2, characterized in that: S2 includes: S2-1. For each display segment, extract the boundary point set of the display segment in the order of video frames, determine the end point group at the beginning of the time series and the end point group at the end of the time series in the boundary point set, and record the midpoint of the end point group as the front position and the back position respectively. Then, record the boundary point group where the distance between the front position and the back position is not zero as the middle raised position, and output the position record group. S2-2. Based on each position recording group, connect the front end position and the back end position to form a body axis, and use the body axis as the boundary to allocate the central bulge position to both sides of the body axis. Calculate the order of appearance of the side where the central bulge position is located in each video frame, and determine the side with the same appearance order as the body side. Output the body axis and body side.
4. The method for identifying Yangtze finless porpoise mothers and calves based on aerial imagery according to claim 3, characterized in that: S2 also includes: S2-3. For adjacent display segments in time, connect the rear end of the previous display segment with the front end of the next display segment to form a connecting line segment. Determine the turning direction of the connecting line segment with the body axis of the previous display segment and the body axis of the next display segment respectively. When the two turning directions are consistent and the body sides of the two display segments are consistent, the corresponding display segments are recorded as a connecting segment pair, and the connecting segment pair set is output. S2-4. Based on the set of continuation segments, perform concatenation in chronological order. Merge continuation segments in the previous continuation segment pair where the next displayed segment in the previous continuation segment pair is the same displayed segment as the previous displayed segment in the next continuation segment pair into the same connection chain. Merge the displayed segments in the same connection chain into the parent segment and output the parent segment set.
5. The method for identifying Yangtze finless porpoise mothers and calves based on aerial imagery according to claim 4, characterized in that: S3 includes: S3-1. Read the accompanying display areas on both sides of the boundary of each parent segment frame by frame, with the body axis as the boundary. Extract the boundary bonding position sequence, body axis projection interval and frame order position for each accompanying display area. Record the accompanying display areas with continuous boundary bonding position sequences, body axis projection intervals that do not cross the body axis and frame order positions that are consistent with the corresponding video frames of the parent segment as bonding units, and output the bonding unit set. S3-2. Establish a pairing relationship between the attaching unit sets according to the same parent segment, the same body side and adjacent frame sequence positions. Perform a continuation solution on the last attaching position of the previous attaching unit and the first attaching position of the next attaching unit. Establish a continuation edge when the two are located on the same side boundary of the parent segment and the corresponding body axis projection intervals overlap. Then delete the pairing relationship that does not satisfy the same side continuation or overlapping continuation and output the attaching continuation diagram.
6. The method for identifying Yangtze finless porpoise mothers and calves based on aerial imagery according to claim 5, characterized in that: S3 also includes: S3-3. Perform chain expansion using the attachment units in the attachment continuation diagram as nodes and the continuation edges as directed edges. Count the first frame position, the last frame position, the number of boundary attachment interruptions, and the number of body side rewrites for each expanded chain. Keep the expanded chains with zero boundary attachment interruptions and zero body side rewrites as accompanying record chains. At the same time, record the first frame position as the first occurrence time, the last frame position as the last retention time, and the body side where it is located as the record body side. Output the accompanying record set. S3-4. Establish record succession relationships for the accompanying record set according to the same parent segment and the same record body side. Perform frame sequence continuity check on the last retention time of the previous accompanying record and the first appearance time of the next accompanying record. Perform boundary coincidence check on the tail attachment position of the previous accompanying record and the first attachment position of the next accompanying record. When the frame sequence continuity check and the boundary coincidence check are successful, merge the corresponding accompanying records into the same young candidate segment and output the young candidate segment set.
7. The method for identifying Yangtze finless porpoise mothers and calves based on aerial imagery according to claim 6, characterized in that: S4 includes: S4-1. Read the front end position, back end position, and middle bulge position of each video frame according to the frame order position of the mother segment, and construct the front end sequence, back end sequence, and bulge sequence respectively. Perform linear interpolation on the front end sequence and back end sequence to fill in the missing frame positions. Perform three-point difference on the filled front end sequence and back end sequence to solve the volume axis change sequence. Then, record the first frame where the front end sequence, back end sequence, and bulge sequence all exist simultaneously as the complete appearance time, and the next frame after the last frame where they all exist simultaneously as the disappearance time. Record the first frame of the young body candidate segment as the first appearance time and the last frame as the last retention time. Output the time sequence record group. S4-2. Based on the time sequence record group, calculate the entry difference between the first appearance time and the complete appearance time, and the departure difference between the last retention time and the disappearance time, respectively. Construct the entry difference vector and the departure difference vector. Perform symbol decomposition on the entry difference vector and the departure difference vector. Retain the young candidate segments where all components of the entry difference vector are positive and all components of the departure difference vector are negative. Perform correlation backreading on the adjacent frames of the first frame and the last frame. When the correlation backreading result is consistent with the symbol decomposition result, output the time sequence establishment group.
8. The method for identifying Yangtze finless porpoise mothers and calves based on aerial imagery according to claim 7, characterized in that: S4 also includes: S4-3. Read the boundary bonding positions and body axis projection intervals of the juvenile candidate segments frame by frame according to the time sequence, and construct the body side sequence, bonding sequence, and cross-axis sequence. Perform sequence alignment on the body side sequence and record the positions where the body side codes of adjacent frames are inconsistent as overwrite bits. Perform interval intersection operation on the bonding sequence and record the positions where the bonding intervals of adjacent frames intersect empty as disconnection bits. Perform projection sign determination on the cross-axis sequence and record the positions where the body axis projection interval falls on both sides of the body axis as cross-axis positions. Then construct a conflict matrix composed of overwrite bits, disconnection bits, and cross-axis positions. Perform Gram product and eigenvalue decomposition on the conflict matrix and record the frame positions where the eigenvalue is not zero as conflict frames. Delete the juvenile candidate segments containing conflict frames and output the side-based group.
9. The method for identifying mother and calf Yangtze finless porpoises based on aerial imagery according to claim 8, characterized in that: S4 also includes: S4-4. Using each candidate segment of the juvenile in the lateral group as a node, construct a decision graph with the entry difference vector, departure difference vector, body side sequence, attachment sequence, and cross-axis sequence as edge attributes. Perform adjacency matrix expansion and dynamic programming pathfinding on the decision graph. Accumulate the number of positive entry components, negative departure components, consecutive frames on the same side, and consecutive frames attached in each path. Delete paths containing cross-axis positions or conflicting frames. Then, sort the path in lexicographical order according to the order of the number of positive entry components, negative departure components, consecutive frames on the same side, and consecutive frames attached. Determine the parent segment corresponding to the first sorted path as the parent juvenile segment and output the set of parent juvenile segments.
10. The method for identifying Yangtze finless porpoise mothers and calves based on aerial imagery according to claim 9, characterized in that: S5 includes: S5-1. Construct a continuation relationship table according to the parent segment identifier, first frame position and last frame position of each parent segment. Perform parent segment continuation check and frame sequence continuation check on parent segments that are adjacent in time. When the last frame position of the previous parent segment and the first frame position of the next parent segment are adjacent in time and the corresponding parent segments belong to the same continuation chain, establish inter-segment connection and output parent-child connection diagram. S5-2. Using the mother-child segments in the mother-child connection graph as nodes and the inter-segment connections as directed edges, perform connection chain expansion. Record the first frame position of the first mother-child segment in the same connection chain as the fragment start point and the last frame position of the last mother-child segment as the fragment end point. Write the mother segment identifiers in the same connection chain into the fragment record in order and output the mother-child fragment set. S5-3. Generate the Yangtze finless porpoise mother-feather identification results according to the mother-feather segment set. Write the target water area identifier, segment start point, segment end point, mother segment identifier order and the number of mother-feather segments corresponding to each mother-feather segment into the identification result table, and output the Yangtze finless porpoise mother-feather identification results.