The invention discloses a reference video object segmentation method based on
semantic consistency and
motion perception. The method comprises the following steps: 1, constructing semantic prompt information and video frame information of a reference video object segmentation
data set; 2, preprocessing the reference video object segmentation
data set; 3, establishing a reference video object segmentation model based on
semantic consistency and
motion perception: designing a double-
branch decoupling strategy for decoupling feature information at semantic and visual levels so as to extract static and motion information of text description and visual features; a hierarchical
motion sensing module is designed to capture and align motion information between different frames, and analyze short-term and long-term motion information, so that the model obtains a sensing ability for a long-term
motion mode; the
semantic consistency module is designed to align semantic description and video features, so that the accuracy of target selection and the integrity of masks are improved, and
false detection of negative samples is avoided; a perceptual dynamic
fusion mechanism is designed to be used for embedding text information into a visual feature space, so that visual features can obtain text
semantic information, and the cross-
modal understanding ability of the model is enhanced; 4, constructing a
loss function, updating
model parameters, setting training parameters, and training to obtain an
optimal weight; and 5, detecting the
test set image based on the
optimal weight to obtain a final segmentation result. According to the method, static and dynamic information is effectively decoupled, the sensing ability to an
object motion mode is enhanced, and the segmentation performance is improved.