Motion estimation method, motion estimation device and equipment thereof
By performing semantic segmentation on video frame sequences and matching foreground targets with the same semantic category, and combining camera pose change information to generate virtual targets, the accuracy problem of motion estimation in complex scenes in existing technologies is solved, achieving higher accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AUTOCHIPS
- Filing Date
- 2025-11-24
- Publication Date
- 2026-04-24
AI Technical Summary
Existing motion estimation methods are easily affected by complex background changes, occlusion, and illumination changes, leading to inaccurate estimations.
By performing semantic segmentation on each frame of the video frame sequence to be processed, the semantic category and position mask of the foreground target are determined, and foreground targets with the same semantic category are matched in the video frame sequence. Virtual targets are generated by combining camera pose change information, and motion estimation results are calculated.
It effectively overcomes errors caused by target occlusion and appearance changes, improves environmental understanding and decision-making capabilities in complex scenarios, and enhances the accuracy and robustness of motion estimation.
Smart Images

Figure CN121921337A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision processing technology, and in particular to a motion estimation method, a motion estimation device, and a computer storage medium / computer program product. Background Technology
[0002] In the field of computer vision, motion estimation is a key technology, widely used in video analysis, object tracking, intelligent monitoring, autonomous driving, and other scenarios. Currently, most mainstream motion estimation methods rely on optical flow or feature matching. These methods are susceptible to interference, leading to inaccurate estimations, especially in situations with complex background changes, occlusion, lighting variations, or large parallax. Summary of the Invention
[0003] To address the aforementioned technical problems, this application proposes a motion estimation method, a motion estimation device, and a computer storage medium / computer program product.
[0004] To address the aforementioned technical problems, this application proposes a motion estimation method, which includes: performing semantic segmentation on each frame of a video frame sequence to be processed, determining the semantic category and position mask of a foreground target in each frame; matching foreground targets of the same semantic category in the video frame sequence to be processed; and determining the motion estimation result of the foreground target based on the changes of the foreground target in the video frame sequence to be processed.
[0005] The step of determining the semantic category and position mask of the foreground target in each frame image includes: determining the position mask distance of foreground targets with the same semantic category in adjacent frame images; when the position mask distance is greater than a preset threshold, assigning different target numbers to two foreground targets with the same semantic category; the step of matching foreground targets with the same semantic category in the video frame sequence to be processed includes: matching foreground targets with the same semantic category and the same target number in the video frame sequence to be processed.
[0006] The step of determining the semantic category and location mask of the foreground target in each frame image includes: determining the location mask distance of the foreground target with the same semantic category in adjacent frame images; For a foreground target in one of the adjacent frame images, if a foreground target of the same semantic category cannot be determined in another frame image, other neighboring frame images of the image where the foreground target is located are acquired until the target neighboring frame image of the foreground target of the same semantic category is determined; a virtual foreground target is generated in the image between the image where the foreground target is located and the target neighboring frame image according to the camera pose change information.
[0007] The step of matching foreground targets of the same semantic category in the video frame sequence to be processed includes: sequentially obtaining a set of foreground targets of the same semantic category in adjacent frame images in the video frame sequence to be processed; calculating the intersection-union ratio (CIU) of each pair of foreground targets from adjacent frame images in the foreground target set; determining the matching score of each pair of foreground targets from adjacent frame images based on the CIU; and establishing a matching relationship for foreground targets with matching scores higher than a preset threshold.
[0008] Wherein, after sequentially obtaining the set of foreground objects of the same semantic category in adjacent frame images in the video frame sequence to be processed, the motion estimation method further includes: Calculate the center point distance between each pair of foreground targets in the foreground target set, traversing from adjacent frame images; the step of determining the matching score of each pair of foreground targets from adjacent frame images based on the intersection-union ratio includes: determining the comprehensive matching score of each pair of foreground targets from adjacent frame images based on the intersection-union ratio and the center point distance.
[0009] The step of determining the motion estimation result of the foreground target based on the changes of the foreground target in the video frame sequence to be processed includes: calculating the change in the center point distance of the foreground target in adjacent frame images, and / or the change in the mask region, to determine the motion estimation result of the foreground target.
[0010] The step of determining the motion estimation result of the foreground target based on the changes of the foreground target in the video frame sequence to be processed includes: determining the world coordinates of the foreground target based on the two-dimensional position and depth information of the center point of the foreground target in adjacent frame images; and determining the motion estimation result of the foreground target based on the world coordinate motion vector of the foreground target in adjacent frame images.
[0011] The step of semantic segmentation of each frame in the video frame sequence to be processed includes: inputting a target detection model into each frame in the video frame sequence to obtain a detection box for each foreground target; performing semantic segmentation on the image region within the detection box in each frame in the video frame sequence; and removing the foreground target when the semantic category obtained by the semantic segmentation is different from the semantic category obtained by the target detection.
[0012] To address the aforementioned technical problems, this application proposes a motion estimation device, which includes a memory and a processor coupled to the memory; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the aforementioned motion estimation method.
[0013] To address the aforementioned technical problems, this application proposes a computer storage medium / computer program product, wherein the computer storage medium is used to store program data, which, when executed by a computer, is used to implement the aforementioned motion estimation method; and / or, The computer program product includes a computer program that, when executed by a processor, implements the method described above.
[0014] Compared with existing technologies, the beneficial effects of this application are as follows: the motion estimation device performs semantic segmentation on each frame of the video frame sequence to be processed, determines the semantic category and position mask of the foreground target in each frame; matches foreground targets of the same semantic category in the video frame sequence to be processed; and determines the motion estimation result of the foreground target based on the changes of the foreground target in the video frame sequence to be processed. Through the above method, errors caused by target occlusion and appearance changes are effectively overcome, improving environmental understanding and decision-making capabilities in complex scenes, and further improving the accuracy of motion estimation. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating the first embodiment of the motion estimation method provided in this application; Figure 2 This is a flowchart illustrating a second embodiment of the motion estimation method provided in this application; Figure 3 This is a flowchart illustrating the third embodiment of the motion estimation method provided in this application; Figure 4 This is a flowchart illustrating the fourth embodiment of the motion estimation method provided in this application; Figure 5 This is a diagram comparing the effects of this application with existing technologies; Figure 6 This is a schematic diagram of the structure of an embodiment of the motion estimation device provided in this application; Figure 7 This is a schematic diagram of the structure of an embodiment of the computer storage medium / computer program product provided in this application. Detailed Implementation
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0017] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0018] Please refer to the details. Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the motion estimation method provided in this application.
[0019] The motion estimation method of this application is applied to a motion estimation device, which can be a server, a terminal device, or a system in which the server and the terminal device cooperate with each other. Accordingly, the various parts of the motion estimation device, such as each unit, subunit, module, and submodule, can all be set in the server, all in the terminal device, or separately in the server and the terminal device.
[0020] Furthermore, the aforementioned server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software modules, such as software or software modules used to provide distributed server functionality, or as a single software program or software module; no specific limitations are made here.
[0021] Step S11: Perform semantic segmentation on each frame of the video frame sequence to be processed, and determine the semantic category and location mask of the foreground target in each frame.
[0022] By using semantic segmentation technology, each image in the video sequence is analyzed frame by frame to accurately identify the specific semantic category of foreground targets, such as pedestrians and vehicles, and generate corresponding pixel-level position masks to mark their spatial distribution.
[0023] In one specific embodiment of this application, deep neural networks, such as DeepLabv3+, PSP Net, SegFormer, etc., can be used to perform semantic segmentation on each frame of the video image to obtain the category and location mask of the foreground object in each frame.
[0024] In one embodiment of this application, the motion estimation method performs semantic segmentation on each frame of a video image, extracts foreground objects (such as people, vehicles, etc.) from each frame, and assigns semantic labels. The specific method is as follows: Given an input frame image, a deep semantic segmentation network S(F) is used to calculate the category probability map for each pixel:
[0025] Where H and W are pixel dimensions, and C is the number of categories.
[0026] The argmax operation is used to obtain the segmentation image, as follows:
[0027] The final output image contains the category label map of each pixel. .
[0028] Remove the background category from the segmentation image, retaining only the foreground object. The specific method is as follows: Let the background category be Then the foreground mask is defined as:
[0029] Finally, output the foreground mask image. This is used for subsequent target localization and matching.
[0030] In one embodiment of this application, when performing semantic segmentation on each frame of an image, foreground targets are removed. The motion estimation device inputs a target detection model into each frame of the video frame sequence to be processed, and obtains a detection box for each foreground target. Semantic segmentation is performed on the image region within the detection box in each frame of the video frame sequence to be processed. If the semantic category obtained from semantic segmentation is different from the semantic category obtained from target detection, the foreground target is removed.
[0031] Specifically, the motion estimation device uses an object detection model to perform preliminary analysis on each frame of the image, obtaining bounding boxes for all foreground objects. Subsequently, the system performs refined semantic segmentation only on the image region defined by each bounding box, rather than processing the entire image, thereby improving segmentation efficiency and focus. Finally, a decision verification mechanism is introduced—if the semantic category given by the object detection model is inconsistent with the category obtained by the semantic segmentation model, the object recognition result is deemed unreliable and removed. This process, through dual model cross-validation, effectively filters out noisy objects caused by misjudgments from a single model, thus significantly improving the accuracy and robustness of foreground object extraction in complex scenes.
[0032] This application proposes an embodiment for determining the semantic category and position mask of foreground targets in each frame of an image. By calculating the centroid distance or overlap of the position masks of similar targets between adjacent frames and comparing it with a preset motion consistency threshold, when the spatial displacement of two similar targets exceeds a reasonable range, the motion estimation device assigns them different target identifiers. This design not only achieves basic target detection and localization but also effectively distinguishes between objects that are similar in appearance but different in substance through motion logic discrimination. This provides a target matching foundation with both semantic consistency and spatiotemporal continuity for subsequent trajectory tracking, thereby significantly improving the accuracy and robustness of multi-target tracking in complex scenes.
[0033] Please see details. Figure 2 , Figure 2 This is a flowchart illustrating a second embodiment of the motion estimation method provided in this application.
[0034] like Figure 2 As shown, the specific steps are as follows: Step S21: Determine the position mask distance of foreground targets of the same semantic category in adjacent frame images.
[0035] Specifically, the motion estimation device selects two consecutive image frames and pairs all foreground targets with the same semantic category. For each pair of targets, the spatial distance between their position masks is calculated. This distance objectively reflects the magnitude of possible displacement of the same target between the two frames, providing a crucial metric for subsequent determination of whether they are continuous motions of the same individual or independent occurrences of different individuals.
[0036] Step S22: When the location mask distance is greater than a preset threshold, set different target numbers for two foreground targets of the same semantic category.
[0037] The motion estimation device compares the calculated position mask distance with a preset threshold based on reasonable motion speed. If the distance is less than or equal to the threshold, the two targets are determined to be the same entity, and they are assigned the same target number to maintain identity consistency. If the distance is significantly greater than the threshold, it indicates that the displacement exceeds the reasonable motion range of the object between adjacent frames, thus determining that they belong to different individual instances. In this case, the motion estimation device assigns a completely new and different target number to the target in the next frame, effectively preventing multiple targets that look similar but are substantially different from each other from being mistakenly tracked as the same entity, ensuring the accuracy of multi-target tracking.
[0038] Furthermore, this application also proposes an embodiment for determining a foreground target; please refer to the details below. Figure 3 , Figure 3 This is a flowchart illustrating the third embodiment of the motion estimation method provided in this application.
[0039] like Figure 3 As shown, the specific steps are as follows: Step S31: Determine the position mask distance of foreground targets of the same semantic category in adjacent frame images.
[0040] Specifically, the motion estimation device first selects two adjacent images, such as frame t and frame t+1, and then performs pairwise comparisons on all foreground objects belonging to the same semantic category in these two frames.
[0041] For each pair of targets of the same category, the motion estimation device calculates the spatial distance between their position masks. This distance can be measured by calculating the Euclidean distance between the centroids of the two masks or by calculating the distance between the centers of their bounding boxes. The core purpose is to obtain a quantitative indicator to determine whether the two targets are the same object in physical space and the positional change caused by motion.
[0042] Step S32: For a foreground target in one of the adjacent frame images, if a foreground target of the same semantic category cannot be determined in another frame image, continue to acquire other neighboring frame images of the image where the foreground target is located until the target neighboring frame image where the foreground target of the same semantic category is located is determined.
[0043] Specifically, when a target cannot find any matching target of the same category in the next frame, it is not immediately considered lost; instead, the search scope is expanded. The search continues to examine more frames after the frame containing the target until a matching target with the same semantic category as target A is successfully found in a certain frame. This process essentially involves a broader open search along the timeline to recover the briefly interrupted tracking trajectory.
[0044] Step S33: Generate a virtual foreground target in the image between the image where the foreground target is located and the neighboring frame image of the target, according to the camera pose change information.
[0045] After finding a matching relationship spanning several frames in step S32, step S33 aims to fill in the target information in all intermediate frames between these two frames.
[0046] The motion estimation device uses known camera pose change information to calculate the most likely position and shape of the target in each intermediate frame through motion interpolation or three-dimensional geometric transformation, based on the target's position in the starting frame and the target's position in the ending frame.
[0047] These generated targets are called virtual foreground targets. They fill in the missing trajectory segments caused by the temporary disappearance of the targets, thus forming a continuous, complete and smooth motion trajectory in time, which greatly improves the robustness of the tracking system and the continuity of the output trajectory.
[0048] Furthermore, in other embodiments of this application, an instance segmentation network can be combined to directly distinguish different individual instances while identifying semantic categories and location masks, providing a more granular foundation for subsequent tracking and matching. The core of these methods lies in combining visual recognition with spatiotemporal context to achieve accurate and stable foreground target perception in dynamic scenes.
[0049] Step S12: Match foreground targets of the same semantic category in the video frame sequence to be processed.
[0050] Matching foreground objects with the same semantic category in a video frame sequence refers to establishing a temporal correspondence between object instances with the same semantic labels, such as "pedestrian" or "vehicle," in consecutive frames of the video.
[0051] Specifically, the motion estimation device extracts all candidate targets belonging to the same category in adjacent frames and performs correlation matching by calculating their appearance feature similarity, spatial position overlap, or motion consistency, thereby associating foreground targets that actually belong to the same physical entity in different frames to form a continuous and stable tracking trajectory.
[0052] This process effectively solves the problem of identity maintenance in complex scenarios such as target occlusion and shape changes, providing a reliable data foundation for subsequent behavior analysis and motion prediction.
[0053] In a specific embodiment of this application, foreground targets with the same semantic category and the same target number can be matched by using a sequence number.
[0054] Specifically, in sequences where target detection and ID assignment have been completed, stable cross-frame tracking is achieved by matching foreground targets with the same semantic category and the same target sequence number. The core advantage of this mechanism lies in combining high-order semantic information with individual identity information, creating a double-checked matching condition. Semantic consistency ensures the correctness of the matched target type, while the continuity of the target sequence number guarantees the spatiotemporal consistency of the individual identity. Its direct benefit is a significant improvement in tracking robustness in complex scenarios—effectively resisting identity confusion caused by objects with similar appearances, and overcoming the ID switching problem caused by the reappearance of a target after brief occlusion or disappearance. Ultimately, it provides continuous, accurate, and identity-consistent target trajectory data for high-level applications such as behavior analysis and trajectory prediction.
[0055] Furthermore, this application proposes an embodiment for matching foreground targets of the same semantic category in a sequence of video frames to be processed. For details, please refer to [link to specific embodiments]. Figure 4 , Figure 4 This is a flowchart illustrating the fourth embodiment of the motion estimation method provided in this application.
[0056] like Figure 4 As shown, the specific steps are as follows: Step S41: Sequentially obtain the set of foreground targets of the same semantic category in adjacent frame images in the video frame sequence to be processed.
[0057] After sequentially obtaining the set of foreground targets of the same semantic category in adjacent frame images in the video frame sequence to be processed, the motion estimation device calculates the distance between the center points of each pair of foreground targets in the set of foreground targets traversing adjacent frame images.
[0058] Step S42: Calculate the intersection-union ratio of each pair of foreground targets in the foreground target set, traversing from adjacent frame images.
[0059] For each frame of the image, connected regions are extracted as target candidates, and a target set is defined:
[0060] Each Includes: Category c, mask area Center point .
[0061] The Intersection over Union (IoU) is calculated as follows:
[0062] Step S43: Determine the matching score of each pair of foreground targets from adjacent frame images based on the cross-union ratio.
[0063] Specifically, the motion estimation device determines the comprehensive matching score of each pair of foreground targets from adjacent frame images based on the cross-union ratio and the center point distance.
[0064] Specifically, the calculation method for category consistency is as follows:
[0065] The calculation method for the center point distance constraint is as follows:
[0066] Step S44: Establish matching relationships for foreground targets whose matching scores are higher than a preset threshold.
[0067] Specifically, the overall matching score is:
[0068] Furthermore, in the embodiments of this application, the Hungarian algorithm is used. Algorithm or greedy matching yields a set of matching pairs. .
[0069] Step S13: Determine the motion estimation result of the foreground target based on the changes of the foreground target in the video frame sequence to be processed.
[0070] Specifically, in the embodiments of this application, the motion estimation device calculates the change in the center point distance of the foreground target in adjacent frame images, and / or the change in the mask region, to determine the motion estimation result of the foreground target.
[0071] Specifically, the motion estimation device estimates the motion state of the foreground target by analyzing the spatiotemporal evolution of the target in consecutive video frames. Specifically, it calculates the displacement vector of the target's center point between adjacent frames to capture its translational motion and velocity, while analyzing the shape and area changes of the target's position mask region to infer its radial motion, such as scale scaling or shape changes caused by moving closer to or away from the camera.
[0072] This dual mechanism, which combines centroid trajectory analysis and contour deformation detection, can comprehensively calculate the target's motion vector, velocity, and motion trend, thereby providing accurate kinematic dynamics basis for subsequent behavior analysis, trajectory prediction, or collision risk assessment.
[0073] In one embodiment of this application, the motion estimation device determines the world coordinates of the foreground target based on the two-dimensional position and depth information of the center point of the foreground target in adjacent frame images, and determines the motion estimation result of the foreground target based on the world coordinate motion vector of the foreground target in adjacent frame images.
[0074] Specifically, for matching target pairs Its motion vector is defined as:
[0075] If shape changes are taken into account, the change in region contour can be further introduced:
[0076] If depth information is fused This can then be converted into a world coordinate motion vector:
[0077] The motion estimation device first fuses the two-dimensional pixel coordinates of the center point of the foreground target in adjacent frames with its corresponding depth information. This information is then back-projected into three-dimensional physical space using the camera's imaging geometry model, thereby calculating the target's precise position in the real-world coordinate system. Based on this, by analyzing the displacement of the same target in world coordinates across consecutive frames, its motion vector and instantaneous velocity in three-dimensional space are calculated. This process achieves motion mapping from the two-dimensional image plane to the three-dimensional physical world, enabling the motion estimation results to more realistically reflect the target's actual motion state. This provides crucial depth perception and motion dynamics parameters for applications such as autonomous driving and robot navigation.
[0078] Please continue reading Figure 5 , Figure 5 This is a diagram comparing the effects of this application with existing technologies. Figure 5 The paper 30 demonstrates a motion estimation method based on semantic segmentation. By segmenting different categories of targets in an image, such as cyclists and bicycles, at the pixel level and modeling their motion information separately, it achieves a fine understanding of complex multi-target motion scenes. Figure 5 The method described in section 31 is a traditional motion estimation method that relies solely on the overall optical flow estimation of pixels or image blocks. It fails to distinguish semantic categories, leading to confusion in the motion description of structurally complex objects. This comparison highlights the crucial role of semantic information in improving the accuracy and interpretability of motion estimation in this application.
[0079] Please refer to the details. Figure 6 , Figure 6 This is a schematic diagram of an embodiment of the motion estimation device provided in this application.
[0080] The motion estimation device 700 of this embodiment includes a processor 71, a memory 72, an input / output device 73, and a bus 74.
[0081] The processor 71, memory 72, and input / output device 73 are connected to the bus 74. The memory 72 stores program data, and the processor 71 is used to execute the program data to implement the motion estimation method described in the above embodiments.
[0082] In this embodiment, processor 71 can also be referred to as a CPU (Central Processing Unit). Processor 71 may be an integrated circuit chip with signal processing capabilities. Processor 71 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor, or processor 71 can be any conventional processor.
[0083] This application also provides a computer storage medium / computer program product; please refer to the following: Figure 7 , Figure 7 This is a schematic diagram of a structure of an embodiment of the computer storage medium / computer program product provided in this application. The computer storage medium 600 stores a computer program 61, which, when executed by a processor, implements the motion estimation method of the above embodiment. The computer program product 600 includes the computer program 61, which, when executed by a processor, implements the method described above.
[0084] When the embodiments of this application are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0085] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A motion estimation method, characterized in that, The motion estimation method includes: Semantic segmentation is performed on each frame of the video frame sequence to be processed to determine the semantic category and location mask of the foreground target in each frame; Match foreground targets of the same semantic category in the video frame sequence to be processed; The motion estimation result of the foreground target is determined based on the changes of the foreground target in the video frame sequence to be processed.
2. The motion estimation method according to claim 1, characterized in that, Determining the semantic category and location mask of the foreground object in each frame of the image includes: Determine the location mask distance of foreground targets of the same semantic category in adjacent frame images; When the location mask distance is greater than a preset threshold, different target numbers are assigned to two foreground targets of the same semantic category. The step of matching foreground targets of the same semantic category in the video frame sequence to be processed includes: In the sequence of video frames to be processed, foreground targets with the same semantic category and the same target index are matched.
3. The motion estimation method according to claim 1 or 2, characterized in that, Determining the semantic category and location mask of the foreground object in each frame of the image includes: Determine the location mask distance of foreground targets of the same semantic category in adjacent frame images; For a foreground target in one of the adjacent frame images, if a foreground target of the same semantic category cannot be determined in another frame image, other neighboring frame images of the image where the foreground target is located are acquired until the target neighboring frame image of the foreground target of the same semantic category is determined. A virtual foreground target is generated in the image between the image where the foreground target is located and the neighboring frame image of the target, according to the camera pose change information.
4. The motion estimation method according to claim 1, characterized in that, The step of matching foreground targets of the same semantic category in the video frame sequence to be processed includes: In the video frame sequence to be processed, the set of foreground targets with the same semantic category in adjacent frame images is obtained sequentially; Calculate the intersection-union ratio (IUGR) of each pair of foreground targets in the foreground target set, traversing from adjacent frame images; The matching score of each pair of foreground targets from adjacent frame images is determined based on the intersection-union ratio; Establish matching relationships for foreground targets whose matching scores are higher than a preset threshold.
5. The motion estimation method according to claim 4, characterized in that, After sequentially obtaining the set of foreground objects of the same semantic category in adjacent frame images in the video frame sequence to be processed, the motion estimation method further includes: Calculate the distance between the center points of each pair of foreground objects in the foreground object set, traversing from adjacent frame images; The step of determining the pairwise foreground target matching score from adjacent frame images based on the intersection-union ratio includes: The overall matching score of each pair of foreground targets from adjacent frame images is determined based on the intersection-union ratio and the center point distance.
6. The motion estimation method according to claim 1, characterized in that, The step of determining the motion estimation result of the foreground target based on the changes of the foreground target in the video frame sequence to be processed includes: Calculate the change in the center point distance of the foreground target in adjacent frame images, and / or the change in the mask region, to determine the motion estimation result of the foreground target.
7. The motion estimation method according to claim 6, characterized in that, The step of determining the motion estimation result of the foreground target based on the changes of the foreground target in the video frame sequence to be processed includes: The world coordinates of the foreground target are determined based on the two-dimensional position and depth information of the center point of the foreground target in adjacent frame images; The motion estimation result of the foreground target is determined based on the world coordinate motion vector of the foreground target in adjacent frame images.
8. The motion estimation method according to claim 1, characterized in that, The semantic segmentation of each frame in the video frame sequence to be processed includes: Input the target detection model into each frame of the video frame sequence to be processed, and obtain the detection box of each foreground target; Semantic segmentation is performed on the image region within the detection box in each frame of the video frame sequence to be processed; When the semantic category obtained from semantic segmentation is different from the semantic category obtained from target detection, the foreground target is removed.
9. A motion estimation device, characterized in that, The motion estimation device includes a memory and a processor coupled to the memory; The memory is used to store program data, and the processor is used to execute the program data to implement the motion estimation method as described in any one of claims 1 to 8.
10. A computer storage medium / computer program product, characterized in that, The computer storage medium is used to store program data, which, when executed by the computer, is used to implement the motion estimation method as described in any one of claims 1 to 8; and / or, The computer program product includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-8.