Main body motion characterization extractor training method, device, medium and product
By training the subject motion representation extractor, combining the loss function of the perspective video and the sub-network reconstruction technology, the problem of insufficient capture of the intelligent body's motion information is solved, efficient motion information perception and driving strategy-related feature extraction are achieved, and the accuracy of perception is improved.
Patent Information
- Application Number
- CN202510855047.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies are unable to accurately capture the motion information of the intelligent body, resulting in insufficient accuracy in motion information perception. Especially in first-person perspective videos, traditional methods are unable to effectively extract features related to driving strategies.
The subject motion representation extractor training method is adopted. Through the first-person video combined with the loss between the actual next frame of the training frame and the reconstructed frame, as well as the loss between the predicted six degrees of freedom and the standard six degrees of freedom, the trained external parameter detection subnetwork, internal parameter detection subnetwork and depth estimation subnetwork are used for photometric reconstruction to train a neural network model that can capture the motion characteristics of embodied intelligent bodies.
It improves the accuracy of motion information perception, can filter out noise information, focus on areas related to driving strategies, accurately capture the motion information of the intelligent body, and enhance the accuracy of driving decisions.
Smart Images

Figure CN120708135A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a subject motion representation extractor training method, equipment, medium and product. Background Art
[0002] Embodied intelligence (Embodied Intelligence) is a cutting-edge field at the intersection of artificial intelligence and robotics. It emphasizes the autonomous learning and evolution of intelligent agents through dynamic interactions between their bodies and their environments. Its core lies in the deep integration of perception, action, and cognition. Embodied agents must fully understand human intent expressed in verbal commands, proactively explore their surroundings, fully perceive multimodal elements from both the virtual and physical environments, and execute appropriate actions to complete complex tasks. Therefore, embodied agent perception is crucial to achieving embodied intelligence.
[0003] Video is an important carrier of multimodal information and one of the foundations of motion perception. Related technologies use methods such as Residual Networks (ResNet) and Inflated 3D ConvNets (I3D) to extract features from video frames or clips. However, these methods only work with third-person videos and cannot capture the motion of the agent itself. Summary of the Invention
[0004] In view of this, the present invention aims to provide a method, device, medium, and product for training a subject motion representation extractor, which can accurately capture the motion information of intelligent agents and improve the accuracy of motion information perception. The specific solution is as follows: In a first aspect, the present application discloses a method for training a subject motion representation extractor, comprising: Inputting a training frame into a motion representation extractor of a subject to be trained, and obtaining a predicted six degrees of freedom according to the output; the training frame is a video frame in a first-person perspective video; Determining a total loss using a first loss function; the total loss includes a loss between an actual next frame of the training frame and a reconstructed frame, and a loss between the predicted six degrees of freedom and a standard six degrees of freedom corresponding to the training frame; the reconstructed frame is a next frame of the training frame reconstructed based on the predicted six degrees of freedom; The parameters of the subject motion representation extractor to be trained are adjusted according to the total loss until the training conditions are met to obtain a subject motion representation extractor that can capture the motion characteristics of the embodied intelligent body, so as to use the six degrees of freedom predicted by the subject motion representation extractor to perceive the motion information of the embodied intelligent body.
[0005] In a second aspect, the present application discloses a method for perceiving motion information of an embodied intelligent body, comprising: Obtain the target image frame from the first perspective currently captured by the embodied intelligent body; Inputting the target image frame into the aforementioned subject motion representation extractor to obtain the six degrees of freedom between the target image frame and the next image frame predicted by the subject motion representation extractor; The motion information of the embodied intelligent body is perceived based on the six degrees of freedom.
[0006] In a third aspect, the present application discloses an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned subject motion representation extractor training method or embodied intelligent body motion information perception method.
[0007] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein when the computer program is executed by a processor, it implements the aforementioned subject motion representation extractor training method, or embodied intelligent body motion information perception method.
[0008] In a fifth aspect, the present application discloses a computer program product, including a computer program, which, when executed by a processor, implements the aforementioned subject motion representation extractor training method, or embodied intelligent body motion information perception method.
[0009] In the present application, a training frame is input into a motion representation extractor of a subject to be trained, and a predicted six degrees of freedom is obtained based on the output; the training frame is a video frame in a first-person perspective video; a total loss is determined using a first loss function; the total loss includes a loss between the actual next frame of the training frame and a reconstructed frame, and a loss between the predicted six degrees of freedom and the standard six degrees of freedom corresponding to the training frame; the reconstructed frame is the next frame of the training frame reconstructed based on the predicted six degrees of freedom; parameters of the motion representation extractor of the subject to be trained are adjusted according to the total loss until the training conditions are met to obtain a subject motion representation extractor that can capture the motion characteristics of the embodied intelligent body, so as to perceive the motion information of the embodied intelligent body using the six degrees of freedom predicted by the subject motion representation extractor.
[0010] The beneficial effects of the present invention are: based on the first-person perspective video combined with the loss between the actual next frame of the training frame and the reconstructed frame of the training frame, as well as the loss between the predicted six degrees of freedom and the standard six degrees of freedom corresponding to the training frame, the subject motion representation extractor is trained to enable the subject motion representation extractor to learn whether various detailed information within the field of view has an impact on driving behavior. The trained subject motion representation extractor can perform efficient feature extraction on the first-person perspective video, obtain accurate and effective six degrees of freedom, filter out noise information irrelevant to driving and focus on areas directly related to driving strategies, accurately capture the motion information of the intelligent body in the first-person perspective video, and thus improve the accuracy of motion information perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work. Figure 1 A flow chart of a subject motion representation extractor training method provided in this application; Figure 2 A schematic diagram of photometric reconstruction training for a specific extrinsic detection subnetwork, intrinsic detection subnetwork, and depth estimation subnetwork provided in this application; Figure 3 A training diagram of a specific subject motion representation extractor to be trained provided in this application; Figure 4 A specific subject motion representation extractor optimization schematic diagram provided in this application. DETAILED DESCRIPTION
[0012] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0013] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0014] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0015] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the subject motion representation extractor training method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0016] In the related art, features are extracted from video frames or clips through methods such as residual networks and dilated 3D convolutional networks. However, these methods are only applicable to third-person videos and cannot capture the motion information of the intelligent body itself. In addition, the feature extraction model for vehicle videos in the related art relies on pseudo-labels, but the pseudo-labels generated by the model are less accurate, which reduces the accuracy of the model. To overcome the above technical problems, this application proposes a subject motion representation extractor training method that can accurately capture the motion information of the intelligent body itself and improve the accuracy of motion information perception.
[0017] The present application embodiment discloses a method for training a subject motion representation extractor, see Figure 1 As shown, the method may include the following steps: Step S11: inputting a training frame into the subject motion representation extractor to be trained, and obtaining predicted six degrees of freedom according to the output; the training frame is a video frame in the first-person perspective video.
[0018] The training frames are images captured from the first-person perspective of the embodied agent. This means that first-person video refers to video content captured by the embodied agent using its own image acquisition device. The trained subject motion representation extractor is capable of capturing the subject motion characteristics of the embodied agent.
[0019] Understandably, third-person perspective scenes focus on capturing detailed information about objects in the video. In embodied intelligent devices (e.g., robots and self-driving cars), onboard cameras record the vehicle's own motion and the evolving surrounding scene as it moves, thereby forming the corresponding video. When understanding a video, it is necessary to infer motion and scene changes based on the video information. Due to the complexity and variability of video scenes, traditional methods struggle to extract effective video representations. For example, while many objects may appear in a driving scene, such as moving vehicles and stationary buildings, these do not affect the vehicle's current motion state. However, traffic lights and other objects play a decisive role in vehicle movement. A vehicle's trajectory is determined by its driving strategy. Only with driving decision-making awareness can one focus more on visual cues that are crucial to driving operations. Therefore, this application proposes a method for training a subject motion representation extractor for first-person perspective videos. The trained subject motion representation extractor is capable of demonstrating driving decision-making awareness.
[0020] The output of the subject motion representation extractor is the predicted six degrees of freedom between the input image frame and the next image frame, where the next image frame is the frame immediately following the input image frame. The trained subject motion representation extractor possesses driving decision-making awareness and can predict changes in the intelligent agent between the current image frame and the next. Six degrees of freedom (6-DoF) refers to the six independent directions of motion a rigid body can have in three-dimensional space, including three translational degrees of freedom (movement along the X, Y, and Z axes) and three rotational degrees of freedom (rotation about the X, Y, and Z axes). This six-DoF allows the intelligent agent to accurately determine its steering changes and positional offsets.
[0021] Inputting a training frame into the subject motion representation extractor to be trained includes: reconstructing a video frame based on the training frame using the subject motion representation extractor to be trained, a post-trained intrinsic parameter detection subnetwork, and a post-trained depth estimation subnetwork to obtain a reconstructed frame; inputting the training frame and the actual next frame of the training frame into the post-trained extrinsic parameter detection subnetwork, and obtaining the standard six degrees of freedom corresponding to the training frame based on the output. In other words, the post-trained extrinsic parameter detection subnetwork, the post-trained intrinsic parameter detection subnetwork, and the post-trained depth estimation subnetwork serve as prior knowledge to guide the subject motion representation extractor in learning how to capture and represent the subject motion characteristics of the embodied intelligent body.
[0022] Step S12: Determine the total loss using the first loss function; the total loss includes the loss between the actual next frame of the training frame and the reconstructed frame, and the loss between the predicted six degrees of freedom and the standard six degrees of freedom corresponding to the training frame; the reconstructed frame is the next frame of the training frame reconstructed based on the predicted six degrees of freedom.
[0023] The first loss function is constructed based on the loss between the actual next frame of the training frame input to the motion representation extractor of the subject to be trained and the reconstructed frame obtained after photometric reconstruction, and the loss between the six degrees of freedom extracted by the motion representation extractor of the subject to be trained based on the training frame and the six degrees of freedom extracted by the external parameter detection subnetwork after training based on the training frame.
[0024] According to the loss between the actual next frame of the training frame and the reconstructed frame, training constraints are performed with frame-level loss. According to the loss between the predicted six degrees of freedom and the standard six degrees of freedom corresponding to the training frame, training constraints are performed with six-degree-of-freedom loss. The loss function is constructed by combining the frame level and six-degree-of-freedom level to improve the training effect and the accuracy of the subject motion representation extractor.
[0025] Specifically, the training steps of the subject motion representation extractor may include: performing photometric reconstruction based on the first-person video using the subject motion representation extractor to be trained, the trained internal parameter detection subnetwork, and the trained depth estimation subnetwork, and training the subject motion representation extractor to be trained using a first loss function; wherein the first loss function is constructed based on the loss between the actual next frame of the training frame input to the subject motion representation extractor to be trained and the reconstructed frame obtained after reconstruction based on the training frame, and the loss between the six degrees of freedom extracted by the subject motion representation extractor to be trained based on the training frame and the six degrees of freedom extracted by the trained external parameter detection subnetwork based on the training frame. That is, the trained external parameter detection subnetwork, the trained internal parameter detection subnetwork, and the trained depth estimation subnetwork are used as prior knowledge to guide the subject motion representation extractor to learn how to capture and represent the subject motion characteristics of the embodied intelligent body.
[0026] Step S13: Adjust the parameters of the subject motion representation extractor to be trained according to the total loss until the training conditions are met to obtain a subject motion representation extractor that can capture the motion characteristics of the embodied intelligent body, so as to use the six degrees of freedom predicted by the subject motion representation extractor to perceive the motion information of the embodied intelligent body.
[0027] Among them, the subject motion representation extractor is trained using first-person video combined with photometric reconstruction technology. The subject motion representation extractor is a neural network model that can capture the motion characteristics of embodied intelligent bodies.
[0028] As can be seen from the above, in this embodiment, the training frame is input into the motion representation extractor of the subject to be trained, and the predicted six degrees of freedom are obtained according to the output; the training frame is a video frame in the first-person perspective video; the total loss is determined using a first loss function; the total loss includes the loss between the actual next frame of the training frame and the reconstructed frame, and the loss between the predicted six degrees of freedom and the standard six degrees of freedom corresponding to the training frame; the reconstructed frame is the next frame of the training frame reconstructed based on the predicted six degrees of freedom; the parameters of the motion representation extractor of the subject to be trained are adjusted according to the total loss until the training conditions are met to obtain a subject motion representation extractor that can capture the motion characteristics of the embodied intelligent body, so as to use the six degrees of freedom predicted by the subject motion representation extractor to perceive the motion information of the embodied intelligent body.
[0029] It can be seen that based on the first-person perspective video combined with the loss between the actual next frame of the training frame and the reconstructed frame of the retraining frame, as well as the loss between the predicted six degrees of freedom and the standard six degrees of freedom corresponding to the training frame, the subject motion representation extractor is trained to enable the subject motion representation extractor to learn whether various detailed information within the field of view has an impact on driving behavior. The trained subject motion representation extractor can perform efficient feature extraction on the first-person perspective video, obtain accurate and effective six degrees of freedom, filter out noise information irrelevant to driving and focus on areas directly related to driving strategies, accurately capture the motion information of the intelligent body in the first-person perspective video, and thus improve the accuracy of motion information perception.
[0030] The training of the subject motion representation extractor relies on the post-training extrinsic parameter detection sub-network, the post-training intrinsic parameter detection sub-network, and the post-training depth estimation sub-network. Therefore, before training the subject motion representation extractor, the steps for generating the post-training extrinsic parameter detection sub-network, the post-training intrinsic parameter detection sub-network, and the post-training depth estimation sub-network may specifically include: S21: Generate training set based on first-person perspective video; Specifically, multiple first-person videos are downloaded from the agent's own perspective. These videos are then preprocessed to create a training set containing multiple video frames. Preprocessing includes frame rate adjustment and frame extraction. These videos are uniformly processed, for example, to 30 frames per second, and frame extraction is performed every two frames to create a training set for training. For example, using the embodied agent as a vehicle, multiple high-definition, unedited first-person driving videos are downloaded from a video website. These videos cover a variety of driving scenarios.
[0031] S22: Performing video frame reconstruction training on the external parameter detection subnetwork, the internal parameter detection subnetwork, and the depth estimation subnetwork based on the training set to obtain a trained external parameter detection subnetwork, a trained internal parameter detection subnetwork, and a trained depth estimation subnetwork.
[0032] Among them, based on the training set, video frame reconstruction training is performed on the external parameter detection subnetwork, the internal parameter detection subnetwork and the depth estimation subnetwork, including: taking every 3 consecutive training frames as a training unit, video frame reconstruction training is performed on the external parameter detection subnetwork, the internal parameter detection subnetwork and the depth estimation subnetwork; the external parameter detection subnetwork is used to determine the camera external parameter change data between two frames, the internal parameter detection subnetwork is used to determine the camera internal parameter, and the depth estimation subnetwork is used to determine the depth map corresponding to the image frame.
[0033] Specifically, we can use self-supervised photometric reconstruction technology to achieve scene reconstruction by standardizing the color consistency between video frames. In this process, we introduce three sub-network modules: external parameter detection sub-network, internal parameter detection sub-network and depth estimation sub-network. The three work together to complete the accurate reconstruction of the scene. Specifically, the external parameter detection sub-network can be The network and internal reference detection sub-network can be The network and deep detection sub-network can be network. PoseNet is a human pose estimation method based on deep learning, and DepthNet is a recurrent neural network architecture for monocular depth prediction. Both are common network models in the industry, and of course other achievable networks can also be used. The function of the depth estimation subnetwork is to calculate the distance from each point in the current scene to the lens, thereby converting the current two-dimensional image into a three-dimensional real scene; the two parameter detection subnetworks focus on estimating the internal and external parameters of the on-board camera, where the internal parameters record the inherent properties of the on-board camera, while the external parameters reveal the angle change and position offset of the camera between two consecutive frames, which can be used to estimate the steering angle and travel speed of the vehicle. This embodiment uses 6-DoF (six degrees of freedom) representation, thereby providing key spatial dynamic information for accurate reconstruction of the scene.
[0034] By collaborating with the external parameter estimation network, the internal parameter estimation network, and the depth estimation sub-network, photometric reconstruction from the current video frame to the next frame is achieved. This process does not rely on external supervision, but instead uses a large amount of unsupervised video to achieve reconstruction through the interaction between networks.
[0035] Specifically, taking every three consecutive training frames as a training unit, the external parameter detection subnetwork, the internal parameter detection subnetwork and the depth estimation subnetwork are trained for video frame reconstruction, including: S221: Input the t-th training frame and the t-1-th training frame into the external reference detection sub-network and the internal reference detection sub-network respectively, and input the t-th training frame into the depth estimation sub-network; For example Figure 2 As shown, three adjacent frames: 、 、 , that is, the training is carried out in units of t-1 frame, t frame and t+1 frame. 、 Input external parameter detection sub-network separately and internal reference detection subnetwork , and get the output of the external parameter detection sub-network: ; Represents the change in camera extrinsic parameters from frame t-1 to frame t, that is, the relative pose between the two frames; The output of the intrinsic parameter detection subnetwork is the predicted camera intrinsic parameter: ; represents the camera intrinsic parameters extracted based on the t-1th frame, represents the camera intrinsic parameters extracted based on the t-th frame; Will Input the depth detection subnetwork to get the depth map of the tth frame : .
[0036] S222: Obtain a first reconstructed frame by photometric reconstruction based on outputs of the extrinsic parameter detection subnetwork, the intrinsic parameter detection subnetwork, and the depth estimation subnetwork, where the first reconstructed frame is a t-th reconstructed frame reconstructed based on the t-1-th training frame; according to To rebuild , the t-th reconstructed frame obtained based on the t-1th training frame is expressed as: ; Among them, K represents the camera internal parameter, Indicates the change value of the camera external parameter from the t-1 frame to the t frame, Represents the depth map of the t-th frame, and the proj() function represents mapping the original pixel space to the reconstructed two-dimensional pixel space, using bilinear interpolation in The pixel space is downsampled to form the first reconstructed frame after reconstruction .
[0037] S223: Input the t-th training frame and the t+1-th training frame into the external parameter detection sub-network and the internal parameter detection sub-network respectively, and input the t-th training frame into the depth estimation sub-network; Will 、 Input external parameter detection sub-network separately and internal reference detection subnetwork , according to the output of the external parameter detection sub-network: ; Represents the change value of the camera extrinsic parameter predicted from frame t+1 to frame t; The output of the intrinsic parameter detection subnetwork is the predicted camera intrinsic parameter: ; represents the camera intrinsic parameters extracted based on the t+1th frame, represents the camera intrinsic parameters extracted based on the t-th frame; Similarly, Input the depth detection subnetwork to get the depth map of the tth frame : .
[0038] S224: Obtain a second reconstructed frame by photometric reconstruction based on outputs of the extrinsic parameter detection subnetwork, the intrinsic parameter detection subnetwork, and the depth estimation subnetwork, where the second reconstructed frame is a t-th reconstructed frame reconstructed based on the t+1-th training frame; according to To rebuild , the t-th reconstructed frame obtained based on the t+1-th training frame is expressed as: ; Among them, K represents the camera internal parameter, Indicates the change value of the camera extrinsic parameter from the t+1th frame to the tth frame, Represents the depth map of the t-th frame, and the proj() function represents mapping the original pixel space to the reconstructed two-dimensional pixel space, and then using bilinear interpolation to Pixel space downsampling forms the second reconstructed frame after reconstruction .
[0039] S225: Perform photometric reconstruction training on the extrinsic parameter detection subnetwork, the intrinsic parameter detection subnetwork, and the depth estimation subnetwork using the second loss function; the second loss function is constructed based on the difference between the tth training frame and the first reconstructed frame, and the difference between the tth training frame and the second reconstructed frame.
[0040] In a specific implementation, the second loss function is constructed based on the structural similarity metric (SSIM) and L1 loss between the t-th training frame and the first reconstructed frame, the structural similarity metric and L1 loss between the t-th training frame and the second reconstructed frame, and the parallax smoothness loss of the t-th training frame. The above-mentioned second loss function can be specifically constructed as follows: ; in, and are all preset weight hyperparameters, is the frame reconstruction loss, represents parallax smoothness loss; It consists of the structural similarity metric SSIM and the L1 loss term: ; in, is the coefficient; ; in, Represents the gradient in the horizontal direction, Represents the gradient in the vertical direction, Represents the mean normalized inverse depth map.
[0041] The Structural Similarity (SSIM) metric between two frames captures the similarity of images or features in terms of structure, brightness, and contrast. By utilizing the structural similarity metric loss, more natural and visually plausible motion estimates can be generated. This is particularly suitable for tasks that require maintaining texture and edge continuity, providing constraints on motion or change that are more consistent with human perception. In vehicle motion estimation, combining SSIM can improve robustness to complex scenes (such as shadows and dynamic objects). The L1 loss provides pixel-level photometric consistency constraints and is robust to noise. L1 provides low-order pixel accuracy, while SSIM provides high-order structural perception. Combined with the parallax smoothness loss, the motion field is forced to conform to physical laws, avoiding cluttered estimation results and improving the model's robustness and visual quality.
[0042] After the first stage of training, three sub-networks are obtained: the post-training extrinsic parameter detection sub-network, the post-training intrinsic parameter detection sub-network, and the post-training depth estimation sub-network. Among them, the post-training extrinsic parameter detection sub-network is closest to the requirement, but it essentially captures the relative motion difference between two adjacent frames. In fact, what is needed is the difference information on a certain frame. This difference is manifested in how driving should be performed to form the ideal next frame. This means that it is necessary to simulate the driving / action decision that should be made when encountering a certain observation frame scene, that is, the network needs to learn driving / action strategy information. When entering the second stage, that is, for the training of the subject motion representation extractor to be trained, the post-training extrinsic parameter detection sub-network, the post-training intrinsic parameter detection sub-network, and the post-training depth estimation sub-network obtained from the previous training are frozen, that is, the model parameters of these sub-networks remain unchanged in the second stage training, and they no longer participate in training. The subject motion representation extractor to be trained can specifically be a posture detection network ,This network is also used to extract six degrees of freedom in order to estimate the camera extrinsic parameters.
[0043] In a specific implementation, the training frame is input into the subject motion representation extractor to be trained, including: S31: reconstructing a video frame based on the training frame using the subject motion representation extractor to be trained, the trained intrinsic reference detection subnetwork, and the trained depth estimation subnetwork to obtain a reconstructed frame; The training frame can be any frame in the training set generated based on the first-person perspective video.
[0044] S32: Inputting the training frame and the actual next frame of the training frame into the trained external parameter detection sub-network, and obtaining the standard six degrees of freedom corresponding to the training frame according to the output; Specifically, the subject motion representation extractor training process includes: ) are respectively input to the subject motion representation extractor to be trained, the internal reference detection subnetwork after training, and the depth estimation subnetwork after training; according to the outputs of the subject motion representation extractor to be trained, the internal reference detection subnetwork after training, and the depth estimation subnetwork after training, the p-th reconstructed frame (i.e. ); that is, the p-th reconstructed frame is obtained based on the p-1th training frame; the first loss function is used to train the subject motion representation extractor; wherein, during the training process, the network parameters of the trained internal reference detection subnetwork and the trained depth estimation subnetwork are frozen and unchanged. For example Figure 3 As shown, the training process of the subject motion representation extractor is to be trained, and only input , reconstructed to At the same time, the p-1th training frame and the pth training frame ( ) Input the trained external parameter detection sub-network and obtain the standard six degrees of freedom of the p-1th training frame based on the output.
[0045] Specifically, the first loss function includes a first sub-loss function and a second sub-loss function; the first sub-loss function is constructed based on the structural similarity measure and L1 loss between the p-th training frame and the p-th reconstructed frame, and the parallax smoothness loss of the p-th training frame; that is, based on and The first sub-loss function Lc is generated by the structural similarity measure and L1 loss between them; the second sub-loss function is constructed based on the cross entropy loss between the predicted six degrees of freedom output by the motion representation extractor of the subject to be trained according to the p-1 training frame and the standard six degrees of freedom output by the external parameter detection sub-network after training according to the p-1 training frame and the p training frame.
[0046] Input the p-1th training frame and the pth training frame into the trained external parameter detection network, and use their output as the true label for training Since the external parameter detection network processes the motion information changes between two adjacent frames after training, When we set the parameters of , we are actually learning to obtain the correct motion strategy information by observing the p-1th training frame to generate the ideal pth reconstruction frame. It must be able to perceive rich motion information and obtain the vehicle subject motion representation extractor after training.
[0047] When estimating the motion state, only input , and use the subject motion representation extractor Extract the corresponding motion status information , when calculating the loss, the frozen external parameters of the detection network are used Calculated relative pose , and get the second sub-loss function Lm: ; Finally, Lc+Lm is used as the first loss function to train the subject motion representation extractor parameters.
[0048] In a specific implementation, the training process of the subject motion representation extractor includes: S41: inputting a training frame into a motion representation extractor of a subject to be trained, and obtaining predicted six degrees of freedom according to the output; the training frame is a video frame in a first-person perspective video; S42: Determine a total loss using a first loss function; the total loss includes a loss between an actual next frame of the training frame and a reconstructed frame, and a loss between the predicted six degrees of freedom and a standard six degrees of freedom corresponding to the training frame; the reconstructed frame is a next frame of the training frame reconstructed based on the predicted six degrees of freedom; S43: adjusting parameters of the subject motion representation extractor to be trained according to the total loss until a training condition is met to obtain a subject motion representation extractor capable of capturing motion characteristics of the embodied intelligent body; S44: Obtain an initial trained subject motion representation extractor, and optimize the initial trained subject motion representation extractor through visual odometry estimation to obtain a final trained subject motion representation extractor.
[0049] That is, after the subject motion representation extractor to be trained is trained using the first loss function, the subject motion representation extractor is further optimized. It can be understood that after the first two stages, an initially trained vehicle subject motion representation extractor is obtained, but since all unsupervised training is used, only primary motion features can be extracted. Therefore, supervised training using a large amount of open annotated data is continued, and the feature extractor is further refined using the visual odometer estimation task. Visual odometer estimation aims to estimate the changes in the moving subject captured by the camera between adjacent frames. For example, the public KITTI odometry dataset is used, which is recorded at 10 frames per second by a visual system installed on a vehicle, capturing street images during driving. These data are often used for visual odometer estimation; in order to maintain the use of monocular vision, only the video data captured from the left camera can be used.
[0050] By using publicly available annotated datasets, we further optimize and adjust the pre-trained feature extractor, enabling it to better adapt to diverse driving scenarios and improving its generalization and accuracy in practical applications.
[0051] Specifically, the initially trained subject motion representation extractor is optimized through visual odometry estimation, including: inputting the odometry dataset into the initially trained subject motion representation extractor, wherein the output of the initially trained subject motion representation extractor is the input of an encoder; obtaining encoding features corresponding to each image frame according to the output of the encoder, inputting the encoding features into a multi-layer perceptron to obtain motion difference information; and optimizing the training of the initially trained subject motion representation extractor using a third loss function; the third loss function is constructed based on the encoding feature difference between two adjacent frames and the motion difference information between the two frames.
[0052] For example Figure 4As shown, a basic Transformer encoder structure is used to further encode the features extracted by the vehicle subject motion representation extractor. The features are then used through a multi-layer perceptron to estimate the difference between two adjacent frames. This difference can be represented using six degrees of freedom (DOF). Six degrees of freedom, as a specific form of camera extrinsic parameter, captures the ability of a three-dimensional object to autonomously move along three dimensions and rotate at three angles in space from a first-person perspective. It can be understood that the output of the vehicle subject motion representation extractor is a low-level motion representation, while the encoder output is a higher-level temporal motion representation. The motion difference information obtained by the multi-layer perceptron enables the model to further learn how to accurately predict motion changes between adjacent frames from the features.
[0053] For example, at time t, the car's motion state matrix is for: ; in Represents the rotation matrix, describing the steering angle of the car; Represents the translation vector, then the motion difference information between time t-1 and time t Expressed as: ; For a video containing N consecutive frames, N-1 relative motion difference information will be obtained after calculation. Input N video frames, and then the encoder will get N features , calculate the difference between adjacent features to get The third loss function constructed based on the difference in encoding features between two adjacent frames and the motion difference information between the two frames is: ; Among them, CrossEntropy represents cross entropy loss, and Encoder represents encoding.
[0054] The subject motion representation extractor obtained through the above training method can focus on the subject motion information and the areas that affect driving strategies, so that the extracted features contain rich motion information, and realize the capture and representation of generalized vehicle subject motion features, avoiding the influence of details of other objects in the video on perception.
[0055] Moreover, in the related art, the feature extraction model of the first-person video of the vehicle relies on the pseudo-label method, and there is a problem that the pseudo-label generated by the model has poor accuracy. The present application is based on a subject motion representation extractor of self-supervised photometric reconstruction, which realizes the training of the video feature extractor through self-supervision without relying on manually labeled data. In the first two stages, a subject motion feature extraction model that can accurately capture the area that has a substantial impact on the driving strategy was constructed through photometric reconstruction technology. Then, in the third stage, the model was carefully adjusted and optimized using labeled data to ensure that the model can learn the appropriate driving decisions to be taken in specific scenarios. Finally, a vehicle subject motion representation extractor that conforms to the driving strategy was obtained.
[0056] The present application discloses a specific method for perceiving motion information of an embodied intelligent body, which may include the following steps: Step S51: obtaining a target image frame of the first perspective currently captured by the embodied intelligent body; In this embodiment, the target video frame is the one currently captured by the embodied agent, and is captured from the embodied agent's first-person perspective. This refers to video content captured by the embodied agent using its own image capture device, with the embodied agent as the primary perspective. Step S52: inputting the target image frame into the aforementioned subject motion representation extractor to obtain the six degrees of freedom between the target image frame and the next image frame predicted by the subject motion representation extractor; In this embodiment, the target image frame is input into a pre-trained subject motion representation extractor. The output of the subject motion representation extractor is the predicted six degrees of freedom between the target image frame and the next image frame. The next image frame is the next image frame of the target image frame. The trained subject motion representation extractor has driving decision-making awareness and can predict the changes of the intelligent body between the current image frame and the next image frame.
[0057] Step S53: Perceiving motion information of the embodied intelligent body based on the six degrees of freedom. Finally, motion information perception of the matrix intelligent body is realized based on the six degrees of freedom output by the subject motion representation extractor. Motion information perception refers to understanding the dynamic changes of objects or scenes by analyzing the motion characteristics in the video, but the motion information perception of the intelligent body is a complex process. In this process, not only the six degrees of freedom will be used, but also multiple perception methods and multimodal data fusion technology will be combined to achieve more comprehensive and accurate environmental perception and motion control. This embodiment does not limit this, and any method that can achieve motion perception in the relevant technology can be used. The specific process of the above-mentioned subject motion representation extractor training method can refer to the corresponding content disclosed in the above-mentioned embodiment, and will not be repeated here.
[0058] As can be seen from the above, in this embodiment, a target image frame of the first perspective currently captured by the embodied intelligent body is obtained; the target image frame is input into the aforementioned subject motion representation extractor to obtain the six degrees of freedom between the target image frame and the next image frame predicted by the subject motion representation extractor; and the motion information perception of the embodied intelligent body is performed based on the six degrees of freedom.
[0059] It can be seen that based on the first-person perspective video combined with the loss between the actual next frame of the training frame and the reconstructed frame of the retraining frame, as well as the loss between the predicted six degrees of freedom and the standard six degrees of freedom corresponding to the training frame, the subject motion representation extractor is trained to enable the subject motion representation extractor to learn whether various detailed information within the field of view has an impact on driving behavior. The trained subject motion representation extractor can perform efficient feature extraction on the first-person perspective video, obtain accurate and effective six degrees of freedom, filter out noise information irrelevant to driving and focus on areas directly related to driving strategies, accurately capture the motion information of the intelligent body in the first-person perspective video, and thus improve the accuracy of motion information perception.
[0060] Accordingly, an embodiment of the present application further discloses a subject motion representation extractor training device, the device comprising: An input module is used to input a training frame into a motion representation extractor of a subject to be trained, and obtain a predicted six degrees of freedom according to the output; the training frame is a video frame in a first-person perspective video; a loss determination module, configured to determine a total loss using a first loss function; the total loss comprising a loss between an actual next frame of the training frame and a reconstructed frame, and a loss between the predicted six degrees of freedom and a standard six degrees of freedom corresponding to the training frame; the reconstructed frame being a next frame of the training frame reconstructed based on the predicted six degrees of freedom; A parameter adjustment module is used to adjust the parameters of the subject motion representation extractor to be trained according to the total loss until the training conditions are met to obtain a subject motion representation extractor that can capture the motion characteristics of the embodied intelligent body, so as to use the six degrees of freedom predicted by the subject motion representation extractor to perceive the motion information of the embodied intelligent body.
[0061] As can be seen from the above, in this embodiment, the training frame is input into the motion representation extractor of the subject to be trained, and the predicted six degrees of freedom are obtained according to the output; the training frame is a video frame in the first-person perspective video; the total loss is determined using a first loss function; the total loss includes the loss between the actual next frame of the training frame and the reconstructed frame, and the loss between the predicted six degrees of freedom and the standard six degrees of freedom corresponding to the training frame; the reconstructed frame is the next frame of the training frame reconstructed based on the predicted six degrees of freedom; the parameters of the motion representation extractor of the subject to be trained are adjusted according to the total loss until the training conditions are met to obtain a subject motion representation extractor that can capture the motion characteristics of the embodied intelligent body, so as to use the six degrees of freedom predicted by the subject motion representation extractor to perceive the motion information of the embodied intelligent body. It can be seen that based on the first-person perspective video combined with the loss between the actual next frame of the training frame and the reconstructed frame of the retraining frame, as well as the loss between the predicted six degrees of freedom and the standard six degrees of freedom corresponding to the training frame, the subject motion representation extractor is trained to enable the subject motion representation extractor to learn whether various detailed information within the field of view has an impact on driving behavior. The trained subject motion representation extractor can perform efficient feature extraction on the first-person perspective video, obtain accurate and effective six degrees of freedom, filter out noise information irrelevant to driving and focus on areas directly related to driving strategies, accurately capture the motion information of the intelligent body in the first-person perspective video, and thus improve the accuracy of motion information perception.
[0062] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0063] An embodiment of the present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned subject motion representation extractor training methods or embodied intelligent body motion information perception method embodiments.
[0064] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned subject motion representation extractor training methods or embodied intelligent body motion information perception method embodiments when running.
[0065] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0066] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned subject motion representation extractor training methods or embodied intelligent body motion information perception method embodiments.
[0067] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned subject motion representation extractor training methods or embodied intelligent body motion information perception method embodiments.
[0068] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0069] The above is a detailed introduction to a subject motion representation extractor training method or an embodied intelligent body motion information perception method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.
Claims
1. A method for training a subject motion representation extractor, characterized in that: include: Inputting a training frame into a motion representation extractor of a subject to be trained, and obtaining a predicted six degrees of freedom according to the output; the training frame is a video frame in a first-person perspective video; Determining a total loss using a first loss function; the total loss includes a loss between an actual next frame of the training frame and a reconstructed frame, and a loss between the predicted six degrees of freedom and a standard six degrees of freedom corresponding to the training frame; the reconstructed frame is a next frame of the training frame reconstructed based on the predicted six degrees of freedom; The parameters of the subject motion representation extractor to be trained are adjusted according to the total loss until the training conditions are met to obtain a subject motion representation extractor that can capture the motion characteristics of the embodied intelligent body, so as to use the six degrees of freedom predicted by the subject motion representation extractor to perceive the motion information of the embodied intelligent body.
2. The subject motion representation extractor training method according to claim 1, characterized in that: Input the training frame into the subject motion representation extractor to be trained, including: Reconstructing the video frame based on the training frame using the subject motion representation extractor to be trained, the trained internal reference detection subnetwork, and the trained depth estimation subnetwork to obtain the reconstructed frame; The training frame and the actual next frame of the training frame are input into the trained external parameter detection subnetwork, and the standard six degrees of freedom corresponding to the training frame are obtained according to the output.
3. The subject motion representation extractor training method according to claim 2, characterized in that: Reconstructing video frames based on the training frames using a subject motion representation extractor to be trained, a trained intrinsic reference detection subnetwork, and a trained depth estimation subnetwork, including: Input the p-1th training frame into the subject motion representation extractor to be trained, the trained internal reference detection subnetwork and the trained depth estimation subnetwork respectively; Obtaining a p-th reconstructed frame by photometric reconstruction based on outputs of the subject motion representation extractor to be trained, the trained intrinsic reference detection subnetwork, and the trained depth estimation subnetwork; During the training of the subject motion representation extractor, the network parameters of the trained internal reference detection subnetwork and the trained depth estimation subnetwork remain frozen.
4. The subject motion representation extractor training method according to claim 3, characterized in that: The first loss function includes a first sub-loss function and a second sub-loss function; The first sub-loss function is constructed based on a structural similarity measure and an L1 loss between the p-th training frame and the p-th reconstructed frame, and a disparity smoothness loss of the p-th training frame; The second sub-loss function is constructed based on the cross entropy loss between the predicted six degrees of freedom output by the subject motion representation extractor to be trained based on the p-1th training frame and the standard six degrees of freedom output by the trained external parameter detection sub-network based on the p-1th training frame and the pth training frame.
5. The subject motion representation extractor training method according to claim 2, characterized in that: Before the training frame is input into the subject motion representation extractor to be trained, it also includes: Generate a training set based on first-person perspective videos; Based on the training set, video frame reconstruction training is performed on the external parameter detection subnetwork, the internal parameter detection subnetwork and the depth estimation subnetwork to obtain the trained external parameter detection subnetwork, the trained internal parameter detection subnetwork and the trained depth estimation subnetwork.
6. The subject motion representation extractor training method according to claim 5, characterized in that: Generate a training set based on the first-person perspective video, including: The first perspective video is preprocessed to obtain a training set containing multiple video frames based on the processed video; the preprocessing includes frame rate adjustment and frame extraction.
7. The subject motion representation extractor training method according to claim 5, characterized in that: Video frame reconstruction training is performed on the external parameter detection sub-network, the internal parameter detection sub-network, and the depth estimation sub-network based on the training set, including: The extrinsic parameter detection subnetwork, the intrinsic parameter detection subnetwork, and the depth estimation subnetwork are trained for video frame reconstruction with every three consecutive training frames as a training unit; the extrinsic parameter detection subnetwork is used to determine the camera extrinsic parameter change data between two frames, the intrinsic parameter detection subnetwork is used to determine the camera intrinsic parameter, and the depth estimation subnetwork is used to determine the depth map corresponding to the image frame.
8. The subject motion representation extractor training method according to claim 7, characterized in that: Taking every three consecutive training frames as a training unit, the extrinsic parameter detection subnetwork, the intrinsic parameter detection subnetwork, and the depth estimation subnetwork are trained for video frame reconstruction, including: Input the t-th training frame and the t-1-th training frame into the external parameter detection subnetwork and the internal parameter detection subnetwork respectively, and input the t-th training frame into the depth estimation subnetwork; Obtaining a first reconstructed frame by photometric reconstruction according to outputs of the extrinsic parameter detection subnetwork, the intrinsic parameter detection subnetwork, and the depth estimation subnetwork, where the first reconstructed frame is a t-th reconstructed frame reconstructed based on the t-1-th training frame; Input the tth training frame and the t+1th training frame into the external parameter detection subnetwork and the internal parameter detection subnetwork respectively, and input the tth training frame into the depth estimation subnetwork; Obtaining a second reconstructed frame by photometric reconstruction based on outputs of the extrinsic parameter detection subnetwork, the intrinsic parameter detection subnetwork, and the depth estimation subnetwork, where the second reconstructed frame is a t-th reconstructed frame reconstructed based on the t+1-th training frame; The extrinsic parameter detection subnetwork, the intrinsic parameter detection subnetwork, and the depth estimation subnetwork are trained for photometric reconstruction using a second loss function; the second loss function is constructed based on the difference between the tth training frame and the first reconstructed frame, and the difference between the tth training frame and the second reconstructed frame.
9. The subject motion representation extractor training method according to claim 8, characterized in that: The second loss function is constructed based on the structural similarity measure and L1 loss between the tth training frame and the first reconstructed frame, the structural similarity measure and L1 loss between the tth training frame and the second reconstructed frame, and the disparity smoothness loss of the tth training frame.
10. The subject motion representation extractor training method according to any one of claims 1 to 9, characterized in that: After adjusting the parameters of the subject motion representation extractor to be trained according to the total loss, the method further includes: An initial trained subject motion representation extractor is obtained, and the initial trained subject motion representation extractor is optimized by visual odometry estimation to obtain a final subject motion representation extractor.
11. The subject motion representation extractor training method according to claim 10, characterized in that: Optimizing the initially trained subject motion representation extractor via visual odometry estimation includes: Inputting the odometry dataset into the initially trained subject motion representation extractor, wherein the output of the initially trained subject motion representation extractor is the input of the encoder; Obtaining encoding features corresponding to each image frame according to the output of the encoder, and inputting the encoding features into a multi-layer perceptron to obtain motion difference information; The subject motion representation extractor after the initial training is optimized and trained using a third loss function; the third loss function is constructed based on the encoding feature difference between two adjacent frames and the motion difference information between the two frames.
12. A method for perceiving motion information of an embodied intelligent body, characterized in that: include: Obtain the target image frame from the first perspective currently captured by the embodied intelligent body; Inputting the target image frame into a subject motion representation extractor to obtain six degrees of freedom between the target image frame and a next image frame predicted by the subject motion representation extractor; the subject motion representation extractor is trained by the subject motion representation extractor training method according to any one of claims 1 to 11; The motion information of the embodied intelligent body is perceived based on the six degrees of freedom.
13. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the subject motion representation extractor training method according to any one of claims 1 to 11, or the subject motion representation extractor training method according to claim 12.
14. A computer-readable storage medium, characterized in that Used to store a computer program; wherein when the computer program is executed by a processor, the subject motion representation extractor training method according to any one of claims 1 to 11 or the subject motion representation extractor training method according to claim 12 is implemented.
15. A computer program product, characterized in that The method comprises a computer program, which, when executed by a processor, implements the subject motion representation extractor training method according to any one of claims 1 to 11, or the subject motion representation extractor training method according to claim 12.
Citation Information
Patent Citations
Depth estimation model training method and device, depth estimation method and device, equipment and medium
CN115578704A
Lightweight monocular depth estimation method and device based on self-supervised deep learning
CN119494866A