Training data set generation method and apparatus, electronic device, and storage medium

By obtaining the original videos taken by multiple cameras from different perspectives, determining the moving trajectory and background of the target, and generating a training data set using the self-training method, the problem of lack of multi-objective pose estimation data in social behavior research is solved, and efficient multi-objective pose estimation is achieved.

WO2025091167A1PCT designated stage expired Publication Date: 2025-05-08SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Patent Information

Application Number
PCT/CN2023/127833
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

In social behavior research, traditional three-box social experiments and image processing techniques are difficult to effectively conduct multi-objective pose estimation, resulting in the problem of data scarcity.

Method used

By acquiring the original videos taken by multiple cameras from different perspectives, determining the moving trajectory and background of the target, using a self-training method to generate a training data set of multi-objective pose estimation, and fusing outlines and motion backgrounds to form the data set.

Benefits of technology

The problem of lack of data for multi-objective pose estimation is solved. The generated training data set can reflect different occlusion relationships, avoid a large number of manual annotations, and improve data volume and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023127833_08052025_PF_FP_ABST
    Figure CN2023127833_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, and provides a training data set generation method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring a raw video; on the basis of the raw video, determining movement trajectories and movement backgrounds of objects in a set scene, and on the basis of a first number of video frames in the raw video, determining the contours of the objects in the first number of video frames; on the basis of the contours of the objects in the first number of video frames, performing self-training on a second number of video frames in the raw video to obtain the contours of the objects in all the video frames; and on the basis of the movement trajectories of the objects in the set scene, fusing the contours and the movement backgrounds of the objects in the corresponding video frames to form a training data set, wherein training data in the training data set is video frames in which objects having different occlusion relationships all have undergone instance labeling. The present application solves the problem in the prior art that data for multi-object pose estimation is deficient.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, electronic device and storage medium for generating training data set Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to a method, device, electronic device, and storage medium for generating a training data set. Background Art

[0002] Currently, research on single-target social behavior relies primarily on data from three-box social experiments. For example, using actions as the target, three-box social experiments can be used to test social preferences for actions. However, these experiments oversimplify the quantification of social behavior, resulting in limited data and crude behavioral quantification. This not only fails to provide effective reference for later clinical trials, but also wastes significant drug development resources.

[0003] In order to make up for the shortcomings of the three-box social experiment, research on multi-target social behavior was proposed. However, traditional image processing and machine learning are still unable to accurately track and estimate the posture of multi-target social behavior, resulting in a very limited increase in the amount of data.

[0004] From the above, we can see that there is still a problem of lack of data for multi-target pose estimation in social behavior research.

[0005] Summary of the Invention

[0006] This application provides a method, device, electronic device, and storage medium for generating a training dataset, which can solve the problem of data scarcity for multi-target pose estimation in related technologies. The technical solution is as follows:

[0007] According to one aspect of the present application, a method for generating a training data set, wherein the training data set is used for multi-target posture estimation, the method comprising: obtaining an original video; each video frame in the original video is shot and collected by multiple cameras arranged at different perspectives for multiple targets performing social behaviors in a set scene; based on the original video, determining the motion trajectory and motion background of each of the targets in the set scene, and based on a first number of video frames in the original video, determining the outline of each of the targets in a first number of video frames; based on the outline of each of the targets in the first number of video frames, self-training is performed on a second number of video frames in the original video to obtain the outline of each of the targets in all video frames; based on the motion trajectory of each of the targets in the set scene, the outline and motion background of each of the targets in the corresponding video frames are fused to form the training data set; each training data in the training data set is a video frame with instance annotations for each of the targets with different occlusion relationships.

[0008] According to one aspect of the present application, a device for generating a training data set, wherein the training data set is used for multi-target posture estimation, and the device includes: a video acquisition module, used to acquire an original video; each video frame in the original video is shot and collected by multiple cameras arranged at different perspectives for multiple targets performing social behaviors in a set scene; a contour determination module, used to determine the motion trajectory and motion background of each target in the set scene based on the original video, and determine the contour of each target in a first number of video frames based on a first number of video frames in the original video; a self-training module, used to self-train each second number of video frames in the original video based on the contour of each target in the first number of video frames, to obtain the contour of each target in all video frames; a data fusion module, used to fuse the contour and motion background of each target in the corresponding video frame based on the motion trajectory of each target in the set scene to form the training data set; each training data in the training data set is a video frame with instance annotation of each target with different occlusion relationships.

[0009] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the method for generating a training data set as described above.

[0010] According to one aspect of the present application, a storage medium stores computer-readable instructions thereon, and the computer-readable instructions are executed by one or more processors to implement the method for generating a training dataset as described above.

[0011] According to one aspect of the present application, a computer program product includes computer-readable instructions, which are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the method for generating a training data set as described above.

[0012] The beneficial effects of the technical solution provided by this application are:

[0013] In the above technical solution, multiple cameras arranged at different perspectives are used to capture and capture multiple targets engaging in social activities in a set scene, obtaining video frames in the original video. On the one hand, the motion trajectory and motion background of each target in the set scene are determined based on each video frame of the original video. On the other hand, the outline of each target in a first number of video frames in the original video is determined based on a first number of video frames in the original video. Then, based on the outline of each target in the first number of video frames, a second number of video frames in the original video are self-trained to obtain the outline of each target in all video frames. Finally, guided by the motion trajectory of the target in the set scene, the outline of each target in the corresponding video frames and the motion background are fused to form a training dataset for multi-target pose estimation. Each training data in the training dataset can reflect the different occlusion relationships between the targets. In other words, the training data for multi-target pose estimation no longer relies primarily on manual annotation, but can be created through self-training, thereby enriching the training dataset and effectively solving the problem of data scarcity for multi-target pose estimation in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts.

[0015] FIG1 is a schematic diagram of an implementation environment according to the present application;

[0016] FIG2 is a flow chart showing a method for generating a training data set according to an exemplary embodiment;

[0017] FIG2a is a schematic diagram of a video frame with contour marking according to the embodiment of FIG2 ;

[0018] FIG2 b is a schematic diagram of a video frame carrying a contour label according to the embodiment of FIG2 ;

[0019] FIG2c is a schematic diagram of training data with instance annotations of multiple targets with occlusion relationships involved in the embodiment of FIG2;

[0020] FIG3 is a flow chart of step 350 in the embodiment corresponding to FIG2 in one embodiment;

[0021] FIG4 is a flowchart of a method for generating a training data set according to an exemplary embodiment;

[0022] FIG4a is a schematic diagram of a video frame with gesture marking according to the embodiment of FIG4 ;

[0023] FIG5 is a flow chart of step 450 in one embodiment of the embodiment corresponding to FIG4 ;

[0024] FIG5 a is a schematic diagram of a multi-target three-dimensional posture estimation result involved in the embodiment of FIG5 ;

[0025] FIG6 is a schematic diagram of a specific implementation of a method for generating a training data set in an application scenario;

[0026] FIG7 is a structural block diagram of a device for generating a training data set according to an exemplary embodiment;

[0027] FIG8 is a hardware structure diagram of a server according to an exemplary embodiment;

[0028] Fig. 9 is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0029] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.

[0030] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0031] As mentioned above, the target social behavior study is mainly based on the three-box social experiment. The three-box social experiment is aimed at the social nature of animals. The experimental box designed includes three parts: left, middle and right. The animal to be socialized is fixed in the experimental box on the left or right side, and the test animal that needs to undergo social behavior research is placed in the middle experimental box. The test animal can move freely and socialize freely in the experimental box. Based on this, the social behavior of the test animal can be analyzed based on the time the test animal stays in the left box, the time it stays in the right box, and the number of times it enters and exits the left and right boxes. Since the social preferences shown by the test animals in the three-box social experiment are actually related to many factors, such as the familiarity of the test animals with the experimental box, the athletic ability of the test animals, and the social status of the test animals, etc., this makes the three-box social experiment too simplistic in considering the social preferences of the test actions, making it too simple to quantify social behavior, and thus making the small amount of data generated unable to provide an effective reference for later clinical trials, and will waste a lot of drug research and development resources.

[0032] To address the data scarcity associated with the three-chamber social experiment, a method for directly observing the social behavior of the subjects was proposed. This method allows multiple animals to move freely and socialize within the experimental chamber, thereby addressing the limitations of the three-chamber social experiment. However, animals are non-rigid bodies, and particularly when engaging in social behavior, they can interact or come into close contact with one another. Traditional image processing and machine learning methods are unable to effectively identify and distinguish multiple animals in close interaction or contact, and still rely heavily on manual visual correction, resulting in limited data enrichment.

[0033] From the above, we can see that the related technology still has the defect of lack of training data for multi-target pose estimation.

[0034] To this end, the method for generating a training data set provided in this application can effectively enrich the training data used for multi-target posture estimation. Accordingly, the method for generating a training data set is suitable for a device for generating a training data set, and the device for generating a training data set can be deployed in an electronic device, which can be a computer device configured with a von Neumann architecture, for example, the computer device includes a desktop computer, a laptop computer, a server, etc.

[0035] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0036] Figure 1 is a schematic diagram of an implementation environment involved in a method for generating a training data set. It should be noted that this implementation environment is only an example adapted to the present invention and should not be considered as providing any limitation on the scope of application of the present invention.

[0037] In FIG1 , the implementation environment includes a collection end 110 and a service end 130 .

[0038] Specifically, the acquisition terminal 110 can also be considered as an image acquisition device, including but not limited to electronic devices with shooting functions such as cameras, cameras, and camcorders. In this application, the acquisition terminal 110 includes multiple cameras arranged at different viewing angles, so that the acquisition terminal 110 can capture and capture multiple targets engaging in social activities in a set scene from different viewing angles.

[0039] Server 130 can be an electronic device such as a desktop computer, laptop computer, or server. It can also be a server cluster composed of multiple servers, or even a cloud computing center composed of multiple servers. Server 130 is used to provide background services, such as, but not limited to, training dataset generation services and multi-target pose estimation services.

[0040] A network communication connection is pre-established between the server 130 and the acquisition terminal 110 via a wired or wireless method, and data is transmitted between the server 130 and the acquisition terminal 110 via the network communication connection. The transmitted data includes, but is not limited to, individual video frames in the original video, individual videos to be processed in the video to be processed, and the like.

[0041] In one application scenario, through the interaction between the acquisition terminal 110 and the server terminal 130, the acquisition terminal 110 shoots and collects original videos of multiple targets performing social behaviors in a set scene, and uploads each video frame in the original video to the server terminal 130 to request the server terminal 130 to provide a training data set generation service.

[0042] For the server 130, after receiving the original video uploaded by the acquisition end 110, it calls the training dataset service, combines the annotation of a small number of video frames and the self-training of most video frames to form a training dataset, so as to solve the problem of lack of training data for multi-target pose estimation in related technologies.

[0043] In another application scenario, through the interaction between the acquisition end 110 and the server end 130, the acquisition end 110 uses cameras with multiple different perspectives to shoot and collect multiple targets in any natural scene to obtain a video to be processed, and uploads each video frame to be processed in the video to be processed to the server end 130 to request the server end 130 to provide a multi-target posture estimation service.

[0044] After receiving the video to be processed from the acquisition terminal 110, the server 130 can use the trained spatiotemporal instance segmentation model and the trained single-view object pose estimation model to provide multi-target pose estimation services for multiple targets in the video to be processed. The spatiotemporal instance segmentation model and the single-view object pose estimation model are trained based on the training data in the training dataset. Given a rich training dataset, the trained models can be generalized and effectively identify and distinguish multiple animals that are closely interacting or in contact in various natural scenes.

[0045] Please refer to FIG. 2 . An embodiment of the present application provides a method for generating a training data set. The method is applicable to an electronic device, which may be the server 130 in the implementation environment shown in FIG. 1 .

[0046] In the following method embodiments, for ease of description, the execution subject of each step of the method is taken as an electronic device as an example for illustration, but this does not constitute a specific limitation.

[0047] As shown in FIG2 , the method may include the following steps:

[0048] Step 310: Obtain the original video.

[0049] Each frame in the original video is captured by multiple cameras positioned at different viewing angles, targeting multiple targets engaging in social behaviors within a set scene. The targets can be any object with social abilities, such as humans, animals (e.g., mice, birds, dogs), etc. Accordingly, the set scene refers to any setting that can accommodate multiple targets engaging in social behaviors, such as an experimental setting, such as an experimental box, or any natural setting, such as a park or a room.

[0050] Video frames are captured and collected using multiple cameras at different viewing angles, targeting multiple subjects engaging in social interactions in a set scene. These cameras can be deployed around the setting. For example, if the subjects are people, multiple cameras can be deployed on different pillars in a room where they are located; if the subjects are animals, multiple cameras can be deployed on multiple lampposts in a park where the animals are located.

[0051] It is understood that the shooting can be a single multiple-shot or a continuous shooting. Then, for multiple targets engaging in social behavior in a set scene, for continuous shooting, a video can be obtained, while for multiple shooting, multiple photos can be obtained. In other words, the video frame in this embodiment comes from a dynamic image, such as any frame in a video as a video frame. Of course, in other embodiments, the video frame can also come from a static image, such as any one of multiple photos as a video frame. This does not constitute a specific limitation. Accordingly, the multi-target posture estimation in this embodiment can be performed in frames.

[0052] Regarding the acquisition of the original video, the original video can be sourced from real-time captured and collected video, or it can be pre-stored in the electronic device and captured and collected during a historical period. After capturing and collecting the original video, the electronic device can process it in real time, or it can pre-store and process it later, for example, when the electronic device's CPU is low, or according to instructions from a staff member. Therefore, the generation of the training dataset in this embodiment can be based on real-time captured original video or on historical time periods, without any specific limitation here.

[0053] Step 330 : determining the motion trajectory and motion background of each target in the set scene based on the original video, and determining the outline of each target in each of the first number of video frames based on the first number of video frames in the original video.

[0054] It should be understood that the accuracy of multi-target pose estimation is directly related to the amount of training data: the larger the data volume, the higher the accuracy. However, as mentioned above, traditional image processing and machine learning are still unable to effectively identify and distinguish multiple animals that are closely interacting or touching, and they still rely mainly on a large amount of manual intervention for visual correction, resulting in insufficient training data for multi-target pose estimation. On the other hand, creating large datasets for multi-target pose estimation is extremely difficult. These datasets are easily affected by factors such as complex experimental scenarios, variable experimental conditions, diverse target types, and, in particular, the variable hair morphology of animals as targets. This makes it almost impossible for even a large dataset to include all situations for pose estimation of any type of target in any scenario. Therefore, the deep learning models trained based on this training often lack generalization, making it difficult to identify the poses of multiple animals in certain real experimental scenarios, and ultimately unable to guarantee the accuracy of multi-target pose estimation. To this end, in this embodiment, a training dataset is generated by combining the labeling of a small number of video frames with self-training of most video frames. This training dataset not only does not require a large amount of manual participation, but also can include complex occlusion relationships between multiple targets and can cover all situations of posture estimation of any type of target in any scene.

[0055] The following is a detailed description of the generation process of the above training data set:

[0056] First, based on the original video, the motion trajectory and motion background of each target in the set scene are determined. The motion trajectory is used to describe the location of each target when performing social behavior in the set scene, and the motion background is used to describe the environment in which each target is in the set scene when performing social behavior.

[0057] In one possible implementation, a tracking algorithm is used to extract the motion trajectory of each target in a set scene from each video frame of the original video. The tracking algorithm includes, but is not limited to, a background subtraction algorithm. In another possible implementation, an image segmentation algorithm is used to extract the motion background of each target in the set scene from each video frame of the original video. The image segmentation algorithm includes, but is not limited to, a frequency maximization algorithm. It is worth noting that the position and environment of each target in the set scene when performing social behaviors may continuously change throughout the entire original video shooting and acquisition process. Therefore, the extraction of motion trajectory and motion background is based on all video frames in the original video. The motion trajectory and motion background extracted from adjacent video frames may be the same or may change. For example, if the target remains stationary, the target's position and environment remain unchanged, and the corresponding motion trajectory and motion background remain unchanged. However, as the target moves with other targets during social behaviors, the target's position and environment gradually change, causing the corresponding motion trajectory and motion background to change.

[0058] Secondly, based on a first number of video frames in the original video, the outline of each target in each of the first number of video frames is determined.

[0059] Specifically, based on the occlusion relationship between the targets when performing social behaviors in the set scene, keyframe extraction is performed on each video frame in the original video to obtain a first number of video frames; the outline of each target in each of the first number of video frames is annotated to determine the outline of each target in each of the first number of video frames. It is explained here that annotation refers to marking the outline of the target in the video frame. As shown in Figure 2a, 201 and 202 represent the marked outlines of mouse a and mouse b, respectively. Accordingly, the video frame with the marked outline can be considered as a video frame carrying the outline label, and the outline label can be considered as being used to indicate the outline of the target in the corresponding video frame.

[0060] That is, only a first number of video frames are included in the annotation. Of course, in other embodiments, the first number of extracted key frames can be further reduced based on video redundancy. For example, for relatively similar adjacent key frames, only one key frame can be retained. Annotation can be performed manually or using the Segment Anything Model (SAM). As long as the first number is significantly smaller than the total number of all video frames in the original video, extensive manual effort can be effectively avoided.

[0061] Step 350 : self-training is performed on a second number of video frames in the original video according to the contours of each target in each of the first number of video frames to obtain the contours of each target in all video frames.

[0062] As mentioned above, in order to avoid a large amount of manual participation in labeling, only a first number of video frames in the original video are labeled. Therefore, in this embodiment, the second number of video frames in the original video are actually unlabeled video frames, and these video frames are labeled through self-training.

[0063] Among them, self-training refers to using a first number of video frames that have been labeled to predict the contours of each target in a second number of video frames that have not been labeled, and then combining the prediction results to complete the training of the deep self-training model. The deep self-training model can be a lightweight image instance segmentation model. When the training is completed, it is possible to output the contours of each target in all video frames based on the trained deep self-training model. It should be noted that the contours of each target in all video frames can be represented by corresponding contour labels, as shown in Figure 2b.

[0064] In step 370 , according to the motion trajectory of each target in the set scene, the outline of each target in the corresponding video frame and the motion background are fused to form a training data set.

[0065] The training data in the training dataset are video frames with instance annotations for each target with different occlusion relationships. It is worth mentioning that in this embodiment, instance annotation refers to marking the target in the video frame, which essentially means marking the outline of the target in the video frame.

[0066] After determining the outline of each target in each video frame, it is possible to mix multiple targets with different outlines but the same or similar motion background into the same video frame based on the trajectory overlap relationship reflected by the different motion trajectories of each target in the set scene, as shown in Figure 2c. Ultimately, a large amount of training data that can reflect the different occlusion relationships between multiple targets is obtained.

[0067] Through the above process, the training data for multi-target pose estimation no longer relies mainly on manual labeling, but can be created through self-training, thereby enriching the training data set, thereby effectively solving the problem of data scarcity for multi-target pose estimation in related technologies.

[0068] Referring to FIG. 3 , in an exemplary embodiment, step 350 may include the following steps:

[0069] Step 351 : Perform a first training on the deep self-training model based on the contours of each target in a first number of video frames.

[0070] That is, the deep self-training model is initially trained using a first number of video frames carrying contour labels as training data. The contour labels are used to indicate the true contours of each object in the video frames. Then, with the initial deep self-training model constructed and its parameters randomly initialized, the initial deep self-training model is trained using the current video frame to obtain a training result indicating the predicted contours of each object in the current video frame.

[0071] At this point, the error between the predicted contour in the training result and the true contour in the contour label can be used to determine whether the deep self-training model has completed the first training. For example, if the error is less than the error threshold, the deep self-training model has completed the first training.

[0072] Since the amount of the first amount of training data is small, at the end of the first training, the deep self-training model initially has the ability to predict the contours of each target in the video frame.

[0073] Step 353: Call the deep self-training model that has completed the first training to predict the contours of each target in each of the second number of video frames to obtain the contours of each target in each of the second number of video frames.

[0074] After obtaining a deep self-training model that initially has the ability to predict the contours of each target in the video frames, the second number of video frames can be input into the deep self-training model for prediction, thereby obtaining the contours of each target in the second number of video frames.

[0075] Step 355: Based on the contours of each target in each of the second number of video frames, the deep self-training model is trained for a second time until a fully trained deep self-training model is obtained.

[0076] After obtaining the contours of each target in each of the second number of video frames, the deep self-training model is trained a second time using the second number of video frames carrying contour labels as training data. The contour labels are used to indicate the true contours of each target in the video frames.

[0077] Similar to the first training, the error between the predicted contours in the training results and the true contours in the contour labels can be used to determine whether the deep self-training model has completed the second training. For example, if the error is less than the error threshold, the deep self-training model has completed the second training.

[0078] Of course, in other embodiments, the second training is not limited to using all of the second number of video frames. Inter-frame similarity and model output likelihood can also be considered, and accurately predicted video frames can be selected from the second number of video frames for the second training. This can also enrich the amount of data for model training and help improve the accuracy of the training model. It should be noted here that inter-frame similarity refers to the similarity of the contours of each object in adjacent video frames. Then, adjacent video frames are considered similar video frames. Suppose one video frame is in the first number of video frames, and the other video frame is in the second number of video frames. Because the former is manually annotated or annotated using SAM and has a higher accuracy, then the latter video frame can be considered an accurately predicted video frame. Participating in the second training is conducive to improving the accuracy of the training model. Model output likelihood measures the prediction accuracy of each of the second number of video frames by calculating probability. The greater the probability, the more accurately predicted the video frame is considered. Participating in the second training is conducive to improving the accuracy of the training model.

[0079] Based on this, the second training process can also include the following steps: based on the similarity of the target contours in adjacent video frames, selecting at least one video frame to be trained from the second number of video frames; or, based on the model output likelihood of the deep self-training model for predicting the target contours, selecting at least one video frame to be trained from the second number of video frames; inputting each video frame to be trained into the deep self-training model, and performing a second training of the deep self-training model until the second training is completed to obtain a deep self-training model that has completed the training.

[0080] Since the amount of data in the second number of video frames is larger, much larger than that in the first number of video frames, at the end of the second training, the deep self-training model can accurately predict the contours of each target in the video frames.

[0081] In step 357 , the trained deep self-training model is called to predict the contours of each target in all video frames contained in the original video to obtain the contours of each target in all video frames.

[0082] Under the influence of the above-mentioned embodiments, self-training of most video frames is achieved, so that the labeling of most video frames can be completed automatically, effectively avoiding a large amount of manual participation in labeling, which is conducive to reducing labor costs and improving labeling efficiency, thereby enabling the creation of large amounts of training data.

[0083] Referring to FIG. 4 , in an exemplary embodiment, after step 370 , the method may further include the following steps:

[0084] Step 410 : training the initial spatiotemporal instance segmentation model according to each training data in the training data set until a trained spatiotemporal instance segmentation model is obtained.

[0085] As previously mentioned, each training data set is composed of video frames with instances of objects in different occlusion relationships annotated. This can also be understood as a video frame carrying a contour label. This contour label is used to indicate the true contour of each object in the training data, and the occlusion relationships between the objects are reflected by the true contour of each object in the training data. It should be noted that the contour labels are generated by annotating the objects. Referring back to Figure 2a, 201 and 202 represent the contours of mice a and b, respectively. Image 200 can be considered a video frame carrying a contour label.

[0086] Among them, the training of the spatiotemporal instance segmentation model is essentially to construct a mathematical mapping relationship between instances (i.e., targets) and contours based on the real contours of multiple targets with different occlusion relationships in each training data. Then, by learning the spatiotemporal relationship of complex occlusions between targets, a mathematical mapping relationship can be constructed, and then based on this mathematical mapping relationship, multiple targets in the video frame to be processed can be accurately segmented into corresponding contours.

[0087] The following is a detailed description of the training process of the spatiotemporal instance segmentation model:

[0088] The first step is to build an initial spatiotemporal instance segmentation model and randomly initialize the parameters of the spatiotemporal instance segmentation model.

[0089] The initial spatiotemporal instance segmentation model is a machine learning model for segmenting instances into contours. For example, the machine learning model can be a Transformer model.

[0090] It is worth mentioning that the current Transformer model can only learn more complex spatiotemporal relationships of occlusion with the support of large amounts of training data. This results in the current Transformer model being unable to obtain a more accurate spatiotemporal instance segmentation model due to lack of training data, and thus cannot be applied to the task of multi-target pose estimation. However, the training dataset created in this application enables the Transformer model to learn the occlusion relationship between each target in each video frame contained in the original video with the support of large amounts of training data, and ultimately achieve multi-target pose estimation.

[0091] In the second step, the initial spatiotemporal instance segmentation model is trained based on the current training data in the training dataset to obtain the training results of the current training data.

[0092] The training results of the current training data are used to indicate the predicted profile of each target in the current training data.

[0093] The third step is to determine whether the spatiotemporal instance segmentation model has completed training based on the error between the predicted contour of each target indicated by the training result and the true contour of each target indicated by the current training data.

[0094] The true contours of the targets indicated by the current training data are determined based on the contour labels of the current training data, and are actually determined based on the targets with different occlusion relationships that are instance-labeled in the current training data.

[0095] Based on the difference between the predicted and true contours of each target, the error between the two can be calculated using a set function. The set function can be an expectation function, a loss function, etc. For example, the loss function includes but is not limited to a square loss function, a logarithmic loss function, an exponential loss function, a Hinge loss function, etc., which are not limited here.

[0096] If the error meets the convergence condition, the trained spatiotemporal instance segmentation model is obtained.

[0097] On the contrary, if the error does not meet the convergence condition, the parameters of the spatiotemporal instance segmentation model are updated, and the second step is returned to continue training the spatiotemporal instance segmentation model based on other training data in the training dataset.

[0098] It is worth mentioning that the convergence condition can be that the error is less than the error threshold, so as to improve the accuracy of the training model, or the number of iterations exceeds the iteration threshold, so as to increase the speed of the training model. That is, the convergence condition can be flexibly set according to the actual needs of the application scenario, and this embodiment does not limit it by comparison.

[0099] Based on the above training process, a spatiotemporal instance segmentation model is obtained that has been trained and has the ability to segment each target in the video frame into corresponding contours.

[0100] Step 430 : Training the initial single-view target pose estimation model according to the pose of each target in each of the third number of video frames until a trained single-view target pose estimation model is obtained.

[0101] First, it should be noted that the training data used to train the single-view object pose estimation model consists of video frames with pose-annotated single objects. In other words, the training data consists of video frames with pose labels that indicate the true pose of the single object in the training data. It should be noted that pose labels are generated by annotating the pose of a single object. Pose annotation essentially refers to labeling the pose of a single object in a video frame. As shown in Figure 4a, the pose of the mouse is actually represented by multiple key points on the mouse. Image 500 can be considered a video frame with pose labels.

[0102] In this embodiment, training data is generated based on each video frame in the original video, taking into account video redundancy and the similarity of object poses between adjacent video frames. Specifically, a keyframe extraction algorithm is used to extract multiple keyframes from each video frame in the original video, forming a third number of video frames. Multiple key points of each object in each of the third number of video frames are annotated to determine the pose of each object in each of the third number of video frames. It is worth noting that the keyframe extraction algorithm can be a temporal clustering algorithm, which is not a specific limitation here. Furthermore, after keyframe extraction, the third number is significantly smaller than the total number of video frames in the original video. Therefore, keypoint annotation can be performed manually or using the Segment Anything Model (SAM). Multiple key points of a single object in a video frame can be annotated, as can multiple key points of multiple objects in a video frame that are not occluded by an occlusion relationship. This is not a limitation here.

[0103] Secondly, the single-view target pose estimation model essentially constructs a mathematical mapping relationship between instances (i.e., targets) and poses based on the poses of single targets or multiple targets without occlusion in each training data set. Based on this mathematical mapping relationship, the pose of the target in the video frame to be processed can be accurately identified. It should be noted that if there are multiple targets in the video to be processed, the single-view target pose estimation model will output multiple prediction results, each of which is used to indicate the pose of one of the targets in the video to be processed.

[0104] The following is a detailed description of the training process of the single-view target pose estimation model:

[0105] The first step is to construct an initial single-view target pose estimation model and randomly initialize the parameters of the single-view target pose estimation model.

[0106] Among them, the initial single-view target pose estimation model is a machine learning model used to identify the pose of the instance.

[0107] In the second step, the initial single-view target pose estimation model is trained based on the current training data to obtain the training results of the current training data.

[0108] The training result of the current training data is used to indicate the predicted posture of the target in the current training data.

[0109] The third step is to determine whether the single-view target pose estimation model has completed training based on the error between the predicted pose of the target indicated by the training result and the true pose of the target indicated by the current training data.

[0110] The true posture of the target indicated by the current training data is determined based on the posture label of the current training data, and is actually determined based on a single target with posture annotations in the current training data or multiple targets without occlusion relationships.

[0111] Based on the difference between the predicted pose and the actual pose of the target, the error between the two can be calculated by setting a function. The set function can be an expectation function, a loss function, etc. For example, the loss function includes but is not limited to a square loss function, a logarithmic loss function, an exponential loss function, a Hinge loss function, etc., which are not limited here.

[0112] If the error meets the convergence condition, the trained single-view target pose estimation model is obtained.

[0113] On the contrary, if the error does not meet the convergence condition, the parameters of the single-view target pose estimation model are updated, and the process returns to the second step to continue training the single-view target pose estimation model based on other training data.

[0114] It is worth mentioning that the convergence condition can be that the error is less than the error threshold, so as to improve the accuracy of the training model, or the number of iterations exceeds the iteration threshold, so as to increase the speed of the training model. That is, the convergence condition can be flexibly set according to the actual needs of the application scenario, and this embodiment does not limit it by comparison.

[0115] Based on the above training process, a single-view target pose estimation model that has been trained and has the ability to recognize the pose of a single target in a video frame is obtained.

[0116] Step 450 : Using the trained spatiotemporal instance segmentation model and the trained single-view target pose estimation model, perform multi-target pose estimation on multiple targets in the video to be processed.

[0117] The video to be processed includes multiple video frames to be processed, and each video frame to be processed is captured by cameras at multiple different perspectives and aimed at multiple targets in any natural scene.

[0118] In a possible implementation, as shown in FIG5 , step 450 may include the following steps:

[0119] Step 451: Obtain each to-be-processed video frame in the to-be-processed video.

[0120] In step 453 , the trained spatiotemporal instance segmentation model is called to predict the contours of multiple targets in each video frame to be processed, and to obtain contour data of different viewing angles.

[0121] The contour data is used to indicate the contours of multiple targets in the video frame to be processed.

[0122] Because each video frame to be processed is captured and observed by multiple cameras at different viewpoints, contour prediction not only determines the contours of each target from different viewpoints (from different frames to be processed), but also determines the contours of each target from the same viewpoint (from the same frame to be processed). In other words, contour data corresponds one-to-one to viewpoint, and for each viewpoint, the contour data can reflect the spatiotemporal relationships between targets, even with complex occlusions.

[0123] In step 455, for the contour data of each perspective, the trained single-view target pose estimation model is called to perform pose estimation on each target in the contour data under the corresponding perspective, and the poses of multiple targets estimated under the corresponding perspective are fused to obtain multi-target pose data under the corresponding perspective.

[0124] The multi-target posture data is used to indicate the posture of each target in the video frame to be processed.

[0125] Based on the silhouette data from each viewpoint, pose estimation can determine the pose of each target at that viewpoint. Then, pose fusion can further determine the poses of multiple targets at that viewpoint. In other words, similar to the silhouette data, the pose data for multiple targets also corresponds one-to-one with the viewpoint. For each viewpoint, the pose data for multiple targets can reflect the different poses of multiple targets that are subject to occlusion.

[0126] Step 457 , obtaining multi-view camera parameters, and reconstructing the three-dimensional pose of the multiple targets based on the multi-view camera parameters and the multi-target pose data at different view angles, to obtain a three-dimensional pose estimation result of the multiple targets.

[0127] Among them, the three-dimensional pose estimation results are used to indicate the poses of multiple targets in three-dimensional space.

[0128] After obtaining the poses of multiple targets from a single viewpoint, the poses of the targets from different viewpoints can be fused based on the multi-view camera parameters. Ultimately, the poses of the targets in 3D space are obtained, i.e., the 3D pose estimation results for the targets, as shown in Figure 5a. It should be noted that the multi-view camera parameters are calculated by calibrating multiple cameras positioned at different viewpoints.

[0129] With the cooperation of the above embodiments, it is possible to accurately track and estimate the posture of multiple targets in any natural scene.

[0130] Figure 6 is a schematic diagram of a multi-target pose estimation system implemented in an application scenario. In this application scenario, the multi-target pose estimation system includes seven components: a multi-camera array behavior acquisition unit, a multi-animal dataset annotation unit, a big data generation unit, a spatiotemporal instance segmentation unit, a single-view animal pose estimation unit, a camera calibration unit, and a 3D pose fusion unit.

[0131] In Figure 6, multi-animal social interaction video data is first collected through a multi-view camera array. Then, a large amount of labeled data is generated from a small amount of manually annotated data through a data generation method. Next, a deep learning method is used to segment the animals and track the animals' postures in a single-view camera. Finally, combined with the calibration parameters of the multi-view camera array, the postures of multiple animals in three-dimensional space are estimated based on the geometric optimization relationship of the two-dimensional animal postures in three-dimensional space, thereby quantifying the social interaction behaviors of multiple animals in three-dimensional space.

[0132] Specifically, the multi-camera array behavior acquisition unit includes a shooting module and a synchronization module. The shooting module contains multiple cameras, and the synchronization module controls the multiple cameras to shoot and collect videos of animal behaviors at the same time.

[0133] The multi-animal data annotation unit includes an instance annotation module and a posture annotation module, which are used to annotate the animal's outline and body key points respectively.

[0134] The big data generation unit includes a trajectory extraction module, a background extraction module, a deep self-training module, and a data hybrid generation module. The trajectory extraction module extracts the movement trajectories of multiple animals, the background extraction module extracts the environmental background, the deep self-training module generates a large amount of labeled animal outline data through small model self-training, and the data hybrid generation module fuses the outputs of the first three modules into a new large dataset.

[0135] The spatiotemporal instance segmentation unit includes a model training module and a model verification module. The spatiotemporal instance segmentation model is trained in the generated big data set through the model training module, and the training results are verified using the model verification module.

[0136] The single-view animal pose estimation unit includes a model training module, a model verification module and a two-dimensional pose fusion module. The model training module is used to train the deep learning pose estimation model in a manually annotated pose dataset. The model verification module is used to verify the effect of the single-view animal pose estimation. The two-dimensional pose fusion module is used to fuse the poses of multiple animals under a single view.

[0137] The camera calibration unit includes a display module, a motion module and an optimization module. The display module displays the checkerboard required for camera calibration, the motion module controls the movement of the checkerboard at different angles and planes, and the optimization module realizes parameter optimization of multi-view camera calibration.

[0138] The three-dimensional posture fusion unit includes a permutation and combination module and a global optimization module. The permutation and combination module obtains a combination of animal postures from multiple perspectives, and then uses the global optimization module to obtain the motion postures of multiple animals in three-dimensional space based on three-dimensional geometric relationships.

[0139] In this application scenario, there is no need to use specific experimental chambers to restrict the social behavior between animals. Instead, the social behavior of multiple animals in any natural scene can be quantified. This not only solves the problem of large data set annotation in the absence of multi-animal posture data sets, but also solves the problem of few behavioral parameters and rough behavior quantification when studying animal social interactions through multi-perspective three-dimensional reconstruction and optimization methods, providing more accurate three-dimensional behavioral data for the study of animal social behavior.

[0140] The following is an embodiment of the device of the present application, which can be used to execute the method for generating a training data set involved in the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the method for generating a training data set involved in the present application.

[0141] Please refer to Figure 7. An embodiment of the present application provides a device 900 for generating a training data set, including but not limited to: a video acquisition module 910, a contour determination module 930, a self-training module 950, and a data fusion module 970.

[0142] The video acquisition module 910 is used to acquire the original video. Each video frame in the original video is captured and collected by multiple cameras arranged at different viewing angles targeting multiple targets performing social behaviors in a set scene.

[0143] The contour determination module 930 is used to determine the motion trajectory and motion background of each target in the set scene based on the original video, and determine the contour of each target in each of the first number of video frames based on the first number of video frames in the original video.

[0144] The self-training module 950 is used to perform self-training on a second number of video frames in the original video according to the contours of each target in the first number of video frames, so as to obtain the contours of each target in all video frames.

[0145] Data fusion module 970 is used to fuse the outline of each target in the corresponding video frame with the moving background based on the target's motion trajectory in the set scene to form a training data set. Each training data in the training data set is a video frame with instance annotations for each target with different occlusion relationships.

[0146] It should be noted that the training data set generation device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when generating the training data set. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the training data set generation device will be divided into different functional modules to complete all or part of the functions described above.

[0147] In addition, the apparatus for generating a training data set provided in the above embodiment and the method for generating a training data set are of the same concept, wherein the specific manner in which each module performs operations has been described in detail in the method embodiment and will not be repeated here.

[0148] Fig. 8 shows a schematic diagram of the structure of a server according to an exemplary embodiment. The server is applicable to the server 130 in the implementation environment shown in Fig. 1 .

[0149] It should be noted that the server is only an example adapted for the present application and should not be considered to provide any limitation on the scope of use of the present application. The server should not be interpreted as needing to rely on or necessarily having one or more components in the exemplary server 2000 shown in FIG8 .

[0150] The hardware structure of the server 2000 may vary greatly due to different configurations or performances. As shown in FIG8 , the server 2000 includes a power supply 210 , an interface 230 , at least one memory 250 , and at least one central processing unit (CPU) 270 .

[0151] Specifically, the power supply 210 is used to provide operating voltage for each hardware device on the server 2000 .

[0152] The interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices, for example, the interaction between the acquisition terminal 110 and the service terminal 130 in the implementation environment shown in FIG1 .

[0153] Of course, in other examples adapted by this application, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input-output interface 235, and at least one USB interface 237, as shown in FIG8 , which is not specifically limited here.

[0154] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon include an operating system 251, application 253 and data 255, etc. The storage method can be temporary storage or permanent storage.

[0155] Among them, the operating system 251 is used to manage and control the various hardware devices and application programs 253 on the server 2000 to enable the central processing unit 270 to calculate and process the massive data 255 in the memory 250. It can be Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0156] The application 253 is a computer-readable instruction that performs at least one specific task based on the operating system 251. It may include at least one module (not shown in FIG8 ), each of which may include computer-readable instructions for the server 2000. For example, the apparatus for generating a training dataset may be considered an application 253 deployed on the server 2000.

[0157] The data 255 may be photos, pictures, etc. stored in a disk, or may be videos to be processed, original videos, etc. stored in the memory 250 .

[0158] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer-readable instructions stored in the memory 250, thereby performing operations and processing on the massive data 255 in the memory 250. For example, the method for generating a training data set may be completed by the central processing unit 270 reading a series of computer-readable instructions stored in the memory 250.

[0159] In addition, the present application can also be implemented through hardware circuits or hardware circuits combined with software. Therefore, the implementation of the present application is not limited to any specific hardware circuits, software, or a combination of the two.

[0160] Please refer to FIG9 . An electronic device 4000 is provided in an embodiment of the present application. The electronic device 4000 may include a desktop computer, a laptop computer, a server, etc.

[0161] In FIG. 9 , the electronic device 4000 includes at least one processor 4001 and at least one memory 4003 .

[0162] Data exchange between the processor 4001 and the memory 4003 can be achieved via at least one communication bus 4002. The communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. The communication bus 4002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG9 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.

[0163] Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0164] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0165] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program instructions or codes in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited to these.

[0166] Computer-readable instructions are stored in the memory 4003 , and the processor 4001 can read the computer-readable instructions stored in the memory 4003 through the communication bus 4002 .

[0167] The computer-readable instructions are executed by one or more processors 4001 to implement the method for generating a training data set in the above embodiments.

[0168] In addition, an embodiment of the present application provides a storage medium having computer-readable instructions stored thereon, and the computer-readable instructions are executed by one or more processors to implement the above method for generating a training data set.

[0169] In an embodiment of the present application, a computer program product is provided. The computer program product includes computer-readable instructions, which are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the above-mentioned method for generating a training data set.

[0170] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0171] The above are only some of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for generating a training data set, wherein the training data set is used for multi-target pose estimation, characterized in that: The method comprises: Acquire an original video; each video frame in the original video is obtained by shooting and collecting multiple targets performing social behaviors in a set scene using multiple cameras arranged at different viewing angles; Based on the original video, determining the motion trajectory and motion background of each of the targets in the set scene, and based on a first number of video frames in the original video, determining the outline of each of the targets in each of the first number of video frames; According to the contours of each of the targets in the first number of video frames, self-training is performed on each of the second number of video frames in the original video to obtain the contours of each of the targets in all video frames; According to the motion trajectory of each target in the set scene, the outline and motion background of each target in the corresponding video frame are fused to form the training data set; each training data in the training data set is a video frame with instance annotations for each target with different occlusion relationships.

2. The method according to claim 1, characterized in that The self-training of the second number of video frames in the original video according to the contours of the targets in the first number of video frames to obtain the contours of the targets in all video frames includes: Based on the contours of each of the targets in each of the first number of video frames, the deep self-training model is trained for the first time; Calling the deep self-training model that has completed the first training, predicting the contour of each of the targets in each of the second number of video frames, and obtaining the contour of each of the targets in each of the second number of video frames; Based on the contours of each of the targets in each of the second number of video frames, training the deep self-training model for a second time until a trained deep self-training model is obtained; The trained deep self-training model is called to predict the contours of each target in all video frames contained in the original video to obtain the contours of each target in all video frames.

3. The method according to claim 2, characterized in that The second training of the deep self-training model based on the contours of each of the targets in each of the second number of video frames until a trained deep self-training model is obtained includes: Based on the similarity of each of the target contours in adjacent video frames, from each of the second number of video frames Select at least one video frame to be trained; or Based on the model output likelihood of the deep self-training model for each of the target contour predictions, selecting at least one video frame to be trained from each of the second number of video frames; Each of the video frames to be trained is input into the deep self-training model, and the deep self-training model is trained for the second time until the second training is completed, thereby obtaining a deep self-training model that has completed the training.

4. The method according to claim 1, characterized in that The determining, based on a first number of video frames in the original video, the contours of each of the targets in each of the first number of video frames comprises: Based on the occlusion relationship between the targets when performing social behaviors in the set scene, extract key frames from each video frame in the original video to obtain a first number of video frames; The contour of each of the objects in each of the first number of video frames is marked to determine the contour of each of the objects in each of the first number of video frames.

5. The method according to any one of claims 1 to 4, characterized in that: After forming the training data set, the method further includes: According to each of the training data in the training data set, the initial spatiotemporal instance segmentation model is trained until a trained spatiotemporal instance segmentation model is obtained; According to the postures of the targets in the third number of video frames, the initial single-view target posture estimation model is trained until a trained single-view target posture estimation model is obtained; Using the trained spatiotemporal instance segmentation model and the trained single-view target pose estimation model, multi-target pose estimation is performed on multiple targets in a video to be processed; the video to be processed includes multiple video frames to be processed, and each of the video frames to be processed is generated by cameras at multiple different perspectives for multiple targets.

6. The method according to claim 5, characterized in that The initial spatiotemporal instance segmentation model is trained according to each of the training data in the training data set until a trained spatiotemporal instance segmentation model is obtained, comprising: Based on the current training data in the training data set, the initial spatiotemporal instance segmentation model is trained to obtain a training result of the current training data; the training result is used to indicate the predicted contour of each of the targets in the current training data; Determining whether the spatiotemporal instance segmentation model has completed training according to the error between the predicted contour of each target indicated by the training result and the real contour of each target indicated by the current training data; the real contour of each target is determined based on each target with different occlusion relationships that has been instance-labeled in the current training data; If so, the spatiotemporal instance segmentation model that has completed training is obtained; otherwise, the parameters of the spatiotemporal instance segmentation model are updated, and the spatiotemporal instance segmentation model is continued to be trained based on other training data in the training data set.

7. The method according to claim 5, characterized in that Before training the initial single-view target pose estimation model according to the poses of each target in each of the third number of video frames, the method further includes: Using a key frame extraction algorithm, extract a plurality of key frames from each video frame included in the original video as each video frame of the third number; A plurality of key points of each of the objects in each of the third number of video frames are marked to determine the posture of each of the objects in each of the third number of video frames.

8. The method according to claim 5, characterized in that The method utilizes the trained spatiotemporal instance segmentation model and the trained single-view target pose estimation model to perform multi-target pose estimation on multiple targets in the processed video, including: Acquire each of the to-be-processed video frames in the to-be-processed video; The trained spatiotemporal instance segmentation model is called to predict the contours of multiple targets in each of the video frames to be processed, and to obtain contour data of different viewing angles; For the contour data of each perspective, the trained single-view target pose estimation model is called to estimate the pose of each target in each contour data under the corresponding perspective, and the poses of multiple targets estimated under the corresponding perspective are fused to obtain the multi-target pose data under the corresponding perspective; The multi-view camera parameters are obtained, and the 3D posture of the multiple targets is reconstructed according to the multi-view camera parameters and the multi-target posture data under different viewing angles to obtain the 3D posture estimation results of the multiple targets.

9. A device for generating a training data set, wherein the training data set is used for multi-target posture estimation, characterized in that: The device comprises: A video acquisition module is used to acquire an original video; each video frame in the original video is obtained by shooting and collecting multiple targets performing social behaviors in a set scene using multiple cameras arranged at different viewing angles; A contour determination module, configured to determine, based on the original video, a motion track and a motion background of each of the targets in the set scene, and to determine, based on a first number of video frames in the original video, a contour of each of the targets in each of the first number of video frames; A self-training module is used to train the target according to the contours of each target in each of the first number of video frames. Performing self-training on each of the second number of video frames in the original video to obtain contours of each of the targets in all the video frames; The data fusion module is used to fuse the outline and motion background of each target in the corresponding video frame according to the motion trajectory of each target in the set scene to form the training data set; each training data in the training data set is a video frame with instance annotations for each target with different occlusion relationships.

10. A storage medium having computer-readable instructions stored thereon, characterized in that: The computer-readable instructions are executed by one or more processors to implement the method for generating a training data set according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Animal individual identification system based on video tracking technology

    CN109377517A

  • Small-sample low-quality image target detection method based on multi-definition integrated self-training

    CN114067173A

  • Animal tracking and attitude estimation method and device, electronic equipment and storage medium

    CN116543006A

  • Training multi-object tracking models using simulation

    US20220092792A1

Cited By

  • Data processing method, device and equipment for hot galvanizing production line

    CN121682270A