Behavior analysis method and apparatus, network device and storage medium
Video data is obtained through multiple camera components, instance segmentation and single-object cropping are performed, and three-dimensional spatial attitude information of the target object is constructed, which solves the problem of inaccurate behavior analysis caused by single camera shooting, and achieves more accurate behavior analysis results.
Patent Information
- Application Number
- PCT/CN2023/138958
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-19
AI Technical Summary
Existing video shots by single cameras result in inaccurate behavioral analysis results and lack of three-dimensional information.
Multiple camera components with different locations are used to obtain multiple video data, and through instance segmentation and single-object cropping, the three-dimensional spatial attitude information of the target object is constructed and behavioral analysis is performed.
More accurate spatial attitude information acquisition and behavior analysis results are achieved, and continuous spatial attitude information can be constructed while some camera components are blocked.
Smart Images

Figure CN2023138958_19062025_PF_FP_ABST
Abstract
Description
Behavior analysis method, device, network equipment and storage medium Technical Field
[0001] The present application relates to the field of computer technology, and more specifically, to a behavior analysis method, apparatus, network device, and storage medium. Background Art
[0002] Dogs are highly social animals with a long history of symbiosis with humans. Working dogs have played a crucial role throughout human history. However, training an excellent working dog requires significant human resources and time. Using the PAT test to screen puppies can significantly improve the success rate of working dog training. The PAT, short for Puppy Aptitude Test, is a series of personality tests designed to identify a puppy's potential for development. Currently, the PAT test typically uses a single camera to capture video of the dog and analyze its behavior.
[0003] However, compared with three-dimensional information, the two-dimensional images captured by a single camera lose the precise information of one dimension, resulting in inaccurate behavior analysis results.
[0004] Summary of the Invention
[0005] This application provides a behavior analysis method, apparatus, network device, and storage medium, which can more accurately analyze the behavior of a target object. The technical solution is as follows:
[0006] According to one aspect of the present application, the present application provides a behavior analysis method, which includes: based on N camera components with different positions, obtaining N video data of a first target object and a second target object, and extracting N video frames with the same acquisition time as a data group, wherein N is greater than or equal to 4; performing instance segmentation on the first target object and the second target object in the video frames in the data group, and after single-target cropping, determining a first image group corresponding to the first target object and a second image group corresponding to the second target object; tracking a first key point of the first target object based on the first image group to determine first spatial posture information; tracking a second key point of the second target object based on the second image group to determine second spatial posture information; and determining behavior analysis results of the first target object and the second target object based on the first spatial posture information and the second spatial posture information.
[0007] In an exemplary embodiment, determining the behavioral analysis results of the first target object and the second target object based on the first spatial posture information and the second spatial posture information includes: determining multiple data groups corresponding to the target time period based on the acquisition time, and determining the behavioral analysis results of the first target object and the second target object based on the first spatial posture information and the second spatial posture information of the multiple data groups.
[0008] In an exemplary embodiment, the instance segmentation of the first target object and the second target object in the video frame in the data group includes: sampling the video data to obtain multiple frames of sampled images; obtaining first contour information of the first target object and second contour information of the second target object in the sampled images; training an instance segmentation model based on the sampled images, the first contour information and the second contour information; and performing instance segmentation on the first target object and the second target object in the video frame in the data group based on the trained instance segmentation model.
[0009] In an exemplary embodiment, after training the instance segmentation model based on the sampled image, the first contour information and the second contour information, the method further includes: extracting video images of the video data to form a time-series image set, so as to perform spatiotemporal continuity training on the instance segmentation model based on the time-series image set to determine the trained instance segmentation model.
[0010] In an exemplary embodiment, the single target cropping includes: based on the image after instance segmentation, removing the background image other than the first target object or the second target object, and replacing it with an empty scene image.
[0011] In an exemplary embodiment, the first target object is a canine animal, and the first key points include the canine animal's: head, nose, eyes, ears, chest, shoulders, front elbows, wrists, front feet, back, hips, hind elbows, hind hooves, hind feet, base of tail, middle of tail, and tip of tail; the second target object is a person, and the second key points include the person's: head, chest, shoulders, elbows, hands, hips, knees, and feet.
[0012] In an exemplary embodiment, the video data also includes a third target object, and the method also includes: obtaining a third image group including the third target object; tracking a third key point of the third target object based on the third image group to determine third spatial trajectory information; and determining behavioral analysis results of the first target object and the third target object based on the first spatial posture information of the first target object and the third spatial trajectory information of the third target object.
[0013] In an exemplary embodiment, the step of acquiring the first image group and the second image group includes: removing part of the images in the first image group or the second image group according to an occlusion relationship between the first target object and the second target object.
[0014] According to one aspect of the present application, the present application provides a behavior analysis device, which includes: a video data acquisition module, which is used to acquire N video data of a first target object and a second target object based on N camera components with different positions, and extract N video frames with the same acquisition time as a data group, where N is greater than or equal to 4; an image group acquisition module, which is used to perform instance segmentation on the first target object and the second target object in the video frames in the data group, and after single target cropping, determine a first image group corresponding to the first target object and a second image group corresponding to the second target object; a key point recognition module, which is used to track the first key point of the first target object based on the first image group to determine the first spatial posture information; and track the second key point of the second target object based on the second image group to determine the second spatial posture information; an analysis result acquisition module, which is used to determine the behavior analysis results of the first target object and the second target object based on the first spatial posture information and the second spatial posture information.
[0015] According to one aspect of the present application, the present application provides a network device comprising: a memory, a transceiver, and a processor; wherein the memory is used to store a computer program; the transceiver is used to send and receive data under the control of the processor; and the processor is used to read the computer program in the memory and execute the behavior analysis method as described above.
[0016] According to one aspect of the present application, the present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the behavior analysis method described above is implemented.
[0017] The beneficial effects of the technical solution provided by this application are:
[0018] The solution of the present application can be applied in the behavior analysis scenario of the target object. It can analyze the spatial posture of at least two target objects based on the video data of multiple camera components in different positions, and perform interactive behavior analysis of at least two target objects based on the spatial posture between the target objects. Compared with the existing method of using a single camera for analysis, this solution can perform video acquisition and posture analysis based on at least four camera components to construct a three-dimensional spatial posture, obtain more accurate spatial posture information, and thus obtain more accurate behavior analysis results. Moreover, when some camera components are blocked, the remaining camera components can also complete the construction of the spatial posture, so that more continuous spatial posture information can be obtained, and more accurate behavior analysis results can be obtained. Specifically, this solution can obtain N video data of the first target object and the second target object based on N camera components with different positions, and extract N video frames with the same acquisition time as the data group according to the acquisition time of each video frame, wherein N is greater than or equal to 4. After installing each camera component, this solution can calibrate the installation position of the camera component; then, the first target object and the second target object in the video frame in the data group can be instance segmented to determine the contours of the first target object and the second target object, and after single target cropping, the first image group corresponding to the first target object and the second image group corresponding to the second target object can be determined, wherein the first image group includes the first target object and does not include the second target object, and the second image group includes the second target object and does not include the first target object; then, based on the first image in the first image group and the installation position of each camera component, the first key point of the first target object is tracked to determine the first spatial posture information; based on the second image in the second image group and the installation position of each camera component, the second key point of the second target object is tracked to determine the second spatial posture information; and then, based on the first spatial posture information and the second spatial posture information of the data group within a period of time, the interactive behavior relationship between the first target object and the second target object is determined to determine the behavior analysis results of the first target object and the second target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0020] FIG1 is a schematic diagram of the steps of a behavior analysis method according to an embodiment of the present application;
[0021] FIG2 is a flow chart of a behavior analysis method according to an embodiment of the present application;
[0022] FIG3 is a flow chart of a behavior analysis method according to another embodiment of the present application;
[0023] FIG4 is a schematic diagram of the structure of a behavior analysis device according to an embodiment of the present application;
[0024] FIG5 is a structural block diagram of a network device according to an embodiment of the present application;
[0025] FIG6 is a structural block diagram of a user equipment according to an embodiment of the present application; DETAILED DESCRIPTION
[0026] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout identify the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.
[0027] Those skilled in the art will understand that, unless otherwise specified, the singular forms "a," "an," "the," and "the" used herein may also include the plural forms, and "a plurality" refers to two or more, and other quantifiers are similar. It should be further understood that the term "comprising" used in the specification of this application refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connection or wireless coupling. The term "and / or" used herein describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0028] The solution of the present application can be applied in the behavior analysis scenario of the target object. It can analyze the spatial posture of at least two target objects based on the video data of multiple camera components in different positions, and perform interactive behavior analysis of at least two target objects based on the spatial posture between the target objects. Compared with the existing method of using a single camera for analysis, this solution can perform video acquisition and posture analysis based on at least four camera components to construct a three-dimensional spatial posture, obtain more accurate spatial posture information, and thus obtain more accurate behavior analysis results. Moreover, when some camera components are blocked, the remaining camera components can also complete the construction of the spatial posture, so that more continuous spatial posture information can be obtained, and more accurate behavior analysis results can be obtained. Specifically, this solution can obtain N video data of the first target object and the second target object based on N camera components with different positions, and extract N video frames with the same acquisition time as the data group according to the acquisition time of each video frame, wherein N is greater than or equal to 4. After installing each camera component, this solution can calibrate the installation position of the camera component; then, the first target object and the second target object in the video frame in the data group can be instance segmented to determine the contours of the first target object and the second target object, and after single target cropping, the first image group corresponding to the first target object and the second image group corresponding to the second target object can be determined, wherein the first image group includes the first target object and does not include the second target object, and the second image group includes the second target object and does not include the first target object; then, based on the first image in the first image group and the installation position of each camera component, the first key point of the first target object is tracked to determine the first spatial posture information; based on the second image in the second image group and the installation position of each camera component, the second key point of the second target object is tracked to determine the second spatial posture information; and then, based on the first spatial posture information and the second spatial posture information of the data group within a period of time, the interactive behavior relationship between the first target object and the second target object is determined to determine the behavior analysis results of the first target object and the second target object.
[0029] This application designs an AI-based method for capturing 3D PAT behavior. It has the following features: It can avoid interference during multi-target tracking. It can compensate for information loss caused by mutual occlusion using multiple cameras. It uses AI to identify targets and continuously track them in the video. It can remove identified targets to avoid tracking interference with the current target. It can also achieve interference-free 3D reconstruction. It can also use behavioral analysis techniques to perform batch analysis and evaluation of canine PAT tests.
[0030] For this purpose, the present application uses a synchronized shooting scheme of eight cameras (camera components) (multiple camera components are distributed almost evenly around the shooting site), and connects the camera to the acquisition host through a USB3.0 interface; during acquisition, a multi-process communication queue is used to control the communication and frame synchronization between processes; the Zhang calibration method is used to estimate the intrinsic parameter matrix and relative position of the camera; machine learning methods are used to track the key points of the body when the animal behavior occurs; a machine learning-based method is used to continuously track and cut out the contours of the target, and the vacancies after cutting are replaced with images of the empty field, and a posture tracking tool such as DeepLabCut is used to track the key points of the operating target. This method is used to track the key points of all targets in the scene, and they are reconstructed in three-dimensional space to obtain three-dimensional behavioral data of multi-target interaction. The PAT test of dogs is analyzed and evaluated by calculating the displacement, speed, and following relationship of the skeleton in three-dimensional space, and using behavioral classification to evaluate the type and duration of behavior during the interaction.
[0031] The processing flow of this application includes: collecting behavioral videos, completing camera calibration, sampling and annotating training sets, training neural networks to identify multi-target contours, removing the background of the identified multi-target contours, tracking key points using posture tracking tools, completing three-dimensional reconstruction, and evaluating and analyzing the PAT test content. Collecting behavioral videos requires the use of multiple cameras to synchronously capture the current scene. The number of cameras is more than 4. The number of cameras used in this system is 8, with a resolution of 1280×720 and a frame rate of 30 frames. The camera calibration adopts Zhang's calibration method, and the checkerboard is used to calibrate the internal and external parameters of the camera. The collected behavioral video is sampled and framed, and the contours are annotated using data annotation tools. This is used as a training set for training in the artificial intelligence network. This system uses an instance segmentation model (such as the YOLACT++ network) for training. In order to ensure the accuracy and continuity of recognition, this application can use the results obtained by YOLACT++ training as the input training set for spatiotemporal correlation training. This system uses the video instance segmentation framework (VisTR) to track continuous frames in the video and ultimately obtain continuous and stable segmentation instances. The resulting segmented instances were individually subtracted from the background, and then a pose tracking tool was used to track key body points. The resulting 2D coordinates were then used for 3D reconstruction based on calibration information. Complete 3D behavioral data was then processed based on the behavioral paradigm to evaluate the dog's performance during the test.
[0032] As shown in Figure 1, this application collects behavioral data (video data includes behavior) using 8-view cameras synchronously, with a resolution of 1280×720 and a frame rate of 30 frames. The collected video is frame-sampled, and the sampled images are contour segmented and annotated. The training set is trained using an instance segmentation model (such as the YOLACT++ neural network), and the verification evaluation is performed on the validation set. The entire network is then applied to all videos to obtain dual-target segmentation instances. In order to obtain more stable instance segmentation, VisTR is used to use the data output by the YOLACT++ network as the training set and validation set for spatiotemporal continuity tracking training, and ultimately obtains stable instance segmentation.
[0033] After obtaining the instance segmentation, to avoid conflicts in the identification of body keypoints, one of the targets was removed from the scene. The resulting gap was replaced with an empty field image, resulting in a single target on the scene and preventing interference between targets. The pose tracking tool DeepLabCut was used to track the keypoints of the single target for 2D pose estimation. For dogs, 27 keypoints were recorded (including: head, nose, eyes, ears, chest, shoulders, front elbows, wrists, front feet, back, hips, hind elbows, hooves, hind feet, tail base, tail midpoint, and tail tip), while for humans, 14 keypoints were recorded (including: head, chest, shoulders, elbows, hands, hips, knees, and feet). The 2D pose estimation results for both the human and dog targets were then processed based on the calibrated camera (camera assembly) installation positions to complete 3D reconstruction, yielding 3D data (spatial pose information for both targets). The 3D reconstruction method used for this 3D reconstruction is triangulation, and the calibration method used is Zhang Zhengyou's.
[0034] After obtaining three-dimensional behavioral data, behavioral analysis and assessment are performed using different methods depending on the behavioral paradigm. For the following test, the dog's ability to follow the person is tested. The displacement between the dog's head center of mass and the tester's foot center of mass is evaluated, with evaluation criteria of non-following, hesitant following, and active following. In the non-following test, the distance between the dog's head center of mass and the tester's foot center of mass continuously increases, with no tendency to approach. In the hesitant following test, the dog tends to approach the person, but the following is not continuous, with large fluctuations in the relative distance and remaining outside the interaction range. In the active following test, the displacement between the dog and the person remains consistent, with the relative distance fluctuating within a small range. For the social attraction test, the dog's closeness to the person is tested. The displacement between the dog's head center of mass and the tester's foot center of mass is evaluated, with the same evaluation criteria as for the following test. For the visual acuity test, the dog's response to a moving object (such as a third target object) is tested, with evaluation criteria including: moving away, standing still and showing no interest, or excitedly playing. An unsupervised behavioral clustering analysis system (Behavior Atlas) is used to perform unsupervised clustering of dog behaviors, and behavioral actions such as walking, staying still, lying down, and biting excitedly are obtained. This is used to evaluate the types and duration of the dog's behavioral performance. For height testing, it is necessary to test the dog's struggle when it is off the ground. The evaluation criteria are: continuous struggling, alternating struggle and calmness, and calmness without struggling. It is necessary to calculate the change in the dog's limb angular velocity. When the dog is struggling continuously, the average change in the angular velocity of the dog's limbs is larger, and when it is calm, the change in the angular velocity of the limbs is smaller or unchanged. People and dogs (the first target object and the second target object) can analyze their interactive behavior based on their spatial posture information. For dogs and third target objects (such as objects), objects usually do not have too many key points, so the movement trajectory of the objects can be analyzed to analyze the dog's behavior.
[0035] As shown in Figure 2, the overall analysis process of this application includes: deploying cameras (camera components) and calibrating the camera positions, then multiple cameras synchronously collecting data, performing instance segmentation, determining the outlines of the first target object and the second target object, and performing two-dimensional posture tracking, identifying the key points of the first target object and the second target object, and performing three-dimensional reconstruction based on the positions of multiple cameras (analyzing the depth of the two-dimensional key points, depth refers to the distance between the key points and the camera, and depth estimation can be performed through multi-eye vision) to obtain three-dimensional spatial posture information. After obtaining the spatial postures of multiple target objects, paradigm assessment and comprehensive evaluation can be performed to determine the behavioral test results.
[0036] Specifically, the present application provides a behavior analysis method, as shown in FIG3 , which includes:
[0037] Step 302: Based on N camera components at different positions, obtain N video data of the first target object and the second target object, and extract N video frames with the same acquisition time as a data group, where N is greater than or equal to 4.
[0038] Step 304 : performing instance segmentation on the first target object and the second target object in the video frames in the data group and performing single-object cropping to determine a first image group corresponding to the first target object and a second image group corresponding to the second target object.
[0039] Step 306: Track the first key points of the first target object based on the first image group to determine first spatial posture information; and track the second key points of the second target object based on the second image group to determine second spatial posture information. Specifically, as an optional embodiment, the first target object is a canine, and the first key points include the canine's head, nose, eyes, ears, chest, shoulders, front elbows, wrists, front feet, back, hips, hind elbows, hind hooves, hind feet, tail base, tail middle, and tail tip; the second target object is a person, and the second key points include the person's head, chest, shoulders, elbows, hands, hips, knees, and feet.
[0040] Step 308: Determine behavior analysis results of the first target object and the second target object based on the first spatial posture information and the second spatial posture information.
[0041] The solution of the present application can be applied in the behavior analysis scenario of the target object. It can analyze the spatial posture of at least two target objects based on the video data of multiple camera components in different positions, and perform interactive behavior analysis of at least two target objects based on the spatial posture between the target objects. Compared with the existing method of using a single camera for analysis, this solution can perform video acquisition and posture analysis based on at least four camera components to construct a three-dimensional spatial posture, obtain more accurate spatial posture information, and thus obtain more accurate behavior analysis results. Moreover, when some camera components are blocked, the remaining camera components can also complete the construction of the spatial posture, so that more continuous spatial posture information can be obtained, and more accurate behavior analysis results can be obtained. Specifically, this solution can obtain N video data of the first target object and the second target object based on N camera components with different positions, and extract N video frames with the same acquisition time as the data group according to the acquisition time of each video frame, wherein N is greater than or equal to 4. After installing each camera component, this solution can calibrate the installation position of the camera component; then, the first target object and the second target object in the video frame in the data group can be instance segmented to determine the contours of the first target object and the second target object, and after single target cropping, the first image group corresponding to the first target object and the second image group corresponding to the second target object can be determined, wherein the first image group includes the first target object and does not include the second target object, and the second image group includes the second target object and does not include the first target object; then, based on the first image in the first image group and the installation position of each camera component, the first key point of the first target object is tracked to determine the first spatial posture information; based on the second image in the second image group and the installation position of each camera component, the second key point of the second target object is tracked to determine the second spatial posture information; and then, based on the first spatial posture information and the second spatial posture information of the data group within a period of time, the interactive behavior relationship between the first target object and the second target object is determined to determine the behavior analysis results of the first target object and the second target object.
[0042] After determining the spatial postures of multiple target objects corresponding to a collection time, a time window can be established to analyze the continuity of the spatial postures, and then determine the behavior, such as determining whether the two are moving away, approaching, or following each other. Specifically, as an optional embodiment, the behavior analysis results of the first target object and the second target object are determined based on the first spatial posture information and the second spatial posture information, including: determining multiple data groups corresponding to the target time period according to the collection time, and determining the behavior analysis results of the first target object and the second target object based on the first spatial posture information and the second spatial posture information of the multiple data groups. Determine the spatial posture information corresponding to the collection time (first spatial posture information or second spatial posture information), and adjust the spatial posture information corresponding to the collection time based on the spatial posture information before or after the collection time, such as learning based on the attention mechanism for optimization. In addition, it is also possible to adjust each other based on the correlation and collision volume between the first spatial posture information and the second spatial posture information at the same moment (or multiple moments).
[0043] This solution can extract images corresponding to some frames in the video data and manually annotate them (annotate the outline of the first target object or the second target object), thereby training a model based on the partial images, and identifying all video images of the video based on the model (instance segmentation) to determine the outline of the target object. Specifically, as an optional embodiment, the instance segmentation of the first target object and the second target object in the video frames in the data set includes: sampling the video data to obtain multiple frames of sampled images; obtaining the first outline information of the first target object and the second outline information of the second target object in the sampled images; training an instance segmentation model based on the sampled images, the first outline information and the second outline information; and performing instance segmentation on the first target object and the second target object in the video frames in the data set based on the trained instance segmentation model.
[0044] Since the behavior of the target object in consecutive frames in the video data is continuous, all the video images of the video data can be formed into a time-series image set to train the instance segmentation model for spatiotemporal continuity. Specifically, as an optional embodiment, after training the instance segmentation model based on the sampled image, the first contour information and the second contour information, the method further includes: extracting the video images of the video data to form a time-series image set, and training the instance segmentation model for spatiotemporal continuity based on the time-series image set to determine the trained instance segmentation model. After instance segmentation, in order to avoid conflicts in the recognition of body key points, one of the targets is cut out in the scene, and the vacancy after cutting is replaced with an empty scene image, so that there is only one target on the scene, avoiding mutual interference between targets. Specifically, as an optional embodiment, the single target cropping includes: based on the image after instance segmentation, removing the background image other than the first target object or the second target object, and replacing it with the empty scene image.
[0045] The interactive behaviors of humans and dogs (first target object and second target object) can be analyzed based on the spatial posture information of the two. For dogs and third target objects (such as objects), objects usually do not have too many key points, so the movement trajectory of the objects can be analyzed to analyze the behavior of dogs. Specifically, as an optional embodiment, the video data also includes a third target object, and the method also includes: obtaining a third image group containing the third target object; tracking the third key point of the third target object based on the third image group to determine the third spatial trajectory information; determining the behavior analysis results of the first target object and the third target object based on the first spatial posture information of the first target object and the third spatial trajectory information of the third target object. For visual sensitivity testing, it is necessary to test the dog's performance towards moving objects (such as the third target object), including: leaving, staying still and not interested, playing excitedly, etc.
[0046] This solution sets up multiple camera components. At a certain moment, in the images captured by some camera components, there may be occlusion between the two target objects. Therefore, it is possible to analyze whether the image can determine the two complete target objects (all key points are complete, or some important key points are complete) to determine the occlusion relationship and select a better image for three-dimensional spatial posture analysis. Specifically, as an optional embodiment, the step of obtaining the first image group and the second image group includes: removing part of the images in the first image group or the second image group based on the occlusion relationship between the first target object and the second target object. The number of remaining images is not less than two.
[0047] Based on the above embodiment, the present application also provides a behavior analysis device, as shown in FIG4 , comprising:
[0048] The video data acquisition module 402 is used to acquire N video data of the first target object and the second target object based on N camera components at different positions, and extract N video frames with the same acquisition time as a data group, where N is greater than or equal to 4.
[0049] The image group acquisition module 404 is used to perform instance segmentation on the first target object and the second target object in the video frame in the data group, and after single target cropping, determine the first image group corresponding to the first target object and the second image group corresponding to the second target object.
[0050] The key point recognition module 406 is used to track the first key point of the first target object based on the first image group to determine the first spatial posture information; and to track the second key point of the second target object based on the second image group to determine the second spatial posture information.
[0051] The analysis result acquisition module 408 is used to determine the behavior analysis results of the first target object and the second target object based on the first spatial posture information and the second spatial posture information.
[0052] The implementation of the embodiment of the present application is similar to the implementation of the above-mentioned method embodiment. The specific implementation can refer to the specific implementation of the above-mentioned method embodiment, which will not be repeated here.
[0053] The solution of the present application can be applied in the behavior analysis scenario of the target object. It can analyze the spatial posture of at least two target objects based on the video data of multiple camera components in different positions, and perform interactive behavior analysis of at least two target objects based on the spatial posture between the target objects. Compared with the existing method of using a single camera for analysis, this solution can perform video acquisition and posture analysis based on at least four camera components to construct a three-dimensional spatial posture, obtain more accurate spatial posture information, and thus obtain more accurate behavior analysis results. Moreover, when some camera components are blocked, the remaining camera components can also complete the construction of the spatial posture, so that more continuous spatial posture information can be obtained, and more accurate behavior analysis results can be obtained. Specifically, this solution can obtain N video data of the first target object and the second target object based on N camera components with different positions, and extract N video frames with the same acquisition time as the data group according to the acquisition time of each video frame, wherein N is greater than or equal to 4. After installing each camera component, this solution can calibrate the installation position of the camera component; then, the first target object and the second target object in the video frame in the data group can be instance segmented to determine the contours of the first target object and the second target object, and after single target cropping, the first image group corresponding to the first target object and the second image group corresponding to the second target object can be determined, wherein the first image group includes the first target object and does not include the second target object, and the second image group includes the second target object and does not include the first target object; then, based on the first image in the first image group and the installation position of each camera component, the first key point of the first target object is tracked to determine the first spatial posture information; based on the second image in the second image group and the installation position of each camera component, the second key point of the second target object is tracked to determine the second spatial posture information; and then, based on the first spatial posture information and the second spatial posture information of the data group within a period of time, the interactive behavior relationship between the first target object and the second target object is determined to determine the behavior analysis results of the first target object and the second target object.
[0054] It should be noted that the division of units and / or modules in the embodiments of the present application is schematic and is merely a logical functional division. In actual implementation, there may be other division methods. In addition, the functional units and / or modules in the various embodiments of the present application may be integrated into one processing unit and / or module, or each unit and / or module may exist physically alone, or two or more units and / or modules may be integrated into one unit and / or module. The above-mentioned integrated units and / or modules may be implemented in the form of hardware or in the form of software functional units and / or modules.
[0055] If the integrated units and / or modules are implemented in the form of software functional units and / or modules and sold or used as independent products, they can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0056] In addition, the data transmission device and data transmission method provided in the above embodiments are based on the same application concept. Since the principles of solving problems by the method and the device are similar, the implementation of the device and the method can refer to each other, and the repeated parts will not be repeated.
[0057] Fig. 5 is a structural block diagram showing a network device according to an exemplary embodiment.
[0058] As shown in FIG. 5 , the network device 1100 includes at least a processor 1110 , a memory 1120 , and a transceiver 1130 .
[0059] The transceiver 1130 is used to receive and send data under the control of the processor 1110 .
[0060] In FIG5 , the bus architecture may include any number of interconnected buses and bridges, specifically linking together various circuits such as one or more processors represented by processor 1110 and memory represented by memory 1120. The bus architecture may also link together various other circuits such as peripherals, voltage regulators, and power management circuits, all of which are well known in the art and, therefore, will not be further described herein. The bus interface provides an interface. The transceiver 1130 may be a plurality of components, namely, a transmitter and a receiver, providing units and / or modules for communicating with various other devices over a transmission medium, such as a wireless channel, a wired channel, an optical cable, or the like.
[0061] The processor 1110 is responsible for managing the bus architecture and general processing, and the memory 1120 can store data used by the processor 1110 when performing operations.
[0062] Optionally, the processor 1110 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor 1110 may also employ a multi-core architecture. The processor 1110 and the memory 1120 may also be physically separated.
[0063] The processor 1110 calls the computer program stored in the memory 1120 to execute any one of the methods for allocating a cell radio network temporary identifier provided in the above embodiments of the present application according to the obtained executable instructions.
[0064] Fig. 6 is a structural block diagram showing a user equipment according to an exemplary embodiment.
[0065] As shown in FIG6 , the user equipment 1300 includes at least a processor 1310 , a memory 1320 , and a transceiver 1330 .
[0066] The transceiver 1330 is used to receive and send data under the control of the processor 1310.
[0067] In Figure 6, the bus architecture may include any number of interconnected buses and bridges, specifically linking together various circuits of one or more processors represented by processor 1310 and memory represented by memory 1320. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are all well known in the art and, therefore, will not be described further herein. The bus interface provides an interface. The transceiver 1330 may be a plurality of components, namely, a transmitter and a receiver, providing units and / or modules for communicating with various other devices on a transmission medium, such as wireless channels, wired channels, optical cables, and the like. For different user devices, the user interface 1340 may also be an interface capable of connecting external or internal devices as required, and the connected devices include but are not limited to a keypad, a display, a speaker, a microphone, a joystick, and the like.
[0068] The processor 1310 is responsible for managing the bus architecture and general processing, and the memory 1320 can store data used by the processor 1310 when performing operations.
[0069] Optionally, the processor 1310 may be a CPU (central processing unit), an ASIC (application-specific integrated circuit), an FPGA (field programmable gate array), or a CPLD (complex programmable logic device). The processor 1310 may also employ a multi-core architecture. The processor 1310 and the memory 1320 may also be physically separated.
[0070] The processor 1310 calls the computer program stored in the memory 1320 to execute any one of the methods for allocating a cell radio network temporary identifier provided in the above embodiments of the present application according to the obtained executable instructions.
[0071] It should be noted here that the above-mentioned device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned method embodiment and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those in the method embodiment will not be described in detail here.
[0072] In addition, an embodiment of the present application provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the data transmission method of each of the above embodiments. The storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as a floppy disk, hard disk, magnetic tape, magneto-optical disk (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)), etc.
[0073] In an embodiment of the present application, a program product is provided. For example, the program product is an FPGA chip or a DSP chip. The program product includes executable instructions stored in a storage medium. A processor reads the executable instructions from the storage medium, so that when the executable instructions are executed by the processor, the data transmission method described in each of the above embodiments is implemented.
[0074] The solution of the present application can be applied in the behavior analysis scenario of the target object. It can analyze the spatial posture of at least two target objects based on the video data of multiple camera components in different positions, and perform interactive behavior analysis of at least two target objects based on the spatial posture between the target objects. Compared with the existing method of using a single camera for analysis, this solution can perform video acquisition and posture analysis based on at least four camera components to construct a three-dimensional spatial posture, obtain more accurate spatial posture information, and thus obtain more accurate behavior analysis results. Moreover, when some camera components are blocked, the remaining camera components can also complete the construction of the spatial posture, so that more continuous spatial posture information can be obtained, and more accurate behavior analysis results can be obtained. Specifically, this solution can obtain N video data of the first target object and the second target object based on N camera components with different positions, and extract N video frames with the same acquisition time as the data group according to the acquisition time of each video frame, wherein N is greater than or equal to 4. After installing each camera component, this solution can calibrate the installation position of the camera component; then, the first target object and the second target object in the video frame in the data group can be instance segmented to determine the contours of the first target object and the second target object, and after single target cropping, the first image group corresponding to the first target object and the second image group corresponding to the second target object can be determined, wherein the first image group includes the first target object and does not include the second target object, and the second image group includes the second target object and does not include the first target object; then, based on the first image in the first image group and the installation position of each camera component, the first key point of the first target object is tracked to determine the first spatial posture information; based on the second image in the second image group and the installation position of each camera component, the second key point of the second target object is tracked to determine the second spatial posture information; and then, based on the first spatial posture information and the second spatial posture information of the data group within a period of time, the interactive behavior relationship between the first target object and the second target object is determined to determine the behavior analysis results of the first target object and the second target object.
[0075] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) that contain computer-usable program code.
[0076] The present application is described with reference to the flowchart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the process and / or box in the flowchart and / or block diagram can be implemented by computer-executable instructions. These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the function specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.
[0077] These processor-executable instructions may also be stored in a processor-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the processor-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0078] These processor-executable instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0079] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0080] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A behavior analysis method, characterized in that, The method includes: Based on N camera components with different positions, N video data of a first target object and a second target object are obtained, and N video frames with the same acquisition time are extracted as a data group, where N is greater than or equal to 4; After performing instance segmentation on the first target object and the second target object in the video frames of the data group and performing single-target cropping, a first image group corresponding to the first target object and a second image group corresponding to the second target object are determined; Based on the first image group, the first key points of the first target object are tracked to determine the first spatial pose information; based on the second image group, the second key points of the second target object are tracked to determine the second spatial pose information; According to the first spatial pose information and the second spatial pose information, the behavior analysis results of the first target object and the second target object are determined.
2. The method according to claim 1, characterized in that, The determining of the behavior analysis results of the first target object and the second target object according to the first spatial pose information and the second spatial pose information includes: According to the acquisition time, multiple data groups corresponding to the target time period are determined, and according to the first spatial pose information and the second spatial pose information of the multiple data groups, the behavior analysis results of the first target object and the second target object are determined.
3. The method according to claim 1, characterized in that, The performing of instance segmentation on the first target object and the second target object in the video frames of the data group includes: Frame sampling is performed on the video data to obtain multiple sampled images; The first contour information of the first target object and the second contour information of the second target object are obtained in the sampled images; Based on the sampled images, the first contour information and the second contour information, an instance segmentation model is trained; According to the trained instance segmentation model, instance segmentation is performed on the first target object and the second target object in the video frames of the data group.
4. The method according to claim 3, characterized in that, After training the instance segmentation model based on the sampled images, the first contour information and the second contour information, the method further includes: The video images of the video data are extracted to form a temporal image set, so as to perform spatio-temporal continuity training on the instance segmentation model according to the temporal image set to determine the trained instance segmentation model.
5. The method according to claim 1, characterized in that, The single-target cropping includes: Based on the image after instance segmentation, the background image other than the first target object or the second target object is removed and replaced according to the empty-scene image.
6. The method according to claim 1, characterized in that, The first target object is a canine animal, and the first key points include: head, nose, both eyes, both ears, chest, both shoulders, both front elbows, both front wrists, both front feet, back, both hips, both rear elbows, both rear hooves, both rear feet, root of the tail, middle of the tail, tip of the tail of the canine animal; the second target object is a human, and the second key points include: head, chest, both shoulders, both elbows, both hands, both hips, both knees, both feet of the human.
7. The method according to claim 1, characterized in that, The video data further includes a third target object, and the method further includes: Obtaining a third image group including the third target object; Based on the third image group, the third key points of the third target object are tracked to determine the third spatial trajectory information; According to the first spatial pose information of the first target object and the third spatial trajectory information of the third target object, the behavior analysis results of the first target object and the third target object are determined.
8. The method according to claim 1, characterized in that, The steps of obtaining the first image group and the second image group include: Remove partial images from the first image group or the second image group according to the occlusion relationship between the first target object and the second target object.
9. A behavior analysis device, characterized in that, The device includes: A video data acquisition module, configured to acquire N video data of a first target object and a second target object based on N camera components with different positions, and extract N video frames with the same acquisition time as a data group, where N is greater than or equal to 4; An image group acquisition module, configured to perform instance segmentation on the first target object and the second target object in the video frames of the data group, and after performing single-target cropping, determine a first image group corresponding to the first target object and a second image group corresponding to the second target object; A key point recognition module, configured to track the first key points of the first target object based on the first image group to determine first spatial pose information; track the second key points of the second target object based on the second image group to determine second spatial pose information; An analysis result acquisition module, configured to determine a behavior analysis result of the first target object and the second target object according to the first spatial pose information and the second spatial pose information.
10. A network device, characterized in that, It includes: A memory, a transceiver, and a processor; wherein, the memory is used to store computer programs; the transceiver is used to transmit and receive data under the control of the processor; The processor is configured to read the computer program in the memory and execute the method as claimed in claims 1-8.
Citation Information
Patent Citations
Testing device and method for evaluating social ability of dogs and people
CN112106674A
Attitude determination method, device and equipment, and storage medium
CN112241731A
Human body posture detection method and device and computer equipment
CN112633196A
Video labeling method, device and equipment and computer readable storage medium
CN112950667A
Behavior and pupil information synchronous analysis method and device, equipment and medium
CN114220168A