Method and apparatus for mapping and classifying free social behaviors of multiple animals, electronic device, and storage medium
By extracting body posture data in multi-animal social behavior videos and applying manifold feature model and Gaussian kernel function, a low-dimensional behavior space and target cluster number model was established, which solved the problems of low efficiency of manual classification and incomplete definition, and achieved efficient and automated classification of multi-animal social behavior.
Patent Information
- Application Number
- PCT/CN2023/135942
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-05
AI Technical Summary
There are problems of inconsistent identification, incomplete definition and inefficient classification when artificially participating in animal behavior classification, especially in the complex situation of multi-animal social behavior.
By obtaining video data of free social behavior of multiple animals, body posture data are extracted and motion sequences, motion sequences and distance sequences are generated. Then, using the manifold feature model and Gaussian kernel function, a low-dimensional behavior space is established and the target cluster number model is determined, and the social behavior original video is finally inputted into the model to obtain the target behavior category.
The automated classification of social behaviors of multiple animals is achieved, the efficiency and consistency of classification is improved, and the problem of incomplete manual definition is avoided.
Smart Images

Figure CN2023135942_05062025_PF_FP_ABST
Abstract
Description
Method, device, electronic device and storage medium for mapping and classifying free social behavior of multiple animals Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device and storage medium for mapping and classifying free social behaviors of multiple animals. Background Art
[0002] The behavioral composition of animals is dynamic, hierarchical, high-dimensional and parallel. Therefore, it is difficult to determine the starting and ending time of a behavior when manually identifying behavior. The starting and ending points of behaviors observed by different experimenters are often different, resulting in inconsistent behavioral identification.
[0003] Secondly, a single behavior consists of movements across multiple parts of an animal's body, making it difficult to pre-define sufficient categories and parameters to precisely define each behavior, leading to incomplete behavioral definitions. To improve the accuracy of manual identification, experimenters need to repeatedly watch the animal's behavior. Identifying a single animal's one-hour behavioral video takes approximately ten hours. Identifying social behaviors across multiple animals is even more complex and time-consuming. Manual behavioral classification results in inconsistent identification, incomplete definitions, and low classification efficiency, a situation that needs further improvement.
[0004] Summary of the Invention
[0005] This application provides a method, device, electronic device, and storage medium for mapping and classifying free social behaviors of multiple animals, which can solve the problems of inconsistent identification, incomplete definition, and low classification efficiency of manual behavior classification in related technologies. The technical solution is as follows:
[0006] According to one aspect of the present application, a method for mapping and classifying the free social behavior of multiple animals includes: obtaining corresponding body posture data based on videos of the free social interaction of multiple animals; obtaining a corresponding sequence set based on the body posture data, wherein the sequence set includes a motion sequence, an action sequence, and a distance sequence; obtaining a two-dimensional time series and a corresponding target time segmentation point corresponding to each sequence based on each sequence in the sequence set; calling a pre-set popular feature model; inputting a manifold feature model based on all time points corresponding to all two-dimensional time series to obtain a corresponding two-dimensional manifold representation; determining a low-dimensional behavior space based on the target time segmentation point and the two-dimensional manifold representation; inputting the coordinates of the low-dimensional behavior space into a Gaussian kernel function to obtain the low-dimensional behavior space point density; determining a target cluster number model based on the stability of the low-dimensional behavior space point density; obtaining original videos of the social behavior of multiple animals; and inputting the original videos of the social behavior into the target cluster number model to obtain the corresponding target behavior category.
[0007] According to one aspect of the present application, a multi-animal free social behavior mapping and classification device includes but is not limited to:
[0008] Posture data acquisition module, which is used to obtain corresponding body posture data based on videos of multiple animals freely socializing;
[0009] A sequence set acquisition module, for acquiring a corresponding sequence set based on body posture data, wherein the sequence set includes a motion sequence, an action sequence, and a distance sequence;
[0010] The time acquisition module is used to obtain the two-dimensional time series and the corresponding target time segmentation points corresponding to each sequence in the sequence set;
[0011] Feature model retrieval module, used to retrieve pre-set popular feature models;
[0012] A two-dimensional manifold representation acquisition module is used to obtain the corresponding two-dimensional manifold representation based on the input manifold feature model of all time points corresponding to the entire two-dimensional time series;
[0013] Behavior space determination module, which is used to determine the low-dimensional behavior space based on the target time segmentation points and the two-dimensional manifold representation;
[0014] The spatial point density acquisition module inputs the coordinates of the low-dimensional behavior space into the Gaussian kernel function to obtain the point density of the low-dimensional behavior space;
[0015] Cluster number model determination module, which is used to determine the target cluster number model based on the stability of the point density in the low-dimensional behavior space;
[0016] The original video acquisition module is used to obtain the original videos of social behaviors of multiple animals;
[0017] The target behavior category acquisition module inputs the original social behavior video into the target clustering model to obtain the corresponding target behavior category.
[0018] In an exemplary embodiment, including but not limited to:
[0019] A two-dimensional representation acquisition module, which is used to acquire a two-dimensional representation corresponding to each sequence based on each sequence in the sequence set;
[0020] The discrete time segment acquisition module decomposes the two-dimensional representation corresponding to each sequence using the dynamic time alignment kernelization algorithm to obtain the discrete time segment corresponding to each sequence;
[0021] A time segmentation point determination module is used to determine the corresponding target time segmentation point based on discrete time segments;
[0022] The target time segmentation point determination module performs a merging operation on all the time segmentation points to determine the target time segmentation point.
[0023] In an exemplary embodiment, including but not limited to:
[0024] A two-dimensional time series acquisition module is used to obtain a two-dimensional time series corresponding to each sequence;
[0025] A six-dimensional time series determination module is used to determine the corresponding six-dimensional time series based on the two-dimensional time series of all sequences, wherein the six-dimensional time series is the two-dimensional time series corresponding to the motion sequence, the two-dimensional time series corresponding to the action sequence, and the two-dimensional time series corresponding to the distance sequence;
[0026] A training sequence acquisition module, used for acquiring a training six-dimensional time sequence based on the six-dimensional time sequence;
[0027] The manifold feature model determination module obtains the corresponding training two-dimensional manifold representation based on the UMAP algorithm operation of the training six-dimensional time series, and is used to determine the manifold feature model between the training six-dimensional time series and the training two-dimensional manifold representation.
[0028] In an exemplary embodiment, including but not limited to:
[0029] A matrix construction module, for constructing a similarity matrix based on the two-dimensional manifold representation and the time segmentation points;
[0030] The matrix representation data acquisition module uses the UMAP algorithm to obtain the matrix representation data of the similarity matrix in two-dimensional space;
[0031] The low-dimensional behavior space establishment module is used to establish a low-dimensional behavior space based on matrix representation data.
[0032] In an exemplary embodiment, including but not limited to:
[0033] Basis vector extraction module, which randomly extracts a portion of the time segments as basis vectors from all time segments;
[0034] A calling module for calling two-dimensional manifold representation and time segmentation points;
[0035] The matrix construction module uses a dynamic time alignment kernelization algorithm to measure the similarity of all time series segments and basis vectors based on the two-dimensional manifold representation and time segmentation points, so as to construct a similarity matrix.
[0036] In an exemplary embodiment, including but not limited to:
[0037] Function set acquisition module, which obtains the corresponding Gaussian kernel function set based on multiple variance parameters;
[0038] A behavior space point density set acquisition module inputs the coordinates of the low-dimensional behavior space into the Gaussian kernel function in the Gaussian kernel function set to obtain the low-dimensional behavior space point density output by each Gaussian kernel function;
[0039] The maximum stable value determination module is used to determine the maximum stable value and the minimum stable value corresponding to the variance parameter based on the stability of the point density of all low-dimensional behavior spaces;
[0040] The target cluster number model determination module defines the maximum stable value and the minimum stable value as the upper bound and the lower bound of the cluster number function, and uses them to determine the target cluster number model.
[0041] In an exemplary embodiment, including but not limited to:
[0042] Input module, used to input the original social behavior video into the target clustering model;
[0043] The segmentation module is used to segment the original social behavior video based on the upper and lower bounds of the cluster number function;
[0044] A segment set acquisition module is used to acquire the corresponding target video segment set;
[0045] The saving module is used to save each target video segment in the target video segment set to a folder named after the behavior category.
[0046] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory, wherein the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the multi-animal free social behavior mapping and classification method as described above.
[0047] According to one aspect of the present application, a storage medium stores computer-readable instructions thereon, wherein the computer-readable instructions are executed by one or more processors to implement the multi-animal free social behavior mapping and classification method as described above.
[0048] According to one aspect of the present application, a computer program product includes computer-readable instructions, which are stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the multi-animal free social behavior mapping and classification method as described above.
[0049] The beneficial effects of the technical solution provided by the present application are: through the implementation plan of the present application, in the process of behavior classification, an unsupervised behavior mapping and classification framework is designed according to the natural structure of social behavior, so that unknown social behaviors can be fully covered. The possible behavior categories are first consistently separated by the algorithm and then manually defined, which not only effectively covers the behavior categories that are difficult to define manually, but also obtains consistent results, and can classify the social behaviors of multiple animals, thereby improving the efficiency of classifying animal social behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts.
[0051] FIG1 is a schematic diagram of an implementation environment according to the present application;
[0052] FIG2 is a flow chart of a method for mapping and classifying free social behaviors of multiple animals according to an exemplary embodiment;
[0053] FIG3 is a flowchart of steps S121 to S123 in another method for mapping and classifying free social behaviors of multiple animals according to an exemplary embodiment;
[0054] FIG4 is a flowchart of steps S131 to S134 in another method for mapping and classifying free social behaviors of multiple animals according to an exemplary embodiment;
[0055] FIG5 is a flowchart of steps S151 to S154 in another method for mapping and classifying free social behaviors of multiple animals according to an exemplary embodiment;
[0056] FIG6 is a flowchart of steps S171 to S174 in another method for mapping and classifying free social behaviors of multiple animals according to an exemplary embodiment;
[0057] FIG7 is a flowchart of steps S181 to S184 in another method for mapping and classifying free social behaviors of multiple animals according to an exemplary embodiment;
[0058] FIG8 is a structural block diagram of a multi-animal free social behavior mapping and classification device according to an exemplary embodiment;
[0059] Fig. 9 is a structural block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0060] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.
[0061] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.
[0062] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0063] Figure 1 is a schematic diagram of an implementation environment involved in a multi-animal free social behavior mapping and classification method. The implementation environment includes a terminal, a server, and a service system configured with a member association database.
[0064] Specifically, the terminal can be used to run a client that provides video classification, and can be an electronic device such as a desktop computer, a laptop computer, a tablet computer, a smart phone, etc., which is not limited here.
[0065] Among them, the client provides video classification functions, such as a media player, a browser, etc., which can be in the form of an application or a web page. Accordingly, the user interface of the client for playing videos can be in the form of a program window or a web page, which is not limited here.
[0066] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. This server is an electronic device used to provide background services. For example, in this implementation, the server provides cloud storage services for audio and video data to terminals.
[0067] The server establishes a communication connection with the service system in advance through wired or wireless means, and realizes linkage with the service system through the communication connection. The service system can be a single server or a server cluster composed of multiple servers.
[0068] Through the interaction between the terminal and the server, the client running on the terminal will initiate a resource usage invitation to the server, requesting the server to determine the resource location and resource usage time through resource allocation, and then issue the invitation.
[0069] For the server, the resource allocation process is executed for the client of the inviter through the resource usage invitation linkage service system, and the invitation result indicating the location of the resource and the resource usage time is returned to the client of the inviter, so that the client of the inviter can further confirm whether to issue a resource usage invitation to the client of the invitee.
[0070] Of course, according to actual operational needs, the server and service system can also be integrated into the same server cluster so that resource allocation is completed by the same server cluster.
[0071] Referring to FIG. 2 , an embodiment of the present application provides a multi-animal free social behavior mapping and classification method, including:
[0072] S100, obtaining corresponding body posture data based on a video of multiple animals freely socializing;
[0073] Among them, a social behavior shooting module is used to shoot videos of multiple animals socializing freely, and then a deep learning posture estimation module is used to estimate the body postures of multiple animals socializing freely.
[0074] In addition, the multi-animal social posture estimation in the DeepLabCut tool can be used to obtain the body posture. First, some key video frames in the video of multiple animals socializing freely are extracted for manual annotation of the body posture of multiple animals. Depending on the type of animal, the predefined number of body points is also different. For example, mice use 16-point annotation, birds use 21-point annotation, and dogs use 17-point annotation; after the annotation is completed, the annotated frames are used to train the deep neural network. After the network training converges, the network is used to estimate the body posture of all animals in the video for subsequent processing; through the above process, the acquisition of body posture data can be achieved, and when using the converged neural network, there is no need for subsequent manual annotation, which can reduce the participation of human factors and improve the annotation efficiency.
[0075] S110: Acquire a corresponding sequence set based on the body posture data.
[0076] Among them, the sequence set includes motion sequence, action sequence and distance sequence. It should be pointed out that the motion sequence represents the instantaneous motion speed of each body point of social animals, and the calculation method is to multiply the absolute value of the difference between adjacent frames by the frame rate; the action sequence represents the movement of the limbs of social animals, and the calculation method is to subtract the coordinate value of the center point of the body from the coordinate value of each body point, and align each animal to the Cartesian coordinate origin; the distance sequence represents the position change between the limbs of social animals during the social process, and the calculation method is the Euclidean distance between the corresponding body points of each animal and other animals.
[0077] S120, based on each sequence in the sequence set, obtaining a two-dimensional time series and a corresponding target time segmentation point corresponding to each sequence;
[0078] The UMAP algorithm can be used to decompose the two-dimensional time series of each sequence, and then obtain the corresponding two-dimensional time series; please refer to Figure 3. The specific process of obtaining the target time segmentation point is as follows:
[0079] S121, obtaining a two-dimensional representation corresponding to each sequence based on each sequence in the sequence set;
[0080] Among them, in the embodiment of the present application, three sequences are included: motion sequence, action sequence and distance sequence, so for two-dimensional representation, there are three sequences of two-dimensional representation.
[0081] S122 , decomposing the two-dimensional representation corresponding to each sequence using a dynamic time alignment kernelization algorithm to obtain discrete time segments corresponding to each sequence.
[0082] S123, determining corresponding target time segmentation points based on discrete time segments;
[0083] The dynamic time alignment kernelization algorithm can be used to calculate the corresponding time segmentation point for each discrete time segment, and the required target time segmentation point can be determined from all the time segmentation points for subsequent use.
[0084] S130, calling a pre-set popular feature model;
[0085] Before calling the manifold feature model, the corresponding manifold feature model required by this application needs to be trained in advance. Please refer to Figure 4, the specific training process for the popular feature model is as follows:
[0086] S131, obtaining a two-dimensional time series corresponding to each sequence and a two-dimensional representation.
[0087] S132, determining a corresponding six-dimensional time series based on the two-dimensional time series of all sequences.
[0088] Among them, the six-dimensional time series is the two-dimensional time series corresponding to the motion sequence, the two-dimensional time series corresponding to the action sequence, and the two-dimensional time series corresponding to the distance sequence. The six-dimensional time series can be determined through the two-dimensional time series corresponding to the three sequences.
[0089] S133, selecting a portion of all six-dimensional time series as a training six-dimensional time series for model training.
[0090] S134, based on performing UMAP algorithm operation on the training six-dimensional time series to obtain the corresponding training two-dimensional manifold representation, the mapping relationship between the training six-dimensional time series and the training two-dimensional manifold representation can be learned. It should be pointed out here that the input here is the time point corresponding to the six-dimensional time series. When the amount of input training data is sufficient and the model based on the UMAP algorithm operation converges, the manifold feature model between the training six-dimensional time series and the training two-dimensional manifold representation can be determined.
[0091] S140 , based on all time points corresponding to all two-dimensional time series, inputting a manifold feature model, obtaining two-dimensional manifold representations corresponding to all time points.
[0092] S150, determining a low-dimensional behavior space based on the target time segmentation point and the two-dimensional manifold representation;
[0093] Among them, in order to determine the low-dimensional behavior space, it is necessary to first construct the corresponding similarity matrix before the corresponding representation data can be passed. To construct the similarity matrix, please refer to Figure 5. The specific process is as follows:
[0094] S151, randomly extract a portion of the time segments from all the time segments as basis vectors, and retrieve the two-dimensional manifold representation and time segmentation points;
[0095] The time segments here are the discrete time segments mentioned above. For the selection of basis vectors, a portion can be randomly selected for determination.
[0096] S152, based on the two-dimensional manifold representation and time segmentation points, a dynamic time alignment kernelization algorithm is used to measure the similarity of all time series segments and basis vectors, and a similarity matrix is constructed.
[0097] S153, after constructing the similarity matrix, a UMAP algorithm may be used to obtain matrix representation data of the similarity matrix in a two-dimensional space.
[0098] S154, establishing a low-dimensional behavior space based on the matrix representation data.
[0099] S160 , inputting the coordinates of the low-dimensional behavior space into a Gaussian kernel function to obtain the point density of the low-dimensional behavior space.
[0100] S170, a model for determining the number of target clusters based on the stability of point density in low-dimensional behavioral space;
[0101] To determine the target cluster number model, refer to Figure 6 and perform the following steps:
[0102] S171, obtaining a corresponding Gaussian kernel function set based on multiple variance parameters;
[0103] S172, inputting the coordinates of the low-dimensional behavior space into a Gaussian kernel function in the Gaussian kernel function set to obtain the low-dimensional behavior space point density output by each Gaussian kernel function;
[0104] S173, determining a maximum stable value and a minimum stable value corresponding to the variance parameter based on the stability of all low-dimensional behavior space point densities;
[0105] S174, defining the maximum stable value and the minimum stable value as the upper bound and the lower bound of the cluster number model, and determining the target cluster number model.
[0106] The execution steps from S171 to S173 are to determine the Gaussian kernel function, and the density of the low-dimensional behavior space output by the Gaussian kernel is stable as the judgment criterion, wherein the variance parameter of the Gaussian kernel is used as input, and the density of the low-dimensional behavior space calculated by the Gaussian kernel is output. Two variance parameters that can make the density of the low-dimensional behavior space stable are selected as boundaries, wherein the large variance parameter is used as the upper bound and the small variance parameter is used as the lower bound, and then the variances corresponding to the lower bound and the upper bound are used as the minimum stable value and the maximum stable value of the cluster number model, thereby determining the cluster number model.
[0107] Through the execution steps S100 to S170 in the embodiment of the present application, based on the body posture data in the initial stage, the difference from the prior art is that it is not only for a single animal, but using a deep neural network after training convergence, multiple animals can be labeled, and no manual participation in labeling is required, which reduces the involvement of human factors and improves the labeling efficiency. In addition, the motion data of the animal is calculated from sequences such as motion sequences, action sequences, and distance sequences. For the stable popular feature model obtained by training, a stable two-dimensional manifold representation can be output. A low-dimensional space model can be established in combination with the target time segmentation point, and a suitable Gaussian function is used to calculate the density of the low-dimensional behavior space. The variance parameter is selected based on the density stability of the low-dimensional behavior space point. After the variance parameter is selected, the corresponding target cluster number model can be determined using the cluster number model. The target cluster number model can be used to distinguish the videos of the social behaviors of multiple animals, and the classification of animal behaviors can be achieved. After the target cluster number model distinguishes the different social behaviors of the animals, the social behaviors distinguished by the target cluster number model can be manually named.
[0108] S180, obtaining original videos of social behaviors of multiple animals, and inputting the original videos of social behaviors into a target clustering model to obtain corresponding target behavior categories.
[0109] In the process of specifically classifying the social behaviors of multiple animals, please refer to Figure 7. The specific behaviors include:
[0110] S181, inputting the original social behavior video into the target clustering model;
[0111] Among them, the original video of social behavior input into the target clustering model can be a video of one animal or multiple animals. For the solution in the embodiment of the present application, different animals can be labeled for different animals in the video.
[0112] S182, segmenting the original social behavior video based on the upper and lower bounds of the cluster number model;
[0113] Among them, when performing social behavior raw video segmentation, the upper bound of the cluster number model is used to distinguish the global differences in animal social behavior, while the lower bound of the cluster number model is used to subdivide the fine differences in animal social behavior.
[0114] S183, obtaining a corresponding target video segment set;
[0115] Among them, after finely distinguishing the upper and lower bounds of the cluster number model, a target video segment set distinguished by the cluster number model can be obtained, and each behavior corresponds to a target video segment set.
[0116] S184, saving each target video segment in the target video segment set into a folder named after the behavior category;
[0117] Among them, after distinguishing the different social behaviors of animals, each social behavior of the animal can be named manually, and after the original video of the social behavior is input, the system can save each target video clip in a corresponding file according to the classification of the target clustering model.
[0118] The following is an embodiment of the device of the present application, which can be used to implement the multi-animal free social behavior mapping and classification method involved in this application. For details not disclosed in the device embodiment of this application, please refer to the method embodiment of the multi-animal free social behavior mapping and classification method involved in this application.
[0119] Please refer to FIG8 . In an embodiment of the present application, a multi-animal free social behavior mapping and classification device is provided, including but not limited to:
[0120] A posture data acquisition module 200 is used to acquire corresponding body posture data based on a video of multiple animals freely socializing;
[0121] A sequence set acquisition module 210 is used to acquire a corresponding sequence set based on the body posture data, wherein the sequence set includes a motion sequence, an action sequence, and a distance sequence;
[0122] A time acquisition module 220 is used to acquire a two-dimensional time series and a corresponding target time segmentation point corresponding to each sequence in the sequence set;
[0123] The feature model retrieval module 230 is used to retrieve a preset popular feature model;
[0124] A two-dimensional manifold representation acquisition module 240 is configured to acquire a corresponding two-dimensional manifold representation based on an input manifold feature model for all time points corresponding to all two-dimensional time series;
[0125] A behavior space determination module 250 is used to determine a low-dimensional behavior space based on target time segmentation points and a two-dimensional manifold representation;
[0126] A spatial point density acquisition module 260 inputs the coordinates of the low-dimensional behavior space into a Gaussian kernel function to obtain the point density of the low-dimensional behavior space;
[0127] A cluster number model determination module 270 is used to determine a target cluster number model based on the stability of the point density in the low-dimensional behavior space;
[0128] The original video acquisition module 280 is used to acquire the original videos of the social behaviors of multiple animals;
[0129] The target behavior category acquisition module 290 inputs the original social behavior video into the target clustering model to obtain the corresponding target behavior category.
[0130] In an exemplary embodiment, including but not limited to:
[0131] A two-dimensional representation acquisition module 300 is configured to acquire a two-dimensional representation corresponding to each sequence based on each sequence in the sequence set;
[0132] The discrete time segment acquisition module 310 decomposes the two-dimensional representation corresponding to each sequence using a dynamic time alignment kernelization algorithm to obtain the discrete time segment corresponding to each sequence;
[0133] A time segmentation point determination module 320 is used to determine corresponding target time segmentation points based on discrete time segments;
[0134] The target time segmentation point determination module 330 performs a merging operation on all the time segmentation points to determine the target time segmentation point.
[0135] In an exemplary embodiment, including but not limited to:
[0136] A two-dimensional time series acquisition module 400 is used to acquire a two-dimensional time series corresponding to each sequence;
[0137] A six-dimensional time series determination module 410 is used to determine a corresponding six-dimensional time series based on the two-dimensional time series of all sequences, wherein the six-dimensional time series is a two-dimensional time series corresponding to the motion sequence, a two-dimensional time series corresponding to the action sequence, and a two-dimensional time series corresponding to the distance sequence;
[0138] A training sequence acquisition module 420 is used to acquire a training six-dimensional time sequence based on the six-dimensional time sequence;
[0139] The manifold feature model determination module 430 obtains the corresponding training two-dimensional manifold representation based on the UMAP algorithm operation on the training six-dimensional time series, and is used to determine the manifold feature model between the training six-dimensional time series and the training two-dimensional manifold representation.
[0140] In an exemplary embodiment, including but not limited to:
[0141] A matrix construction module 500 is used to construct a similarity matrix based on the two-dimensional manifold representation and the time segmentation points;
[0142] The matrix representation data acquisition module 510 uses the UMAP algorithm to obtain matrix representation data of the similarity matrix in two-dimensional space;
[0143] The low-dimensional behavior space establishment module 520 is used to establish a low-dimensional behavior space based on the matrix representation data.
[0144] In an exemplary embodiment, including but not limited to:
[0145] A basis vector extraction module 600 randomly extracts a portion of the time segments as basis vectors from all the time segments;
[0146] A retrieval module 610 is used to retrieve a two-dimensional manifold representation and a time segmentation point;
[0147] The matrix construction module 620 measures the similarity between all time series segments and basis vectors using a dynamic time alignment kernelization algorithm based on the two-dimensional manifold representation and time segmentation points, so as to construct a similarity matrix.
[0148] In an exemplary embodiment, including but not limited to:
[0149] A function set acquisition module 700 acquires a corresponding Gaussian kernel function set based on a plurality of variance parameters;
[0150] The behavior space point density set acquisition module 710 inputs the coordinates of the low-dimensional behavior space into the Gaussian kernel function in the Gaussian kernel function set to obtain the low-dimensional behavior space point density output by each Gaussian kernel function;
[0151] The maximum stable value determination module 720 is used to determine the maximum stable value and the minimum stable value corresponding to the variance parameter based on the stability of the point density of all low-dimensional behavior spaces;
[0152] The target cluster number model determination module 730 defines the maximum stable value and the minimum stable value as the upper bound and the lower bound of the cluster number function, and uses them to determine the target cluster number model.
[0153] In an exemplary embodiment, including but not limited to:
[0154] Input module 800, for inputting the original social behavior video into the target clustering model;
[0155] A segmentation module 810 is used to segment the original social behavior video based on the upper bound and lower bound of the cluster number function;
[0156] The segment set acquisition module 820 is used to acquire the corresponding target video segment set;
[0157] The saving module 830 is configured to save each target video segment in the target video segment set into a folder named after the behavior category.
[0158] It should be noted that the multi-animal free social behavior mapping and classification device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when performing multi-animal free social behavior mapping and classification. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the multi-animal free social behavior mapping and classification device will be divided into different functional modules to complete all or part of the functions described above.
[0159] In addition, the multi-animal free social behavior mapping classification device provided in the above embodiment and the multi-animal free social behavior mapping classification method embodiment belong to the same concept, and the specific way in which each module performs the operation has been described in detail in the method embodiment and will not be repeated here.
[0160] Please refer to FIG9 . An electronic device 4000 is provided in an embodiment of the present application. The electronic device 4000 may include a desktop computer, a laptop computer, a server, etc.
[0161] In FIG. 9 , the electronic device 4000 includes at least one processor 4001 and at least one memory 4003 .
[0162] Data exchange between the processor 4001 and the memory 4003 can be achieved via at least one communication bus 4002. The communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. The communication bus 4002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG9 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0163] Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
[0164] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0165] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program instructions or codes in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited to these.
[0166] Computer-readable instructions are stored in the memory 4003 , and the processor 4001 can read the computer-readable instructions stored in the memory 4003 through the communication bus 4002 .
[0167] The computer-readable instructions are executed by one or more processors 4001 to implement the multi-animal free social behavior mapping and classification method in the above embodiments.
[0168] In addition, an embodiment of the present application provides a storage medium having computer-readable instructions stored thereon, and the computer-readable instructions are executed by one or more processors to implement the multi-animal free social behavior mapping and classification method as described above.
[0169] In an embodiment of the present application, a computer program product is provided, which includes computer-readable instructions stored in a storage medium. One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the multi-animal free social behavior mapping and classification method as described above.
[0170] In the solution in the embodiment of the present application, it is first necessary to use a social behavior shooting module to shoot videos of multiple animals socializing freely, so as to obtain corresponding body posture data. No manual labeling is required during the process, and the training and convergence of the deep neural network is directly used for processing, thereby reducing the involvement of human factors; then, the corresponding motion sequence, action sequence and distance sequence can be obtained based on the body posture data. Correspondingly, for each sequence, there is a corresponding two-dimensional time series and a discrete time segment corresponding to the two-dimensional time series, and the time segmentation point can be obtained through the discrete time segment.
[0171] Then, the pre-trained manifold feature model is called up. Correspondingly, all time points corresponding to all two-dimensional time series are input into the manifold feature model to obtain the two-dimensional manifold representation corresponding to all time points. At the same time, a part of all time segments is randomly extracted as basis vectors, and the dynamic time alignment kernelization algorithm is used to calculate the similarity between the time series segments and the basis vectors, so that a similarity matrix can be constructed. The similarity matrix is obtained using the UMAP algorithm to obtain the matrix representation data of the similarity matrix in two-dimensional space. After obtaining the matrix representation data, a low-dimensional behavior space can be established.
[0172] After establishing the low-dimensional behavior space, the coordinates of the low-dimensional behavior space can be obtained, and the coordinate information can be input into the Gaussian kernel function. The density of the low-dimensional behavior space output by the Gaussian kernel is stable as the judgment criterion, wherein the variance parameter of the Gaussian kernel is used as input, and the density of the low-dimensional behavior space calculated by the Gaussian kernel is output. Two variance parameters that can stabilize the density of the low-dimensional behavior space are selected as boundaries, wherein the large variance parameter is used as the upper bound and the small variance parameter is used as the lower bound. The variances corresponding to the lower bound and the upper bound are then used as the minimum stable value and the maximum stable value of the clustering number model, thereby determining the target clustering number model.
[0173] After determining the target clustering number model, the original videos of social behaviors of multiple animals are input into the target clustering number model. After fine differentiation by the upper and lower bounds of the target clustering number model, a set of target video clips differentiated by the target clustering number model can be obtained, and each target video clip is saved in a corresponding file, thereby realizing the classification of social behaviors of multiple animals.
[0174] To sum up, the scheme in the embodiment of the present application can first classify the social behaviors of multiple animals, and in the process of classifying the original videos of social behaviors, while ensuring the classification accuracy, the problem of inconsistent social behavior classification can be solved through the animal's body posture data and decomposition. The manifold features and low-dimensional space mapping classification solve the problem of incomplete definition of social behavior. The strategy of classifying first and then defining unsupervised social behavior classification solves the problem of low efficiency in social behavior classification. Without human participation, the efficiency of subdividing and classifying the social behaviors of multiple animals can be improved.
[0175] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0176] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for mapping and classifying the free social behaviors of multiple animals, characterized in that, it includes: Obtaining corresponding body pose data based on the videos of multiple animals' free social interactions; Obtaining a corresponding sequence set based on the body pose data, where the sequence set includes motion sequences, action sequences, and distance sequences; Obtaining a corresponding two-dimensional time series and a corresponding target time segmentation point for each sequence in the sequence set; Invoking a pre-set manifold feature model; Inputting all time points corresponding to all two-dimensional time series into the manifold feature model to obtain corresponding two-dimensional manifold representations; Determining a low-dimensional behavior space based on the target time segmentation point and the two-dimensional manifold representation; Inputting the coordinates of the low-dimensional behavior space into a Gaussian kernel function to obtain the point density of the low-dimensional behavior space; Determining a target clustering number model based on the stability of the point density of the low-dimensional behavior space; Obtaining the original video of the social behaviors of multiple animals; Inputting the original video of the social behaviors into the target clustering number model to obtain the corresponding target behavior categories.
2. The method according to claim 1, characterized in that, in the process of obtaining the corresponding target time segmentation point, the method further includes: Obtaining a corresponding two-dimensional representation for each sequence in the sequence set; Decomposing the two-dimensional representation corresponding to each sequence using the dynamic time warping kernelization algorithm to obtain discrete time segments corresponding to each sequence; Determining the corresponding target time segmentation point based on the discrete time segments; Performing a merging operation on all the time segmentation points to determine the target time segmentation point.
3. The method according to claim 2, characterized in that, before invoking the manifold feature model, the method further includes: Obtaining the two-dimensional time series of the two-dimensional representation corresponding to each sequence; Determining a corresponding six-dimensional time series based on the two-dimensional time series of all sequences, where the six-dimensional time series is the two-dimensional time series corresponding to the motion sequence, the two-dimensional time series corresponding to the action sequence, and the two-dimensional time series corresponding to the distance sequence; Obtaining a training six-dimensional time series based on the six-dimensional time series; Performing a UMAP algorithm operation on the training six-dimensional time series to obtain a corresponding training two-dimensional manifold representation, and determining the manifold feature model between the training six-dimensional time series and the training two-dimensional manifold representation.
4. The method according to claim 2, characterized in that, in the process of determining the low-dimensional behavior space based on the target time segmentation point and the two-dimensional manifold representation, the method further includes: Constructing a similarity matrix based on the two-dimensional manifold representation and the time segmentation point; Using the UMAP algorithm to obtain the matrix representation data of the similarity matrix in the two-dimensional space; Establishing a low-dimensional behavior space based on the matrix representation data.
5. The method according to claim 1, characterized in that, in the process of constructing the similarity matrix based on the two-dimensional manifold representation and the time segmentation point, the method further includes: Randomly extracting a part of the segments from all the time segments as basis vectors, and invoking the two-dimensional manifold representation and the time segmentation point; Measuring the similarity between all time series segments and the basis vectors using the dynamic time warping kernelization algorithm based on the two-dimensional manifold representation and the time segmentation point, and constructing a similarity matrix.
6. The method according to claim 1, wherein, in the process of determining the target clustering number model based on the stability of the low-dimensional behavior space point density, the method further includes: obtaining a corresponding Gaussian kernel function set based on a plurality of variance parameters; inputting the coordinates of the low-dimensional behavior space into the Gaussian kernel functions in the Gaussian kernel function set to obtain the low-dimensional behavior space point density output by each Gaussian kernel function; determining the maximum stability value and the minimum stability value corresponding to the variance parameter based on the stability of all the low-dimensional behavior space point densities; defining the maximum stability value and the minimum stability value as the upper bound and the lower bound of the clustering number function, and determining the target clustering number model.
7. The method according to claim 1, wherein, in the process of segmenting the original social behavior video based on the target clustering number model, the method further includes: inputting the original social behavior video into the target clustering number model; segmenting the original social behavior video based on the upper bound and the lower bound of the clustering number function; obtaining a corresponding set of target video segments; saving each target video segment in the set of target video segments to a folder named after the behavior category.
8. A multi-animal free social behavior mapping and classification device, wherein, it includes: a posture data acquisition module, which is used to acquire corresponding body posture data based on the video of multi-animal free social interaction; a sequence set acquisition module, which is used to acquire a corresponding sequence set based on the body posture data, wherein the sequence set includes a motion sequence, an action sequence, and a distance sequence; a time acquisition module, which is used to acquire a corresponding two-dimensional time sequence and a corresponding target time segmentation point for each sequence based on each sequence in the sequence set; a feature model retrieval module, which is used to retrieve a pre-set manifold feature model; a two-dimensional manifold representation acquisition module, which inputs all time points corresponding to all two-dimensional time sequences into the manifold feature model to acquire a corresponding two-dimensional manifold representation; a behavior space determination module, which is used to determine a low-dimensional behavior space based on the target time segmentation point and the two-dimensional manifold representation; a space point density acquisition module, which inputs the coordinates of the low-dimensional behavior space into a Gaussian kernel function to acquire the low-dimensional behavior space point density; a clustering number model determination module, which is used to determine a target clustering number model based on the stability of the low-dimensional behavior space point density; an original video acquisition module, which is used to acquire the original social behavior video of multi-animals; a target behavior category acquisition module, which inputs the original social behavior video into the target clustering number model to acquire a corresponding target behavior category.
9. An electronic device, wherein, it includes: at least one processor and at least one memory, wherein, the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the multi-animal free social behavior mapping and classification method according to any one of claims 1 to 7.
10. A storage medium, on which computer-readable instructions are stored, wherein, the computer-readable instructions are executed by one or more processors to implement the multi-animal free social behavior mapping and classification method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Behavior quantification method based on states and maps and terminal
CN112057079A
Three-dimensional human body movement data dividing method
CN1975779A
Action recognition apparatus, learning apparatus, and action recognition method
US20220076003A1
State and graph-based behavior quantification method and terminal
WO2022027590A1