Behavior recognition method and device, computer equipment, storage medium and storage medium
By extracting and inputting target video clips into the position prediction model, determining the behavioral data of animals for preset actions is solved, and the accuracy problem of traditional image processing methods is solved in identifying animal behaviors, achieving higher recognition accuracy.
Patent Information
- Application Number
- CN202411955159.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional image processing methods are susceptible to interference from factors such as image resolution and noise when identifying animal behavior, resulting in a decrease in recognition of recognition accuracy.
By obtaining the initial behavior video of the target animal, the target video clip including the preset action is extracted, and input it into the position prediction model, the position information of the target animal in the video frame is obtained, thereby determining its behavior data for the preset action.
This method improves the accuracy of animal behavior recognition and avoids the problem of inaccurate recognition caused by image details loss.
Smart Images

Figure CN120014697A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a behavior recognition method, device, computer equipment, storage medium and program product. Background Art
[0002] Generally speaking, different behavioral manifestations can be used to reflect the characteristics of different diseases, and then by studying different behavioral manifestations, the development of drugs to treat different diseases can be promoted. For example, itching is defined as an unpleasant feeling on the skin that leads to the desire to scratch. Generally speaking, scratching behavior is considered the gold standard for assessing the degree of itching. Accurate and rapid identification of scratching behavior plays an important role in promoting the study of itching mechanisms and genes, and the development of new antipruritic and immune drugs.
[0003] In traditional technology, different behaviors are usually identified by collecting animal behavior videos and using traditional image processing methods. For example, traditional image processing methods involve comparing the distance between the animal's hind limbs and back with a predefined threshold distance in each frame to identify scratching behavior. However, when processing complex images, traditional image processing methods may be interfered by factors such as image resolution and noise. These factors may cause the loss or deformation of image details, thereby affecting the accurate measurement of the distance between various parts of the animal's body, thereby reducing the accuracy of animal behavior recognition. Summary of the invention
[0004] Based on this, it is necessary to provide a behavior recognition method, device, computer equipment, storage medium and program product to improve the accuracy of animal behavior recognition in response to the above technical problems.
[0005] In a first aspect, the present application provides a behavior recognition method, the method comprising:
[0006] Obtaining the initial behavior video of the target animal within a preset period of time;
[0007] Extracting a target video segment including a preset action from the initial behavior video; wherein the preset action at least includes a scratching action;
[0008] Inputting the target video clip into a position prediction model to obtain position information of the target animal in each video frame in the target video clip; wherein the position information includes position information of a body part related to the preset action;
[0009] The behavior data of the target animal for the preset action is determined according to the position information of the target animal in each of the video frames in the target video segment.
[0010] In one embodiment, extracting a target video segment including a preset action from the initial behavior video includes:
[0011] Filtering the still video frames in the initial behavior video to obtain motion video clips;
[0012] A target video segment including a preset action is extracted from the motion video segment.
[0013] In one embodiment, filtering the still video frames in the initial behavior video to obtain the motion video clips includes:
[0014] Adjusting the format, size and frame rate of the initial behavior video to obtain an adjusted behavior video corresponding to the initial behavior video;
[0015] The still video frames in the adjustment behavior video are filtered to obtain motion video clips.
[0016] In one of the embodiments, extracting the target video segment including the preset action from the motion video segment includes:
[0017] A target video segment including a preset action is extracted from the motion video segment according to the difference between consecutive video frames in the motion video segment.
[0018] In one of the embodiments, extracting the target video segment including the preset action from the motion video segment includes:
[0019] The motion video segment is input into a motion prediction model to obtain a target video segment including a preset action.
[0020] In one embodiment, the position information includes position information of a body part performing a preset action;
[0021] The step of determining the behavior data of the target animal for the preset action according to the position information of the target animal in each of the video frames in the target video segment includes:
[0022] Acquiring position information of a reference body part of the target animal;
[0023] The behavior data of the target animal for the preset action is determined according to the position information of the body part performing the preset action and the position information of the reference body part.
[0024] In a second aspect, the present application further provides a behavior recognition device, the device comprising:
[0025] An acquisition module is used to acquire the initial behavior video of the target animal within a preset time period;
[0026] An extraction module, used to extract a target video segment including a preset action from the initial behavior video; wherein the preset action at least includes a scratching action;
[0027] A prediction module, used for inputting the target video clip into a position prediction model to obtain the position information of the target animal in each video frame in the target video clip; wherein the position information includes the position information of the body part related to the preset action;
[0028] A determination module is used to determine the behavior data of the target animal for the preset action according to the position information of the target animal in each of the video frames in the target video segment.
[0029] In a third aspect, the present application further provides a computer device, the computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:
[0030] Obtaining the initial behavior video of the target animal within a preset period of time;
[0031] Extracting a target video segment including a preset action from the initial behavior video; wherein the preset action at least includes a scratching action;
[0032] Inputting the target video clip into a position prediction model to obtain position information of the target animal in each video frame in the target video clip; wherein the position information includes position information of a body part related to the preset action;
[0033] The behavior data of the target animal for the preset action is determined according to the position information of the target animal in each of the video frames in the target video segment.
[0034] In a fourth aspect, the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0035] Obtaining the initial behavior video of the target animal within a preset period of time;
[0036] Extracting a target video segment including a preset action from the initial behavior video; wherein the preset action at least includes a scratching action;
[0037] Inputting the target video clip into a position prediction model to obtain position information of the target animal in each video frame in the target video clip; wherein the position information includes position information of a body part related to the preset action;
[0038] The behavior data of the target animal for the preset action is determined according to the position information of the target animal in each of the video frames in the target video segment.
[0039] In a fifth aspect, the present application further provides a computer program product, the computer program product comprising a computer program, which implements the following steps when executed by a processor:
[0040] Obtaining the initial behavior video of the target animal within a preset period of time;
[0041] Extracting a target video segment including a preset action from the initial behavior video; wherein the preset action at least includes a scratching action;
[0042] Inputting the target video clip into a position prediction model to obtain position information of the target animal in each video frame in the target video clip; wherein the position information includes position information of a body part related to the preset action;
[0043] The behavior data of the target animal for the preset action is determined according to the position information of the target animal in each of the video frames in the target video segment.
[0044] The above-mentioned behavior recognition method, device, computer equipment, storage medium and program product obtain the initial behavior video of the target animal within a preset time period; extract the target video segment including the preset action from the initial behavior video; wherein the preset action at least includes the scratching action; input the target video segment into the position prediction model to obtain the position information of the target animal in each video frame in the target video segment; wherein the position information includes the position information of the body part related to the preset action; and determine the behavior data of the target animal for the preset action according to the position information of the target animal in each video frame in the target video segment. The above-mentioned scheme, when identifying the behavior data of the target animal for the preset action, inputs the target video segment including the preset action into the position prediction model, and determines the behavior data of the target animal for the preset action according to the position information of the target animal in each video frame in the target video segment output by the model, and does not rely on the image processing method to perform preset behavior recognition, thereby avoiding the problem of inaccurate preset behavior recognition caused by the loss of image details. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A diagram of an application environment of a behavior recognition method in an embodiment;
[0046] Figure 2 A schematic diagram of a human-computer interaction interface in an embodiment;
[0047] Figure 3 A flowchart of a behavior recognition method in one embodiment;
[0048] Figure 4 This is a schematic diagram of the structure of a cage for storing target animals in one embodiment;
[0049] Figure 5 A schematic diagram of a process of extracting a target video segment including a preset action in one embodiment;
[0050] Figure 6 A schematic diagram of a process for acquiring motion video clips of multiple target animals in one embodiment;
[0051] Figure 7 A schematic diagram of a process of obtaining an adjustment behavior video in one embodiment;
[0052] Figure 8 A schematic diagram of a process for determining the behavior data of a target animal for a preset action in one embodiment;
[0053] Fig. 9 is a schematic diagram of a process of performing behavioral analysis on a target animal in one embodiment;
[0054] Fig.10 is a structural block diagram of a behavior recognition device in an embodiment;
[0055] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] The behavior recognition method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown in the figure, the hardware part includes an edge computing platform, a Mobile Industry Processor Interface (MIPI) camera, a power supply, a small display, a light source board, a detachable cage, and an external structure. The edge computing platform can be NVIDIA Jetson TX2.
[0058] The edge computing platform is a distributed computing architecture that aims to transfer data processing and analysis functions from traditional data centers or cloud computing environments to places closer to the data source, that is, "edge" devices. The edge computing platform can process, store and analyze data on devices close to the data generation point (such as sensors, IoT devices, smart terminals, etc.) without transmitting all data to remote data centers or cloud servers. Edge platforms can be, but are not limited to, other hardware platforms (such as NVIDIA Jetson, Raspberry Pi, Intel NUC, Google Coral, etc.) and edge computing solutions provided by cloud service providers (such as AWS IoT Greengrass, Microsoft Azure IoT Edge, etc.).
[0059] The software part includes human-computer interaction part, video acquisition part, data dimension reduction part and behavior parameter recognition part.
[0060] In the human-computer interaction part, an intuitive human-computer interaction interface is designed to facilitate user operation and real-time observation of the operating status of the edge computing platform. The edge computing platform is connected to the small display through the High Definition Multimedia Interface (HDMI) interface, and connected to the keyboard and mouse through the Universal Serial Bus (USB) or Bluetooth.
[0061] For example, see Figure 2 , Figure 2 A schematic diagram of a human-computer interaction interface is provided. The main functions of the human-computer interaction interface include real-time display of video streams, wherein the right side of the human-computer interaction interface displays the real-time captured animal behavior video stream, and the user can observe the dynamic behavior of the animal during the experiment. The human-computer interaction interface includes a control panel, and the user can preview, control the start, control the timing, or control the end of the experiment through the left control panel, and conveniently control the entire experimental process.
[0062] The video acquisition part is used to collect the behavioral video data of the target animal.
[0063] The data dimension reduction part is used to perform dimension reduction processing on the behavioral video data of the target animal to reduce the amount of data calculation.
[0064] The behavior recognition part is used to identify the behavior of the target animal based on the video data after dimensionality reduction processing.
[0065] In one embodiment, Figure 3As shown, a behavior recognition method is provided, which is described by applying the method to a server, and includes the following steps:
[0066] S301, obtaining an initial behavior video of a target animal within a preset time period.
[0067] Exemplarily, the target animal may be a rodent, such as a mouse. The preset time period may be determined based on the predicted behavior, such as a certain time period of a day, or a few days, etc. The initial behavior video is a behavior video of the target animal within the preset time period, and may include video clips of preset actions, or video clips of non-preset actions. For example, the initial behavior video of the target animal may be obtained every 5 minutes, and the obtained initial behavior video may be saved.
[0068] Optional, see Figure 4 , Figure 4 A schematic diagram of the structure of a cage for storing target animals is provided. The cage is made of five acrylic plates spliced together to form a 2×2 structure. Each cage has four sides and no top or bottom. The overall size can be set according to the size of the target animal, for example, it can be set to 24cm×24cm×37cm, and the internal size of each cage is 10cm×10cm×20cm. A mesh is provided on the top of the excrement recovery box, and the excrement of the target animal can fall into the collection bag below through the mesh. The cage and the excrement recovery box are detachable, which is convenient for capturing the target animal and cleaning the excrement after the experiment, while ensuring that the target animal has enough space to move around.
[0069] Four cameras can be used to shoot the target animals in four cages at 120FPS from the top view to obtain the initial behavior video of the target animals within a preset period of time. Each camera is connected to Jetson TX2. The top uses evenly distributed light emitting diodes (LEDs) as the light source, and a diffuser is provided below to diffuse the light into soft light. The diffuser is about 1 cm away from the top of the cage, which ensures ventilation and prevents the target animals from jumping out of the cage.
[0070] S302: extracting a target video segment including a preset action from the initial behavior video.
[0071] Exemplarily, the initial behavior video may be subjected to image processing to extract a target video segment including a preset action from the initial behavior video, or a model may be used to extract feature data of the initial behavior video, and based on the extracted feature data, a target video segment including a preset action may be extracted from the initial behavior video. The preset action may include at least a scratching action, which is used to analyze the scratching action.
[0072] S303: Input the target video clip into a position prediction model to obtain the position information of the target animal in each video frame in the target video clip.
[0073] The target video clips include bout video clips, wherein a complete preset action is called a bout. For example, the target animal licking or putting its hind paws on the floor is defined as a scratching termination action, and a complete scratching action is called a bout.
[0074] Furthermore, duration can correspond to the corresponding bout signal segment in the cropped motion video segment. Duration refers to the duration of the bout, that is, the time required for the scratching action. The number of touches can correspond to the number of contacts between the claws and the back or cheek in the cropped motion video segment. Touch can be understood as a single contact between the claws and the back of the neck or cheek during scratching.
[0075] Furthermore, the target video clip can be directly input into the position prediction model. The position prediction model extracts features of the target video clip, and can determine the exact key frame where the target animal's paw just touches the skin based on the touch information, and predict the position information of the target animal in each video frame in the target video clip based on the exact key frame where the target animal's paw just touches the skin.
[0076] The accurate key frame where the claw of the target animal just touches the skin can also be input into the position prediction model, which extracts features from the target video clip and predicts the position information of the target animal in each video frame in the target video clip based on the extracted feature data. The position information includes the position information of the body parts related to the preset action; for example, if the preset action is a scratching action, the body parts related to the preset action include the claws, ears, body and other parts of the target animal.
[0077] S304: Determine the behavior data of the target animal for a preset action according to the position information of the target animal in each video frame in the target video segment.
[0078] Furthermore, the relative position information of each body part of the target animal related to the preset action in each video frame in the target video clip can be determined to determine the behavior data of the target animal for the preset action. Among them, the behavior data of the target animal for the preset action can be understood as the specific location of the preset action, the duration of the action, and other data. For example, if the claws of the target animal are in contact with the body and last for a period of time, it can be determined that the target animal has made a scratching action.
[0079] The above behavior recognition method obtains the initial behavior video of the target animal within a preset time period; extracts the target video segment including the preset action from the initial behavior video; wherein the preset action at least includes the scratching action; inputs the target video segment into the position prediction model to obtain the position information of the target animal in each video frame in the target video segment; wherein the position information includes the position information of the body part related to the preset action; and determines the behavior data of the target animal for the preset action according to the position information of the target animal in each video frame in the target video segment. The above scheme, when identifying the behavior data of the target animal for the preset action, inputs the target video segment including the preset action into the position prediction model, and determines the behavior data of the target animal for the preset action according to the position information of the target animal in each video frame in the target video segment output by the model, without relying on traditional image processing methods to perform preset behavior recognition, thereby avoiding the problem of inaccurate recognition of preset behaviors caused by loss of image details.
[0080] In some optional implementations, in order to extract a target video segment including a preset action from an initial behavior video, static video frames in the initial behavior video may be first filtered out to reduce the amount of data processing.
[0081] For example, see Figure 5 , Figure 5 A schematic diagram of a process for extracting a target video segment including a preset action is provided, which specifically includes the following steps:
[0082] S501, filtering the still video frames in the initial behavior video to obtain motion video clips.
[0083] For example, the initial behavior video data can be differentially processed, and the content that does not contain motion information can be filtered out through corrosion and expansion operations, that is, the static video frames in the initial behavior video are filtered. Then, the differential signal is reconstructed to reduce the complexity of the behavior recognition data calculation while capturing the motion information and retaining the timing information.
[0084] S502: extracting a target video segment including a preset action from the motion video segment.
[0085] Furthermore, a target video segment including a preset action can be extracted from the motion video segment. For example, image analysis can be performed on the motion video segment to extract the target video segment including the preset action from the motion video segment. Alternatively, the motion video segment is input into a video extraction model to extract the target video segment including the preset action from the motion video segment.
[0086] In the embodiment of the present application, by filtering the still video frames in the initial behavior video, the data processing amount for extracting the motion video segments is reduced, and the efficiency of extracting the motion video segments can be improved.
[0087] In some optional implementations, when the number of target animals is at least two, if behavior prediction is required for multiple target animals, a camera can be used to obtain the initial behavior video of each target animal respectively, and then the initial behavior video of each target animal can be processed synchronously to speed up the efficiency of processing the initial behavior video of each target animal.
[0088] For example, see Figure 6 , Figure 6 A schematic diagram of a process for obtaining motion video clips of multiple target animals is provided, which specifically includes the following steps:
[0089] S601, adjusting the format, size and frame rate of the initial behavior video to obtain an adjusted behavior video corresponding to the initial behavior video.
[0090] For example, see Figure 7 , Figure 7 A schematic diagram of a process for obtaining an adjustment behavior video is provided.
[0091] First, the nvv4l2camerasrc element on the Jetson TX2 platform can be used to obtain the initial behavior video of the target animal from the camera. Further, the initial behavior video can be converted to UYVY format and the video size can be adjusted to 640×480 and the frame rate can be set to 120 frames / second in NVMM memory.
[0092] Then, the nvvidconv element can be used to convert each adjusted video into NV12 format and compress it to 320×240 size to reduce the amount of data required in the process of extracting motion video clips. Furthermore, the nvv4l2h264enc element can be used to further encode the video into H.264 format to obtain the adjusted behavior video corresponding to the initial behavior video.
[0093] S602: Filter the still video frames in the adjustment behavior video to obtain motion video clips.
[0094] For example, the adjusted behavior video data can be differentially processed, and the content that does not contain motion information can be filtered out through corrosion and expansion operations, that is, the static video frames in the adjusted behavior video are filtered. Then, the differential signal is reconstructed to reduce the complexity of the behavior recognition data calculation while capturing the motion information and retaining the timing information.
[0095] In the embodiment of the present application, by converting the format, size and frame rate of the initial behavior video, it is easier to perform filtering operations on still video frames, thereby improving the efficiency of obtaining motion video clips.
[0096] In some optional implementations, extracting the target video segment including the preset action can be achieved by:
[0097] According to the difference between consecutive video frames in the motion video segment, a target video segment including a preset action is extracted from the motion video segment.
[0098] Exemplarily, a differential signal of a motion video clip can be obtained, wherein the differential signal is used to characterize the difference between consecutive video frames in the motion video clip; and then the differential signal can be subjected to frequency domain analysis to distinguish between a target video clip containing a preset action and a target video clip not containing the preset action.
[0099] Alternatively, the motion video clip can be input into the motion prediction model, and the motion model can be used to predict the target video clip including the preset action. The motion prediction model can be any one of the long short-term memory network (Long Short-Term Memory, LSTM), YOLO model, or time series model (Time Series Model, TSM).
[0100] In the embodiments of the present application, two feasible methods for extracting target video segments including preset actions are provided, which enrich the means for extracting target video segments including preset actions and improve the convenience of extracting target video segments including preset actions.
[0101] In some optional implementations, the position information in the above embodiments may include position information of a body part that performs a preset action; for example, if the preset action is a scratching action, the body part that performs the preset action is a claw.
[0102] For example, see Figure 8 , Figure 8 A schematic diagram of a process for determining the behavior data of a target animal for a preset action is provided, which specifically includes the following steps:
[0103] S801, obtaining position information of a reference body part of a target animal.
[0104] Exemplarily, the position information of a reference body part of the target animal can be obtained, wherein the reference body part can be determined based on the body structure and preset action of the target animal. For example, if the target animal is a mouse and the preset action is a scratching action, the reference body part can be the body and the ear, which are used to distinguish the specific scratching part.
[0105] S802, determining the behavior data of the target animal for the preset action according to the position information of the body part performing the preset action and the position information of the reference body part.
[0106] Exemplarily, image analysis can be used to determine the relative position between position information of a body part that performs a preset action and position information of a reference body part, such as determining the relative position of a claw in the body; and then, based on the relative position between the position information of the body part that performs the preset action and position information of the reference body part, the behavioral data of the target animal with respect to the preset action can be determined, for example, determining the location where the target animal's action with respect to the preset action occurs.
[0107] Exemplarily, the site where the mouse's scratching action occurs can be divided into four areas: the left cheek, the right cheek, the left back of the neck, and the right back of the neck. Thus, the site where the mouse's scratching action occurs and the duration of a scratching action in one area can be determined based on the position of the mouse's claws, the position of the body's center of gravity, and the position of the ears.
[0108] In an embodiment of the present application, by determining the location where the target animal's action occurs for the preset action and data such as the duration of the action based on the location information of the body part performing the preset action and the location information of the reference body part of the target animal, it is convenient to analyze the target action for the preset action in more dimensions to improve the accuracy of the analysis.
[0109] In some optional implementations, see Fig. 9 , Fig. 9 A schematic diagram of a process for performing behavioral analysis on a target animal based on an initial behavioral video is provided, which specifically includes the following steps:
[0110] Step 1: Data dimensionality reduction processing.
[0111] Exemplarily, still video frames in the initial behavior video may be filtered to obtain motion video clips, thereby reducing the complexity of behavior recognition data calculation.
[0112] Step 2: Signal processing.
[0113] Exemplarily, a target video segment including a preset action can be extracted from a motion video segment. For example, a signal analysis method can be used to extract a target video segment including a preset action, for example, firstly, based on the difference between consecutive video frames in the motion video segment, a target video segment including a preset action is extracted from the motion video segment. The motion video segment can also be input into a motion prediction model, and the motion model is used to predict a target video segment including a preset action.
[0114] Step 3: Identification of bout, touch and duration.
[0115] The target video segment includes a video segment corresponding to bout. The duration can be used to correspondingly crop the signal segment of bout in the motion video segment. The exact key frame where the paw of the target animal just touches the skin can also be determined based on the touch information.
[0116] Step 4: Identify the behavioral data of the target animal for the preset action.
[0117] Furthermore, the accurate key frame where the target animal's paw just touches the skin can be input into the position prediction model to predict the position information of the target animal in each video frame in the target video clip; and the relative position between the position information of the body part performing the preset action and the position information of the reference body part can be determined through image analysis, thereby realizing the behavioral data recognition of the target animal for the preset action.
[0118] In the embodiment of the present application, a multi-channel parallel method is used to achieve efficient recognition of the preset behavior of the target animal. It is also possible to simultaneously achieve behavioral video recording and behavior recognition of multiple (such as four or six) target animals. In addition, the embodiment of the present application uses touch and duration to identify and count the areas where the preset behavior occurs, which will help to more comprehensively evaluate the behavioral data of the target animal for the preset behavior.
[0119] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0120] Based on the same inventive concept, the embodiment of the present application also provides a behavior recognition device for implementing the behavior recognition method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more behavior recognition device embodiments provided below can refer to the limitations of the behavior recognition method above, and will not be repeated here.
[0121] In one embodiment, Fig.10 As shown, a behavior recognition device is provided, comprising:
[0122] An acquisition module 10 is used to acquire an initial behavior video of a target animal within a preset period of time;
[0123] An extraction module 20 is used to extract a target video segment including a preset action from the initial behavior video; wherein the preset action at least includes a scratching action;
[0124] The prediction module 30 is used to input the target video clip into the position prediction model to obtain the position information of the target animal in each video frame in the target video clip; wherein the position information includes the position information of the body part related to the preset action;
[0125] The determination module 40 is used to determine the behavior data of the target animal for the preset action according to the position information of the target animal in each video frame in the target video segment.
[0126] The above-mentioned behavior recognition device obtains the initial behavior video of the target animal within a preset time period; extracts the target video segment including the preset action from the initial behavior video; wherein the preset action includes at least the scratching action; inputs the target video segment into the position prediction model to obtain the position information of the target animal in each video frame in the target video segment; wherein the position information includes the position information of the body part related to the preset action; and determines the behavior data of the target animal for the preset action according to the position information of the target animal in each video frame in the target video segment. The above-mentioned scheme, when identifying the behavior data of the target animal for the preset action, inputs the target video segment including the preset action into the position prediction model, and determines the behavior data of the target animal for the preset action according to the position information of the target animal in each video frame in the target video segment output by the model, and does not rely on image processing methods to perform preset behavior recognition, thereby avoiding the problem of inaccurate recognition of preset behaviors caused by loss of image details.
[0127] In one embodiment, the extraction module 20 specifically includes:
[0128] A filtering unit, used for filtering still video frames in the initial behavior video to obtain motion video clips;
[0129] The extraction unit is used to extract a target video segment including a preset action from the motion video segment.
[0130] In one embodiment, the filter unit is specifically used for:
[0131] The format, size and frame rate of the initial behavior video are adjusted to obtain an adjusted behavior video corresponding to the initial behavior video; and the still video frames in the adjusted behavior video are filtered to obtain a motion video clip.
[0132] In one embodiment, the extraction unit is specifically used for:
[0133] According to the difference between consecutive video frames in the motion video segment, a target video segment including a preset action is extracted from the motion video segment.
[0134] In one embodiment, the extraction unit is specifically used for:
[0135] The motion video clip is input into the motion prediction model to obtain a target video clip including a preset action.
[0136] In one embodiment, the position information includes position information of a body part performing a preset action; the determination module 40 is specifically used to:
[0137] Obtaining position information of a reference body part of a target animal; determining behavior data of the target animal for the preset action based on the position information of the body part that performs the preset action and the position information of the reference body part.
[0138] Each module in the above-mentioned behavior recognition device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0139] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.11 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store behavioral video data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a behavior recognition method is implemented.
[0140] Those skilled in the art will understand that Fig.11The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0141] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0142] Obtaining the initial behavior video of the target animal within a preset period of time;
[0143] Extracting a target video segment including a preset action from the initial behavior video; wherein the preset action at least includes a scratching action;
[0144] Inputting the target video clip into the position prediction model to obtain the position information of the target animal in each video frame in the target video clip; wherein the position information includes the position information of the body part related to the preset action;
[0145] According to the position information of the target animal in each video frame in the target video segment, the behavior data of the target animal for the preset action is determined.
[0146] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0147] The still video frames in the initial behavior video are filtered to obtain motion video clips; and the target video clips including the preset actions are extracted from the motion video clips.
[0148] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0149] The format, size and frame rate of the initial behavior video are adjusted to obtain an adjusted behavior video corresponding to the initial behavior video; and the still video frames in the adjusted behavior video are filtered to obtain a motion video clip.
[0150] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0151] According to the difference between consecutive video frames in the motion video segment, a target video segment including a preset action is extracted from the motion video segment.
[0152] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0153] The motion video clip is input into the motion prediction model to obtain a target video clip including a preset action.
[0154] In one embodiment, the position information includes position information of a body part that performs a preset action; when the processor executes the computer program, the following steps are also implemented:
[0155] Obtaining position information of a reference body part of a target animal; determining behavior data of the target animal for the preset action based on the position information of the body part that performs the preset action and the position information of the reference body part.
[0156] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the behavior recognition method described in any of the above embodiments are implemented.
[0157] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps of the behavior recognition method described in any of the above embodiments.
[0158] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0159] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0160] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0161] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the present application. It should be noted that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A behavior recognition method, characterized in that: The method comprises: Obtaining the initial behavior video of the target animal within a preset period of time; Extracting a target video segment including a preset action from the initial behavior video; wherein the preset action at least includes a scratching action; Inputting the target video clip into a position prediction model to obtain position information of the target animal in each video frame in the target video clip; wherein the position information includes position information of a body part related to the preset action; The behavior data of the target animal for the preset action is determined according to the position information of the target animal in each of the video frames in the target video segment.
2. The method according to claim 1, characterized in that The step of extracting a target video segment including a preset action from the initial behavior video comprises: Filtering the still video frames in the initial behavior video to obtain motion video clips; A target video segment including a preset action is extracted from the motion video segment.
3. The method according to claim 2, characterized in that The filtering process of the still video frames in the initial behavior video to obtain the motion video clips includes: Adjusting the format, size and frame rate of the initial behavior video to obtain an adjusted behavior video corresponding to the initial behavior video; The still video frames in the adjustment behavior video are filtered to obtain motion video clips.
4. The method according to claim 2, characterized in that: The step of extracting a target video segment including a preset action from the motion video segment comprises: A target video segment including a preset action is extracted from the motion video segment according to the difference between consecutive video frames in the motion video segment.
5. The method according to claim 2, characterized in that: The step of extracting a target video segment including a preset action from the motion video segment comprises: The motion video segment is input into a motion prediction model to obtain a target video segment including a preset action.
6. The method according to claim 1, characterized in that The position information includes position information of the body part performing the preset action; The step of determining the behavior data of the target animal for the preset action according to the position information of the target animal in each of the video frames in the target video segment includes: Acquiring position information of a reference body part of the target animal; The behavior data of the target animal for the preset action is determined according to the position information of the body part performing the preset action and the position information of the reference body part.
7. A behavior recognition device, characterized in that: The device comprises: An acquisition module is used to acquire the initial behavior video of the target animal within a preset time period; An extraction module, used to extract a target video segment including a preset action from the initial behavior video; wherein the preset action at least includes a scratching action; A prediction module, used for inputting the target video clip into a position prediction model to obtain the position information of the target animal in each video frame in the target video clip; wherein the position information includes the position information of the body part related to the preset action; A determination module is used to determine the behavior data of the target animal for the preset action according to the position information of the target animal in each of the video frames in the target video segment.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.