Image processing method, electronic device, and computer readable medium
By using image processing methods optimized through inter-frame difference and depth information, the problems of resource waste and high energy consumption of fixed frame rate cameras in animal photography are solved, and the camera achieves accurate frame rate switching and improved stability in different scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ADDX (BEIJING) TECH CO LTD
- Filing Date
- 2025-02-05
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, fixed frame rate cameras cannot accurately distinguish different behaviors when filming animals, resulting in resource waste and high energy consumption. Furthermore, they cannot accurately capture images when animals are moving at high intensity, missing details and increasing the camera damage rate.
By acquiring animal image sets and frame rate information, inter-frame difference processing is performed to determine motion information. Based on the motion information and frame rate, the shooting frame rate is switched. Combined with depth information and the GrabCut algorithm, image segmentation is optimized to precisely control the camera frame rate.
It improves the accuracy and stability of camera shooting in different scenarios, reduces energy consumption and damage rate, and optimizes resource utilization.
Smart Images

Figure CN119996604B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of computer technology, and more particularly to image processing methods, electronic devices, and computer-readable media. Background Technology
[0002] Currently, with the development of IoT technology, monitoring and filming pets or wild animals is becoming increasingly common. However, because monitoring and filming typically uses a fixed frame rate, this results in a significant waste of resources. For storing animal information using cameras that switch frame rates, the common approach is to use cameras with fixed frame rates to capture animal images.
[0003] However, the inventors discovered that when using the above method to store animal information from cameras based on frame rate switching, the following technical problems often arise:
[0004] Because it uses a fixed frame rate for image acquisition, it cannot accurately distinguish different animal behaviors and wastes a lot of storage resources. It also cannot accurately capture images when the animal is moving at a high intensity, missing a lot of details. When the animal is still, the images captured contain a lot of repetitive information, resulting in high energy consumption and damage rate of the camera.
[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0007] Some embodiments of this disclosure provide image processing methods, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the background section above.
[0008] In a first aspect, some embodiments of this disclosure provide an image processing method applied to camera shooting, including: acquiring a first animal image set of a target animal set and first shooting frame rate information; performing inter-frame difference processing on the first animal image set to obtain a motion information set of the first animal in the target animal set; and performing shooting frame rate switching processing based on the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information.
[0009] Optionally, the above method further includes: controlling the camera to capture a set of animal images of the target animal set under the shooting frame rate switching information; performing animal recognition on the set of animal images to obtain animal recognition information; and storing the animal recognition information and the set of animal images in the target storage terminal by updating them over time.
[0010] Optionally, the above-mentioned frame rate switching process based on the first animal's motion information set and the first shooting frame rate information to obtain shooting frame rate switching information includes: performing state information filtering processing on the first animal's motion information set to obtain filtered animal state information; and performing frame rate switching processing based on the filtered animal state information and the first shooting frame rate information to obtain shooting frame rate switching information.
[0011] Optionally, the above-mentioned frame rate switching process based on the first animal's motion information set and the first shooting frame rate information to obtain shooting frame rate switching information includes: performing intensity information filtering processing on the animal motion information set to obtain filtered animal motion intensity information; in response to determining that the filtered animal motion intensity information is not greater than a preset motion intensity threshold, performing frame rate switching processing on the first shooting frame rate information to obtain second shooting frame rate information, which is used as shooting frame rate switching information.
[0012] Optionally, the above-mentioned frame rate switching process based on the first animal's motion information set and the first shooting frame rate information to obtain shooting frame rate switching information includes: performing intensity information filtering processing on the animal motion information set to obtain filtered animal motion intensity information; in response to determining that the filtered animal motion intensity information is greater than the preset motion intensity threshold, performing frame rate switching processing on the first shooting frame rate information to obtain third shooting frame rate information, which is used as shooting frame rate switching information.
[0013] Optionally, the above-mentioned animal recognition of the animal image set to obtain animal recognition information includes: in response to determining that the animal image set is an image set based on a first preset shooting frame rate, performing image preprocessing on the animal image set to obtain a preprocessed animal image set; inputting the preprocessed animal image set into an animal feature extraction network included in the animal recognition model to obtain a first animal feature information set, wherein the animal recognition model further includes: a multi-scale fusion network and a head recognition output network, wherein the animal feature extraction network includes: multiple hybrid attention networks, multiple first feature extraction networks, multiple max pooling layers, multiple cascaded aggregation networks, and multiple lightweight aggregation networks; inputting the first animal feature information set into the multi-scale fusion network to obtain a second animal feature information set, wherein the multi-scale fusion network includes: multiple first feature extraction networks, multiple second feature extraction networks, multiple lightweight aggregation networks, and multiple content-aware prediction networks; inputting the second animal feature information set into the head recognition output network to obtain animal recognition information, wherein the head recognition output network includes: multiple third feature extraction networks.
[0014] Optionally, the above-mentioned animal recognition of the animal image set to obtain animal recognition information includes: in response to determining that the animal image set is an image set based on a second preset shooting frame rate, performing downsampling convolution processing on the animal image set to obtain a first animal pose feature information set; inputting the first animal pose feature information set into the feature extraction network included in the animal pose recognition model to obtain a second pose feature information set, wherein the animal pose recognition model further includes: a first branch parallel fusion network, a second branch parallel fusion network, a third branch parallel fusion network, and convolutional layers, and the feature extraction network includes: a residual convolutional network, A multi-path fusion convolutional network is used; the second pose feature information set is input into the first branch parallel fusion network to obtain a first multi-scale pose feature information set; the first multi-scale pose feature information set is input into the second branch parallel fusion network to obtain a second multi-scale pose feature information set; the second multi-scale pose feature information set is input into the third branch parallel fusion network to obtain a third multi-scale pose feature information set; the third multi-scale pose feature information set is input into the convolutional layer to obtain animal pose recognition information, which is used as animal recognition information to determine whether the animal pose recognition information of the target animal set is flight pose information.
[0015] Optionally, the first branch-parallel fusion network includes: a branch-parallel expansion network, multiple multi-scale feature extraction convolutional networks, and multiple channel spatial attention networks; and the input of the second pose feature information set into the branch-parallel fusion network to obtain the first multi-scale pose feature information set includes: inputting the second pose feature information set into the first branch-parallel expansion network to obtain a fourth multi-scale pose feature information set; inputting the fourth multi-scale pose feature information set into the multiple multi-scale feature extraction convolutional networks to obtain a fifth multi-scale pose feature information set; inputting the fifth multi-scale pose feature information set into the multiple channel spatial attention networks to obtain a sixth multi-scale pose feature information set; performing sampling and fusion processing on the sixth multi-scale pose feature information set to obtain a seventh multi-scale pose feature information set; and performing nonlinear processing on the seventh multi-scale pose feature information set to obtain the first multi-scale pose feature information set.
[0016] Optionally, before acquiring the first animal image set and the first shooting frame rate information of the target animal set, the method further includes: in response to detecting infrared signal information sent by a passive infrared sensor, performing signal detection processing on the infrared signal information to obtain infrared detection information; and controlling the camera to shoot animal images based on the infrared detection information to acquire the first animal image set of the target animal set within the shooting area.
[0017] Optionally, the above method further includes: in response to determining that the infrared detection information indicates that no motion information has been detected, controlling the camera to enter a standby state.
[0018] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the implementations of the first aspect above.
[0019] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0020] The above embodiments of this disclosure have the following beneficial effects: The image processing methods of some embodiments of this disclosure, when applied to camera shooting, can improve the accuracy of camera switching, improve the stability of the camera, and reduce the damage rate. Specifically, the high energy consumption and damage rate of cameras are due to the following reasons: Since image acquisition is performed at a fixed frame rate, it is impossible to accurately distinguish different animal behaviors and wastes a lot of storage resources; when the animal's movement intensity is high, the image cannot be accurately captured, resulting in the loss of many details; and when the animal is stationary, the animal images captured contain a lot of repetitive information, leading to high energy consumption and damage rate of the camera. Based on this, the image processing methods of some embodiments of this disclosure, when applied to camera shooting, can first acquire a first animal image set and a first shooting frame rate information for the target animal set. Here, the acquired first animal image set facilitates the subsequent determination of the motion information of the current target animal set, and the acquired first shooting frame rate information facilitates the subsequent switching of the camera's shooting frame rate. Then, inter-frame difference processing is performed on the first animal image set to obtain a motion information set for the first animal in the target animal set. Here, the motion information of the target animal can be accurately determined so as to subsequently determine the switching of the camera's shooting frame rate. Finally, based on the aforementioned motion information set of the first animal and the aforementioned first shooting frame rate information, shooting frame rate switching processing is performed to obtain shooting frame rate switching information. This improves the accuracy of the camera's shooting frame rate switching in different scenarios, reducing camera wear and power waste. Therefore, this image processing method, when applied to camera shooting, can precisely control the switching of the camera's shooting frame rate using the target animal's motion information. It can switch between different shooting frame rates under different conditions to capture animal images that better match the current scene. This means it can capture more details when the target animal is moving vigorously, reduce the collection of a large amount of repetitive information when the target animal is stationary, reduce camera performance and power consumption, and reduce the camera's damage rate. Attached Figure Description
[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0022] Figure 1 This is a schematic diagram of an application scenario of an image processing method according to some embodiments of the present disclosure;
[0023] Figure 2 This is a flowchart of some embodiments of the image processing method according to the present disclosure;
[0024] Figure 3This is a schematic diagram of animal images taken according to some embodiments of the image processing method of this disclosure;
[0025] Figure 4 This is a schematic diagram of a cascaded aggregation network in some embodiments of the image processing method according to the present disclosure;
[0026] Figure 5 This is a flowchart of some other embodiments of the image processing method according to the present disclosure;
[0027] Figure 6 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0028] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0029] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0033] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0034] Figure 1 This is a schematic diagram illustrating an application scenario in which an image processing method according to some embodiments of the present disclosure is applied to camera shooting.
[0035] exist Figure 1In the application scenario, the electronic device 101 can first acquire a first animal image set 102 and a first shooting frame rate information 103 for the target animal set. Then, it performs inter-frame difference processing on the first animal image set 102 to obtain a motion information set 104 for the first animal in the target animal set. Finally, based on the motion information set 104 and the first shooting frame rate information 103, it performs shooting frame rate switching processing to obtain shooting frame rate switching information 105.
[0036] It should be noted that the aforementioned electronic device 101 can be either hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the electronic device is software, it can be installed in the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here.
[0037] It should be understood that Figure 1 The number of electronic devices shown is merely illustrative. Any number of electronic devices can be used depending on the implementation requirements.
[0038] Continue to refer to Figure 2 The diagram illustrates a flow 200 of some embodiments of the image processing method according to the present disclosure, applied to camera capture. This image processing method, applied to camera capture, includes the following steps:
[0039] Step 201: Obtain the first animal image set and the first shooting frame rate information of the target animal set.
[0040] In some embodiments, the above image processing method is applied to the execution subject of the camera (e.g., ...). Figure 1 The electronic device 101 shown can acquire a first set of animal images and a first shooting frame rate information of the target animal set via a wired or wireless connection. The camera mentioned above can be a camera with shooting frame rate switching for capturing images of the target animal set. The target animals in the target animal set can be animals that are close to the feeder and captured by the camera. For example, the target animals can be birds. The first animal images in the first set of animal images can be images of the target animal set captured by the camera at the shooting frame rate corresponding to the first shooting frame rate information at the current moment. The first shooting frame rate information can be information about the number of images that the camera can capture per second when capturing images of the target animals at the current moment.
[0041] In some optional implementations of certain embodiments, before acquiring the first animal image set and the first shooting frame rate information of the target animal set, the method may further include the following steps:
[0042] The first step involves processing the infrared signal information emitted by the passive infrared sensor to obtain infrared detection information. The passive infrared sensor can be a sensor that determines the degree of movement of the target animal through infrared signals. The infrared signal information can be information about the infrared signals emitted and acquired by the passive infrared sensor. The infrared detection information can characterize the degree of movement of the target animal.
[0043] As an example, the aforementioned execution entity can first filter the infrared signal information to obtain filtered infrared signal information. Then, using a Fourier transform algorithm, it can perform spectral analysis on the filtered infrared signal information to obtain infrared frequency information, which serves as infrared detection information.
[0044] The second step involves controlling the camera to capture animal images based on the aforementioned infrared detection information, thereby obtaining a first set of animal images of the target animal group within the capture area. The capture area can be any area that the camera can photograph.
[0045] As an example, the aforementioned execution entity may, in response to determining that the infrared frequency corresponding to the aforementioned infrared detection information is greater than or equal to a preset frequency threshold, determine the current shooting frame rate information as the shooting frame rate of the aforementioned camera, and control the aforementioned camera to shoot animal images at the shooting frame rate corresponding to the aforementioned current shooting frame rate information, so as to obtain a first set of animal images of the aforementioned target animal within the shooting area.
[0046] Step 202: Perform inter-frame difference processing on the first animal image set to obtain the motion information set of the first animal in the target animal set.
[0047] In some embodiments, the executing entity may perform inter-frame difference processing on the first animal image set to obtain a motion information set for the first animal in the target animal set. The animal motion information in the motion information set can characterize the motion of the target animal. For example, the animal motion information set may include, but is not limited to, at least one of the following: stationary, flying, standing, and motion intensity. The first animal may be any animal in the target animal set.
[0048] As an example, the aforementioned executing entity can utilize a visual algorithm to determine the set of pixel differences between any two adjacent first animal images in the aforementioned first animal image set, specifically for the target animal set, as the animal motion information set. The aforementioned visual algorithm can be an algorithm that uses an object detection model to identify and track the position of the target animal.
[0049] In the process of adopting technical solutions to address the first technical problem mentioned above, the following second technical problem often arises: when determining the motion intensity of the target animal through the inter-frame pixel differences of the first animal image set, the background image and lighting conditions can affect the inter-frame difference calculation, leading to low accuracy and consequently causing performance degradation and damage to the camera. A conventional solution to this second technical problem is to use the GrabCut algorithm to determine the animal foreground image set and the animal background image set of the first animal image set. Then, the sum of the pixel differences between the animal foreground image set and the animal background image set is determined as the animal motion intensity information set. However, the aforementioned conventional solutions still suffer from the following technical problems: Since the GrabCut algorithm only utilizes foreground and background image segmentation based on significant texture and color differences between the first animal images of consecutive frames, the segmentation accuracy is low when the differences between foreground and background pixels are small or the background is complex. Furthermore, the GrabCut algorithm has a long computation time, increasing the image segmentation computation time. Additionally, the presence of a large amount of redundant information in the foreground and background pixels leads to a longer inter-frame difference computation time, resulting in lower accuracy in calculating animal motion intensity information. This leads to lower camera performance stability, wasted camera resources, and an increased camera damage rate. Considering the shortcomings of conventional solutions and combining the advantages and current technological status of the inventors' inter-frame difference algorithms, the following solution has been adopted:
[0050] In some optional implementations of certain embodiments, the above-mentioned inter-frame difference processing of the first animal image set to obtain an animal motion intensity information set for the target animal set may include the following steps:
[0051] First, for each first animal image in the first animal image set mentioned above, perform the following foreground / background segmentation steps:
[0052] Sub-step 1 involves determining the animal depth image of the first animal image. Each pixel value in the animal depth image represents the distance from the camera. The executing entity can determine the animal depth image for the first animal image using a depth sensor.
[0053] Sub-step 2 involves performing morphological filtering on the animal depth image to obtain a filtered animal depth image. In practice, the execution entity can utilize a morphological opening and closing operation algorithm to first convert the animal depth image to grayscale before performing morphological filtering to obtain the filtered animal depth image. It should be noted that the opening operation in the morphological opening and closing algorithm can remove some background noise from the animal depth image, resulting in a smoother contour. The closing operation in the morphological opening and closing algorithm can fill in broken parts of the contour in the animal depth image. Therefore, the morphological filtering process can effectively avoid image over-segmentation caused by irregular details and noise in the animal depth image.
[0054] Sub-step 3 involves performing a first image segmentation on the filtered animal depth image to obtain a segmented animal depth image. In practice, the executing entity can select at least one pixel value greater than or equal to a preset pixel threshold from the pixel values in the filtered animal depth image as the set of pixel values to be removed. The preset pixel threshold can be a pre-defined pixel threshold used for image segmentation. For example, the preset pixel threshold could be 160. Removing the set of pixel values to be removed from the pixel values yields the segmented animal depth image.
[0055] Sub-step 4: The segmented animal depth image and the first animal image are fused to obtain the fused animal image.
[0056] Sub-step 5 involves performing superpixel segmentation on the fused animal image to obtain a superpixel-segmented animal image. This superpixel-segmented animal image can be a block-shaped image. The superpixel segmentation process can be a superpixel segmentation process using the SLIC (simple linear iterative cluster) algorithm that incorporates depth information. The superpixel distance value in the SLIC algorithm can be a distance value that includes the depth information, color information, and spatial location information of the pixel. This superpixel distance value can be obtained through the following steps: First, determine the square of the difference between the pixel values of any two corresponding pixels in the fused animal image, as the pixel distance value. Second, determine the sum of the squares of the differences in the horizontal axis values, the vertical axis values, and the depth values of any two pixels, as the depth distance value. Third, determine the sum of the squares of the differences in brightness, the first color, and the second color of any two pixels, as the color distance value. The first color can be color A in the CIELAB (CIELab color model), where color A includes colors ranging from dark green to gray to bright pink. The second color mentioned above can be color B in CIELAB colors, where color B includes colors ranging from bright blue to gray to yellow. Then, the positional distance between any two pixels is determined. Finally, the superpixel distance is determined by the square root of the sum of the ratios of the pixel distance, depth distance, color distance, and the square of a preset threshold, and the ratio of the positional distance to the square of the clustering distance. The preset threshold can be any value in the range [1, 40]. The clustering distance can be the square root of the number of pixels in the fused animal image and the number of clustered pixel centers in the SLIC algorithm's clustering pixel center set.
[0057] Sub-step 6 involves constructing a graph model for the superpixel-segmented animal image to obtain an undirected weighted graph for the superpixel-segmented animal image. Nodes in the undirected weighted graph can be superpixels in the superpixel-segmented animal image. Edges in the undirected weighted graph can be edges between superpixels. The weight values corresponding to the edges in the undirected weighted graph can be base e, with the exponent being the negative of the sum of the ratios of the differences in color values between any two pixels to a preset color threshold, the ratio of the differences in depth values to a preset depth threshold, and the ratio of the differences in the normalized average of the surface normals of the superpixels to a preset superpixel threshold.
[0058] Sub-step 7: Based on the aforementioned undirected weighted graph, perform foreground-background prior processing on the superpixel-segmented animal image to obtain an animal background saliency image and an animal foreground saliency image. The animal foreground saliency image can be a background saliency map obtained by arbitrarily selecting the upper, lower, left, and right boundaries as background query points on the edge background corresponding to the fused animal image and applying a manifold ranking algorithm. The edge background can be an image outside the region selected by the user in the fused animal image. The animal foreground saliency image can be a foreground saliency map obtained by using the center moments of the fused animal image as foreground query points and applying a manifold ranking algorithm.
[0059] As an example, the aforementioned execution entity can utilize popular ranking algorithms to perform foreground-background prior processing on the superpixel segmented animal image based on the aforementioned undirected weighted graph, thereby obtaining an animal background salient image and an animal foreground salient image.
[0060] Sub-step 8: Perform image fusion on the above animal foreground salient image and the above animal background salient image to obtain the fused animal salient image.
[0061] Sub-step 9 involves using the GrabCut algorithm, which incorporates depth information, to segment the fused animal salient image, resulting in an animal foreground image and an animal background image. The energy function in the GrabCut algorithm can be an energy function that incorporates depth information as a constraint term within the original GrabCut algorithm's energy function. This energy function can be achieved by adding a first weight value multiplied by the region data term to the original energy function, and adding a second weight value and a smoothing term incorporating depth information to the smoothing term. The first weight value can be obtained through the following steps: First, determine the KL (Kullback-Leibler divergence) values of the foreground Gaussian mixture model corresponding to the color information and the background Gaussian mixture model corresponding to the color information, as the first KL divergence value. Second, determine the KL divergence values of the foreground Gaussian mixture model corresponding to the depth information and the background Gaussian mixture model corresponding to the depth information, as the second KL divergence value. Then, determine the sum of the first and second KL divergence values as the third KL divergence value. Finally, the ratio of the first KL divergence function value to the third KL divergence function value is determined as the first weight value. The aforementioned second weight value can be obtained through the following steps: First, determine the smoothing term function value that fuses depth information and superpixel transparency information, as the depth smoothing term function value. Second, determine the sum of the function value corresponding to the original smoothing term function and the depth smoothing term function value, as the smoothing term function value. Finally, the ratio of the function value corresponding to the original smoothing term function to the smoothing term function value is determined as the second weight value.
[0062] The second step is to determine the set of foreground pixel differences between the obtained animal foreground images and the previous frame animal foreground images, and to determine the set of background pixel differences between the obtained animal background images and the previous frame animal background images.
[0063] The third step is to perform a weighted summation of each foreground pixel difference in the foreground pixel difference set and the corresponding background pixel difference in the background pixel difference set to obtain an inter-frame pixel difference set, which serves as the animal motion intensity information set.
[0064] The above technical solution and its related content, combined with step "step 203" as an inventive point of this disclosure, solve the second technical problem mentioned in the background: "Because the GrabCut algorithm only uses the foreground and background images with significant texture and color differences between the first animal images of the preceding and following frames for image segmentation, the segmentation accuracy is low when the difference between the foreground and background pixels is small or the background is complex. In addition, the operation time of the GrabCut algorithm is long, which increases the segmentation operation time of the image segmentation. Furthermore, there is a large amount of redundant information in the foreground and background pixels, which makes the inter-frame difference operation time long. The accuracy of animal motion intensity information calculation is low, resulting in low stability of camera performance, waste of a large amount of camera resources and increased damage rate of the camera." Factors leading to low camera performance stability, wasted camera resources, and increased camera damage rate are often as follows: Because the GrabCut algorithm only utilizes foreground and background image segmentation based on significant texture and color differences between the first and last animal images in consecutive frames, the segmentation accuracy is low when the difference between foreground and background pixels is small or the background is complex. Furthermore, the GrabCut algorithm has a long computation time, increasing the segmentation computation time. Additionally, the presence of a large amount of redundant information in foreground and background pixels results in a longer inter-frame difference computation time, leading to lower accuracy in calculating animal motion intensity information. Solving these factors can improve camera performance stability, reduce wasted camera resources, and lower the camera damage rate. To achieve this, this disclosure first performs grayscale conversion, filtering, and coarse segmentation on the acquired animal depth image. This reduces the area of background pixels while preserving foreground information, thus reducing computational resources and shortening computation time during subsequent image segmentation. Secondly, superpixel segmentation is performed on the fused animal image, which integrates depth information and the first animal image. This allows for the formation of superpixels from sets of highly similar pixels, reducing computation time. Furthermore, the fusion of depth information improves the accuracy of superpixel segmentation. Next, foreground and background prior processing based on depth information is applied to the fused animal image to obtain a salient animal image that incorporates depth information, highlighting the differences between foreground and background pixels. Then, the GrabCut algorithm, which integrates depth information, is used to segment the fused salient animal image, improving segmentation accuracy and reducing computation time. This allows for the reduction of redundant information between foreground and background pixels during animal image compression, minimizing waste of limited storage space. Finally, by using precisely segmented foreground and background image sets, a more accurate set of animal motion intensity information is determined. This enables precise control of camera frame rate switching, reducing camera resource waste, improving camera performance stability, and decreasing camera damage rates.
[0065] Step 203: Based on the motion information set of the first animal and the first shooting frame rate information, perform shooting frame rate switching processing to obtain shooting frame rate switching information.
[0066] In some embodiments, the executing entity may perform a shooting frame rate switching process based on the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information. The shooting frame rate switching information may be information obtained after changing the current shooting frame rate information.
[0067] As an example, the aforementioned executing entity may, in response to determining that intensity information representing flight posture exists in the aforementioned animal motion intensity information set, perform frame rate enhancement processing on the shooting frame rate corresponding to the aforementioned current shooting frame rate information to obtain enhanced frame rate information, which serves as shooting frame rate switching information. The enhanced frame rate information may be frame rate information corresponding to the aforementioned animal motion intensity information, obtained through statistical analysis of a large amount of shooting frame rate and motion intensity information data.
[0068] In some optional implementations of certain embodiments, the above-mentioned frame rate switching process based on the first animal's motion information set and the first shooting frame rate information to obtain shooting frame rate switching information may include the following steps:
[0069] The first step is to perform state information filtering on the aforementioned set of movement information of the first animal to obtain filtered animal state information. The filtered animal state information can be information from the aforementioned set of movement information that falls within a movement state. The movement state can include, but is not limited to, at least one of the following: stationary, flying, standing, walking.
[0070] The second step involves performing a shooting frame rate switching process based on the filtered animal status information and the first shooting frame rate information to obtain shooting frame rate switching information.
[0071] As an example, the aforementioned executing entity may, in response to determining that the filtered animal state information represents a state of motion, perform frame rate switching processing on the first shooting frame rate information to obtain fourth shooting frame rate information, which serves as the shooting frame rate switching information. The fourth shooting frame rate information may be information on shooting frame rates capable of accurately capturing the motion details of the first animal. This fourth shooting frame rate information may be a critical value for the motion state obtained through regression analysis of a large amount of collected animal motion state information and the camera's shooting frame rate.
[0072] In some optional implementations of certain embodiments, the above-mentioned frame rate switching process based on the first animal's motion information set and the first shooting frame rate information to obtain shooting frame rate switching information may include the following steps:
[0073] The first step is to filter the above-mentioned motion information set by intensity to obtain filtered motion intensity information. The filtered motion intensity information can be the information indicating the highest motion intensity among the above-mentioned sets.
[0074] The second step involves, in response to the determination that the motion intensity information after the above filtering is not greater than a preset motion intensity threshold, performing a shooting frame rate switching process on the first shooting frame rate information to obtain second shooting frame rate information, which serves as the shooting frame rate switching information. The preset motion intensity threshold can be a pre-set value. This preset motion intensity threshold can be a critical value between varying motion intensity levels, statistically derived by collecting a large amount of motion intensity information and shooting frame rates. The second shooting frame rate information can be pre-set information for shooting the target animal at a lower motion intensity. For example, the first shooting frame rate information could be 10 frames per second.
[0075] In some optional implementations of certain embodiments, the above-mentioned frame rate switching process based on the first animal's motion information set and the first shooting frame rate information to obtain shooting frame rate switching information may include the following steps:
[0076] The first step is to filter the intensity information of the above-mentioned motion information set to obtain the filtered motion intensity information.
[0077] The second step involves, in response to the determination that the motion intensity information after filtering is greater than the preset motion intensity threshold, switching the shooting frame rate of the first shooting frame rate information to obtain the third shooting frame rate information, which serves as the shooting frame rate switching information. The third shooting frame rate information can be a pre-set frame rate that clearly captures the pose of the target animal. For example, the first shooting frame rate information could be 60 frames per second. It should be noted that switching different shooting frame rates for different motion intensities can accommodate the needs of different scenarios, thereby reducing the camera's operating costs and energy consumption, improving the camera's operational stability, and meeting the needs for target animal recognition in different scenarios.
[0078] The above embodiments of this disclosure have the following beneficial effects: The image processing methods of some embodiments of this disclosure, when applied to camera shooting, can improve the accuracy of camera switching, improve the stability of the camera, and reduce the damage rate. Specifically, the high energy consumption and damage rate of cameras are due to the following reasons: Since image acquisition is performed at a fixed frame rate, it is impossible to accurately distinguish different animal behaviors and wastes a lot of storage resources; when the animal's movement intensity is high, the image cannot be accurately captured, resulting in the loss of many details; and when the animal is stationary, the animal images captured contain a lot of repetitive information, leading to high energy consumption and damage rate of the camera. Based on this, the image processing methods of some embodiments of this disclosure, when applied to camera shooting, can first acquire a first animal image set and a first shooting frame rate information for the target animal set. Here, the acquired first animal image set facilitates the subsequent determination of the motion information of the current target animal set, and the acquired first shooting frame rate information facilitates the subsequent switching of the camera's shooting frame rate. Then, inter-frame difference processing is performed on the first animal image set to obtain a motion information set for the first animal in the target animal set. Here, the motion information of the target animal can be accurately determined so as to subsequently determine the switching of the camera's shooting frame rate. Finally, based on the aforementioned motion information set of the first animal and the aforementioned first shooting frame rate information, shooting frame rate switching processing is performed to obtain shooting frame rate switching information. This improves the accuracy of the camera's shooting frame rate switching in different scenarios, reducing camera wear and power waste. Therefore, this image processing method, when applied to camera shooting, can precisely control the switching of the camera's shooting frame rate using the target animal's motion information. It can switch between different shooting frame rates under different conditions to capture animal images that better match the current scene. This means it can capture more details when the target animal is moving vigorously, reduce the collection of a large amount of repetitive information when the target animal is stationary, reduce camera performance and power consumption, and reduce the camera's damage rate.
[0079] Further reference Figure 3 The diagram illustrates a flow 300 of another embodiment of the image processing method according to the present disclosure, applied to camera capture. This image processing method, applied to camera capture, includes the following steps:
[0080] Step 301: Obtain the first animal image set and the first shooting frame rate information of the target animal set.
[0081] Step 302: Perform inter-frame difference processing on the first animal image set to obtain the motion information set of the first animal in the target animal set.
[0082] Step 303: Based on the motion information set of the first animal and the first shooting frame rate information, perform shooting frame rate switching processing to obtain shooting frame rate switching information.
[0083] In some embodiments, the specific implementation of steps 301-303 and the resulting technical effects can be found in [reference needed]. Figure 2 Steps 201-203 in the corresponding embodiments will not be repeated here.
[0084] Step 304: Control the camera to capture a set of animal images of the target animal group under the shooting frame rate switching information.
[0085] In some embodiments, the executing entity can control the camera to capture a set of animal images of the target animal group under the aforementioned frame rate switching information. The animal images in the set can be images of the target animal group captured by the camera at the frame rate corresponding to the aforementioned frame rate switching information. The animal images can be as follows: Figure 4 The diagram shown is shown in the image.
[0086] Step 305: Perform animal identification on the set of animal images to obtain animal identification information.
[0087] In some embodiments, the executing entity may perform animal identification on the aforementioned set of animal images to obtain animal identification information. This animal identification information may be information characterizing the posture and category of the target animal set. For example, the animal identification information may include, but is not limited to, at least one of the following: target animal category information, posture information, and animal key point information.
[0088] As an example, the aforementioned execution entity can input the aforementioned set of animal images into the object detection model to obtain animal identification information. The object detection model can be a pre-trained SSD (Single Shot MultiBox Detector) with a residual network or VisionTransformer as the backbone network.
[0089] In some optional implementations of certain embodiments, the above-mentioned animal identification of the animal image set to obtain animal identification information may include the following steps:
[0090] The first step involves preprocessing the animal image set to obtain a preprocessed animal image set, in response to determining that the animal image set is based on a first preset frame rate. The first preset frame rate can be a frame rate sufficient to capture the animal's movement when its motion is relatively gentle, without requiring a high frame rate. For example, the first preset frame rate could be the second frame rate information mentioned above. The image preprocessing can include: mosaic data augmentation, adaptive anchor box calculation, and adaptive image scaling. The mosaic data augmentation can be a data augmentation process that randomly scales, crops, and arranges four animal images from the animal image set and then stitches them together to generate a new image. The adaptive anchor box calculation can be a calculation that uses a clustering algorithm to determine the location of the target action. The adaptive image scaling can be a process that scales the image size of the animal image set to an image size that conforms to the input size of the animal recognition model.
[0091] The second step involves inputting the preprocessed animal image set into the animal feature extraction network included in the animal recognition model to obtain a first animal feature information set. This animal recognition model further includes a multi-scale fusion network and a head recognition output network. The animal feature extraction network comprises multiple first feature extraction networks, multiple max-pooling layers, multiple cascaded aggregation networks, multiple lightweight aggregation networks, and a pyramid pooling network. The animal recognition model can be a neural network model that identifies and outputs the category of animals in the input preprocessed animal image set. The animal feature extraction network can be a network that performs multi-scale feature extraction on the input preprocessed animal image set to output multi-scale animal feature information. The multi-scale fusion network can be a network that fuses multi-scale animal feature information to obtain feature information containing more detailed information. The head recognition output network can be a network that performs target detection, classification, and regression prediction on the fused feature information output by the multi-scale fusion network to output the category of the target animal. The first feature extraction network can be a network that includes convolutional layers with 3*3 kernels and a stride of 2, BN (Batch Normalization) and LeakyReLU (Leaky Rectified Linear Unit).
[0092] The aforementioned lightweight aggregation networks can be networks with residual structures that perform parallel feature extraction on the input feature information through two channels. These lightweight aggregation networks can include multiple second feature extraction networks and multiple third feature extraction networks. The second feature extraction networks can be networks including convolutional layers with 3x3 kernels and a stride of 1, BN layers, and the LeakyReLU activation function. The third feature extraction networks can be networks including convolutional layers with 1x1 kernels and a stride of 1, BN layers, and the LeakyReLU activation function. The aforementioned lightweight aggregation networks can extract features from the input animal feature information through the following steps: First, the input animal feature information is input into the third feature extraction network included in the first branch to obtain first animal input feature information. This first animal input feature information can be information representing the features of the animal image in the form of a graph. Second, the first animal input feature information is sequentially input into the two second feature extraction networks included in the first branch to obtain second and third animal input feature information. Next, the input animal feature information is fed into the third feature extraction network included in the second branch to obtain the fourth animal input feature information. Then, channel features are concatenated on the first, second, third, and fourth animal input feature information to obtain input channel concatenated feature information. Finally, the input channel concatenated feature information is fed into the third feature extraction network to obtain the fifth animal input feature information.
[0093] The cascaded aggregation network mentioned above can be a network that expands channels through multiple group convolutions and performs channel shuffling and merging. The cascaded aggregation network can be constructed by adding a hybrid attention network to the third and second feature extraction networks in the first branch of the lightweight aggregation network, replacing the second feature extraction network with a third feature extraction network, adding a channel shuffling network in the middle, and adding a residual connection network to the first branch. The network structure diagram of the cascaded aggregation network is shown below. Figure 5 As shown. Figure 5 The two branches in the algorithm split the input feature channels into equal parts. The upper branch preserves the original features to avoid gradient vanishing due to over-attention, while the lower branch uses the idea of dense residual networks to extract spatial and channel features. Figure 5The ChannelShuffle network in the above context can be a network that rearranges the channels between groups of input feature information to enhance the information interaction between channels. The hybrid attention network mentioned above can be a network that introduces feature focusing at both the spatial and channel levels. The hybrid attention network can perform feature extraction through the following steps: First, the input animal feature information is input into SENet (Squeeze-and-Excitation Networks) to obtain animal channel feature information. Second, the above animal channel feature information is input into a channel global average pooling network to obtain animal channel attention feature information. Specifically, the channel global average pooling network can be a network that inputs the input animal channel feature information into a global average pooling layer to perform global average pooling along the channel direction, then passes it through a fully connected layer and uses a sigmoid activation function to input the animal channel attention feature information. Third, the above animal channel attention feature information and the input animal feature information are globally weighted along the channel direction to obtain animal channel weighted feature information. Next, the weighted animal channel features are input into a max-pooling layer and an average-pooling layer, respectively, to obtain animal max-pooling and average-pooling features. Then, these animal max-pooling and average-pooling features are concatenated to obtain concatenated animal pooling features. Next, these concatenated animal pooling features are input into a 1x1 convolutional layer and a Sigmad activation function to obtain animal spatial attention features. The 1x1 convolutional layer can be a single-channel convolutional layer that reduces the dimensionality of the concatenated animal pooling features. Finally, the animal spatial attention features are multiplied by the input animal features to obtain animal spatial channel features.
[0094] The aforementioned pyramid pooling network can be a network that uses different pooling kernels to perform pooling operations to fuse feature information at different scales. The pyramid pooling network performs feature pooling processing through the following steps: First, the input animal feature information is input into a second feature extraction network to obtain input feature information. Second, the input feature information is input into a branch pooling network and a branch convolutional network respectively to obtain pooled feature information and convolutional feature information. Specifically, the branch pooling network can be a branch network that performs max pooling with kernels of 5*5, 9*9, and 13*13 in parallel on the input feature information, then concatenates the features before inputting them into the second feature extraction network for feature extraction. The branch convolutional network can be a branch network that inputs the input feature information into the second feature extraction network. Finally, the pooled feature information and convolutional feature information are concatenated using channel features and then input into the second feature extraction network to obtain animal pooled feature information.
[0095] The third step involves inputting the first animal feature information set into the multi-scale fusion network to obtain the second animal feature information set. The multi-scale fusion network includes multiple first feature extraction networks, multiple second feature extraction networks, multiple lightweight aggregation networks, and multiple content-aware prediction networks. The content-aware prediction network (CARAFE, Content-Aware ReAssembly of Features) can be a network that upsamples features by aggregating contextual information within the receptive domain and using adaptive kernels. The content-aware prediction network can include a network composed of a kernel prediction module and a content-aware feature reconstruction module. The content-aware feature reconstruction module can consist of three parts: a channel compression module, a content encoder, and a kernel normalization module. The channel compression module can be a compression module that uses 1*1 convolution to reduce the number of channels in the input feature map. The content encoder can be a module that takes the compressed feature map as input and encodes the content to generate a reconstructed kernel. The kernel normalization module is a module that applies a Softmax function to activate each reconstructed kernel. The aforementioned content-aware reconstruction module weights each target reconstruction region with the weights obtained from the kernel prediction module, and then concatenates all the reconstructed target regions. For example, if the animal feature information input to the content-aware prediction network is C*W*H, then the output of the aforementioned content-aware prediction network can be C*aW*aH. C can represent the number of channels. W can represent the feature length. H can represent the feature width. a can represent the upsampling ratio.
[0096] The fourth step involves inputting the second set of animal feature information into the head recognition output network to obtain animal identification information. The head recognition output network includes multiple third feature extraction networks. The animal identification information may be information about the animal category label to which the target animal belongs.
[0097] In some optional implementations of certain embodiments, the above-mentioned animal identification of the animal image set to obtain animal identification information may include the following steps:
[0098] The first step involves, in response to determining that the aforementioned animal image set is an image set based on a second preset shooting frame rate, performing downsampling convolution processing on the aforementioned animal image set to obtain an animal pose feature information set. The second preset shooting frame rate can be a shooting frame rate where the target animal's movement is relatively intense, requiring a higher shooting frame rate to capture more detailed pose information. For example, the second preset shooting frame rate can be the aforementioned third shooting frame rate information. The animal pose feature information in the aforementioned animal pose feature information set can be the pose information of the target animal in the animal images. The aforementioned downsampling convolution processing can be a downsampling process where the aforementioned animal image set is sequentially input into two pose recognition convolutional networks. The aforementioned pose recognition convolutional network can be a network including a convolutional layer with a 3*3 kernel, a stride of 2, edge padding of 1, and 64 output channels, a BN layer, and a ReLU activation function.
[0099] The second step involves inputting the aforementioned animal posture feature information set into the feature extraction network of the animal posture recognition model to obtain a second posture feature information set. This animal posture recognition model further includes: a first-branch parallel fusion network, a second-branch parallel fusion network, a third-branch parallel fusion network, and convolutional layers. The feature extraction network includes: a residual convolutional network and a multi-path fusion convolutional network. The animal posture recognition model can be a deep neural network model that identifies the animal posture of an input animal image and outputs the animal's posture information. The feature extraction network can be a network that extracts features from the input first animal posture feature information at different levels and scales. The residual convolutional network can be a network with a residual structure, including a first posture feature extraction network, a second posture feature extraction network, and a third posture feature extraction network as the first branch, and a third posture feature extraction network as the second branch. The first posture feature extraction network can be a network including a convolutional layer with a 1*1 kernel, a stride of 1, and 64 output channels, a BN layer, and a ReLU activation function. The second pose feature extraction network can be a network comprising a standard convolutional layer with a 3x3 kernel, a convolutional layer with a 1x3 kernel capturing features in the horizontal direction, and a convolutional layer with a 3x1 kernel capturing features in the vertical direction, and then performing feature concatenation, a batch normalization (BN) layer, and a ReLU activation function on the pose feature information obtained from the three convolutional layers. The third pose feature extraction network can be a network comprising a convolutional layer with a 1x1 kernel, a stride of 1, and 256 output channels, a BN layer, and a ReLU activation function. The multi-path fusion convolutional network can be a network after removing the second branch from the residual convolutional network. The first-branch parallel fusion network can be a network that adds an animal pose feature information extraction branch and performs multi-scale pose feature fusion. The initial feature information of the added branch in the first-branch parallel fusion network can be feature information whose size is half the size of the second pose feature information, obtained by outputting a convolutional layer with a 3x3 kernel, a stride of 1, and twice the number of output channels as the previous branch, and a BN layer. The aforementioned second-branch parallel fusion network can be a network with one more feature extraction branch than the aforementioned first-branch parallel fusion network, and the initial feature information size of the additional branch can be one-quarter of the size of the aforementioned second pose feature information. The aforementioned third-branch parallel fusion network can be a network with one more feature extraction branch than the aforementioned second-branch parallel fusion network, and the initial feature information size of the additional branch can be one-eighth of the size of the aforementioned second pose feature information. The aforementioned convolutional layer can be a convolutional layer with a 1*1 kernel, a stride of 2, an edge padding of 1, and 17 output channels. The aforementioned animal pose recognition information can be information determining whether the target animal is flying.
[0100] The third step involves inputting the second attitude feature information set into the first branch parallel fusion network to obtain the first multi-scale attitude feature information set. This first multi-scale attitude feature information set may include attitude feature information at different scales.
[0101] The fourth step is to input the first multi-scale pose feature information set into the second branch parallel fusion network to obtain the second multi-scale pose feature information set.
[0102] The fifth step is to input the second multi-scale pose feature information set into the third branch parallel fusion network to obtain the third multi-scale pose feature information set.
[0103] The sixth step involves inputting the aforementioned third multi-scale pose feature information set into the convolutional layer to obtain animal pose recognition information, which is then used as animal identification information to determine whether the animal pose recognition information of the target animal set is flight pose information. The aforementioned animal pose recognition information can be information about the pose of the identified target animal.
[0104] In some optional implementations of certain embodiments, the first branch parallel fusion network mentioned above includes: a branch parallel expansion network, multiple multi-scale feature extraction convolutional networks, and multiple channel spatial attention networks. The first branch parallel expansion network may be a network including an added pose feature extraction branch network that uses a second pose feature information set downsampled by 2x as initial pose information, and an original branch network that extracts the second pose feature information from the first pose feature information. The 2x downsampling can be performed using convolutional layers with 3x3 kernels, a stride of 1, and twice the number of output channels as the previous branch. The multi-scale feature extraction convolutional networks in the multiple multi-scale feature extraction convolutional networks can be networks that combine pose feature information using multiple convolutional layers with different kernels to extract richer feature information and reduce the number of parameters. The multi-scale feature extraction convolutional networks can perform feature extraction through the following steps: First, the input pose feature information is input into the pose feature information after passing through a 3x3 convolutional layer and a 3x3 dilated convolutional layer, respectively, and the features are added together to obtain the added pose feature information. Then, global average pooling is performed on the summed pose feature information to obtain global pose feature information of length L. This global pose feature information is then input into a fully connected layer for aggregation and compression to obtain compressed pose feature information. Next, the compressed pose feature information is input into two fully connected layers respectively to obtain two tiled pose feature information of length L. These two tiled pose feature information are then input into a Softmax activation function to obtain two pose weight values. Finally, these two pose weight values are weighted and summed with the corresponding pose feature information input through convolutional and dilated convolutional layers. The channel spatial attention network in the above-mentioned multi-channel spatial attention network can be a network that extracts channel attention and spatial attention from pose feature information to suppress noise in animal images and improve target detection accuracy. The channel spatial attention network can include both a channel attention module and a spatial attention module. The aforementioned channel attention module can be a network that obtains two 1*1*C feature maps by performing global max pooling and global average pooling on the input pose feature information, then inputs the two 1*1*C pose feature information into a multilayer perceptron (MLP) with shared weights, sums the two 1*1*C pose feature information element by element, and outputs the channel weight coefficients. Finally, the channel weight coefficients and the input pose feature information are multiplied to obtain the output pose feature information.The aforementioned spatial attention module can perform channel-based max pooling and average pooling on the output pose feature information from the channel attention module to obtain pooled pose feature information. Then, the pooled pose feature information is compressed to 1 by passing it through a convolutional layer with a 7*7 kernel. The spatial weight coefficients are obtained by passing them through a Sigmoid activation function layer. Finally, the spatial weight coefficients and the output pose feature information are multiplied together to output a convolutional network of pose feature information.
[0105] Optionally, inputting the second pose feature information set into the branch-parallel fusion network to obtain the first multi-scale pose feature information set may include the following steps:
[0106] The first step is to input the second pose feature information set into the branch parallel extension network to obtain the fourth multi-scale pose feature information set.
[0107] The second step involves inputting the aforementioned fourth multi-scale pose feature information set into the aforementioned multiple multi-scale feature extraction convolutional networks to obtain the fifth multi-scale pose feature information set. The input can be the input from a fourth multi-scale pose feature information set to a multi-scale feature extraction convolutional network.
[0108] The third step involves inputting the aforementioned fifth multi-scale pose feature information set into the aforementioned multi-channel spatial attention network to obtain the sixth multi-scale pose feature information set. The input can be a fifth multi-scale pose feature information set input into a single-channel spatial attention network.
[0109] The fourth step involves sampling and fusing the aforementioned sixth multi-scale attitude feature information set to obtain the seventh multi-scale attitude feature information set. This sampling process can include upsampling and downsampling. Specifically, it can involve upsampling and downsampling the attitude feature information output from each branch, and fusing features from each branch. In practice, the executing entity can fuse the attitude feature information obtained from the second branch (after a 2x upsampling process) with the attitude feature information output from the first branch to obtain the seventh multi-scale attitude feature information; and it can also fuse the attitude feature information obtained from the first branch (after a 2x downsampling process) with the attitude feature information output from the second branch to obtain the seventh multi-scale attitude feature information.
[0110] The fifth step involves performing nonlinear processing on the aforementioned seventh multi-scale attitude feature information set to obtain the first multi-scale attitude feature information set. This nonlinear processing can be performed by inputting the seventh multi-scale attitude feature information set into a ReLU activation function.
[0111] Step 306: Update the animal identification information and the set of animal images over time and store them in the target storage.
[0112] In some embodiments, the executing entity may update the animal identification information and the set of animal images over time and store them in a target storage location. The target storage location may be a local database or a cloud database used to store the animal identification information and the set of animal images.
[0113] As an example, the aforementioned executing entity may first determine the image capture time of the aforementioned animal image set. Then, according to the image capture time, the aforementioned animal image set and the historical animal image set are sorted in reverse chronological order and stored in the aforementioned target storage terminal, so as to access the latest animal images in the target storage terminal. Finally, the aforementioned animal identification information is stored in the storage space corresponding to the aforementioned animal image set in the aforementioned target storage terminal.
[0114] In some optional implementations of certain embodiments, after step 306, the above method may further include the following steps:
[0115] In response to determining that the infrared detection information indicates no motion information has been detected, the camera is controlled to enter a standby state. This standby state can be a state in which the camera stops capturing and recognizing images.
[0116] from Figure 3 It can be seen from this that, with Figure 2 Compared to the description of some corresponding embodiments, Figure 3 The image processing methods in some corresponding embodiments, applied to the camera shooting process 300, embody the extended steps of determining the target animal set's animal image set based on the shooting frame rate switching information, and performing animal recognition on the animal image set. Since the schemes described in these embodiments can first control the camera to shoot the target animal set's animal image set under the aforementioned shooting frame rate switching information, it can improve the capture of the target animal's movement details in the animal image set and remove a large amount of repetitive information when the target animal is stationary. Then, animal recognition is performed on the animal image set to obtain animal recognition information. Due to the accurate capture of different movement information of the target animal by the animal image set, the accuracy of animal recognition can be improved, so as to accurately grasp the target animal's behavioral information. Finally, the animal recognition information and the animal image set are updated and stored in the target storage terminal over time. Here, the waste of storage resources can be reduced, and user access and understanding of the target animal's situation can be facilitated. Therefore, it is possible to improve the target recognition accuracy for different shooting frame rates in different scenarios and reduce the waste of storage resources.
[0117] The following is for reference. Figure 6 It illustrates electronic devices suitable for implementing some embodiments of this disclosure (e.g., Figure 1 A schematic diagram of the structure of electronic device 101)600 in the middle. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0118] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0119] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 6 Each box shown can represent a device or multiple devices as needed.
[0120] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined above in the methods of some embodiments of this disclosure.
[0121] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0122] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0123] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a first set of animal images of the target animal set and first shooting frame rate information; perform inter-frame difference processing on the first set of animal images to obtain a motion information set for the first animal in the target animal set; and perform shooting frame rate switching processing based on the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information.
[0124] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0126] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0127] Some embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the above-described image processing methods and is applied to camera capture.
[0128] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. An image processing method applied to camera capture, comprising: Acquire the first set of animal images and the first shooting frame rate information of the target animal set; The first animal image set is subjected to inter-frame difference processing to obtain a motion information set for the first animal in the target animal set; Based on the motion information set of the first animal and the first shooting frame rate information, a shooting frame rate switching process is performed to obtain shooting frame rate switching information; Control the camera to capture a set of animal images of the target animal group under the shooting frame rate switching information; Animal recognition is performed on the animal image set to obtain animal recognition information, including: in response to determining that the animal image set is an image set based on a second preset shooting frame rate, downsampling convolution processing is performed on the animal image set to obtain a first animal pose feature information set; the first animal pose feature information set is input into the feature extraction network included in the animal pose recognition model to obtain a second pose feature information set, wherein the animal pose recognition model further includes: a first branch parallel fusion network, a second branch parallel fusion network, a third branch parallel fusion network, and convolutional layers, the feature extraction network includes: a residual convolutional network and a multi-path fusion convolutional network, and the first branch parallel fusion network includes: a branch... The system comprises a parallel extended network, multiple multi-scale feature extraction convolutional networks, and multiple channel spatial attention networks. The second pose feature information set is input into the first branch parallel fusion network to obtain a first multi-scale pose feature information set. The first multi-scale pose feature information set is input into the second branch parallel fusion network to obtain a second multi-scale pose feature information set. The second multi-scale pose feature information set is input into the third branch parallel fusion network to obtain a third multi-scale pose feature information set. The third multi-scale pose feature information set is input into the convolutional layer to obtain animal pose recognition information, which is used as animal recognition information to determine whether the animal pose recognition information of the target animal set is flight pose information. The animal identification information and the set of animal images are updated and stored in the target storage terminal over time.
2. The method according to claim 1, wherein, The step of performing frame rate switching processing based on the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information includes: The motion information set of the first animal is filtered to obtain the filtered animal state information. Based on the filtered animal status information and the first shooting frame rate information, a shooting frame rate switching process is performed to obtain shooting frame rate switching information.
3. The method according to claim 1, wherein, The step of performing frame rate switching processing based on the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information includes: The motion information set is subjected to intensity information filtering to obtain filtered motion intensity information; In response to determining that the filtered motion intensity information is not greater than a preset motion intensity threshold, the first shooting frame rate information is processed to switch the shooting frame rate to obtain the second shooting frame rate information, which is used as the shooting frame rate switching information.
4. The method according to claim 1, wherein, The step of performing frame rate switching processing based on the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information includes: The motion information set is subjected to intensity information filtering to obtain filtered motion intensity information; In response to determining that the filtered motion intensity information is greater than a preset motion intensity threshold, the first shooting frame rate information is processed to switch the shooting frame rate to obtain the third shooting frame rate information, which is used as the shooting frame rate switching information.
5. The method according to claim 1, wherein, The process of performing animal identification on the set of animal images to obtain animal identification information includes: In response to determining that the set of animal images is an image set based on a first preset shooting frame rate, the set of animal images is preprocessed to obtain a preprocessed set of animal images; The preprocessed animal image set is input into the animal feature extraction network included in the animal recognition model to obtain a first animal feature information set. The animal recognition model further includes a multi-scale fusion network and a head recognition output network. The animal feature extraction network includes multiple hybrid attention networks, multiple first feature extraction networks, multiple max pooling layers, multiple cascaded aggregation networks, and multiple lightweight aggregation networks. The first animal feature information set is input into the multi-scale fusion network to obtain the second animal feature information set. The multi-scale fusion network includes: multiple first feature extraction networks, multiple second feature extraction networks, multiple lightweight aggregation networks, and multiple content-aware prediction networks. The second animal feature information set is input into the head recognition output network to obtain animal recognition information, wherein the head recognition output network includes multiple third feature extraction networks.
6. The method according to claim 1, wherein, The step of inputting the second pose feature information set into the first branch parallel fusion network to obtain the first multi-scale pose feature information set includes: The second pose feature information set is input into the branch parallel extension network to obtain the fourth multi-scale pose feature information set; The fourth multi-scale pose feature information set is input into the multiple multi-scale feature extraction convolutional networks to obtain the fifth multi-scale pose feature information set. The fifth multi-scale pose feature information set is input into the multi-channel spatial attention network to obtain the sixth multi-scale pose feature information set. The sixth multi-scale pose feature information set is sampled and fused to obtain the seventh multi-scale pose feature information set; The seventh multi-scale attitude feature information set is subjected to nonlinear processing to obtain the first multi-scale attitude feature information set.
7. The method according to claim 1, wherein, Before acquiring the first animal image set and the first shooting frame rate information of the target animal set, the method further includes: In response to the detection of infrared signal information sent by a passive infrared sensor, the infrared signal information is processed to obtain infrared detection information; Based on the infrared detection information, the camera is controlled to capture animal images to obtain a first set of animal images of the target animal set within the shooting area.
8. The method according to claim 7, wherein, The method further includes: In response to determining that the infrared detection information indicates that no motion information has been detected, the camera is controlled to enter a standby state.
9. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.
10. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Photographing Apparatus, And Method For Photographing Moving Object With The Same
CN107205115A
Small target detection method for images acquired by unmanned aerial vehicle based on improved YOLOv8 algorithm
CN118628939A
Video camera and imaging method
JP2012023634A